- GitHub
- GitLab
- privacy
- AI
- data governance
GitHub training models on repos without explicit opt-out. Why we migrated to GitLab

On March 25, 2026, GitHub updated its Privacy Statement and Terms of Service. The core change: interaction data from Copilot Free, Pro and Pro+ users — inputs, outputs, code snippets and the surrounding context — will be used to train GitHub's AI models by default, starting April 24, 2026. If you did not act before that date, your code is already in the pipeline. This post explains exactly what changed, how to check your current setting, and why we concluded that GitLab is the right platform for client work.
What exactly does GitHub collect — and what do they use it for?
GitHub is careful with language here. The policy covers interaction data: what you type into Copilot, what Copilot returns, and the code context it used to generate the response. GitHub states that code stored at rest in private repositories is not used for training. The catch: whenever Copilot is active, it processes your private repository code as interaction context. That session data falls under the interaction policy, not the at-rest policy.
For a team where every developer has Copilot open while working on client code, this is a meaningful distinction that disappears in practice.
Who is affected
- Copilot Free, Pro and Pro+ users — affected by default from April 24
- Copilot Business and Copilot Enterprise — exempt under existing contract terms
- Students and teachers using GitHub Education — also exempt
Most individual developers and small teams sit on Free, Pro or Pro+. The enterprise exemption is commercially sensible for GitHub — large clients negotiate data terms. But it also means the exemption is unavailable to the majority of open-source contributors, freelancers, and small agencies.
How to opt out (if you haven't already)
Navigate to github.com/settings/copilot/features. Under the Privacy section, find "Allow GitHub to use my data for AI model training" and disable it. The toggle defaults to on — you must explicitly turn it off. This covers your personal account only; organisation administrators control the setting for org members separately.
Why the opt-out model is the wrong default
Regulated industries — finance, healthcare, defence, public sector — operate with a clear principle: data does not go anywhere without explicit authorisation. The same applies to client work under NDA. "May be used for training unless you opt out" is not a framework compatible with those requirements. It places the compliance burden on the developer, not on the platform.
GitLab's own analysis called this a "governance wake-up call" — and they are not wrong. The technical opt-out exists. But relying on every developer in your organisation finding the correct settings page before a deadline is not a governance control. It is hoping nothing goes wrong.
GitLab's position: opt-in AI, zero training by default
GitLab does not train AI models on customer code at any tier. AI features are opt-in, not opt-out. Critically, GitLab has a zero-day data retention policy with its AI infrastructure partners (Fireworks AI, AWS, Google): vendor input and output data is discarded immediately after the response is delivered. No storage, no abuse monitoring logs, no training pipeline.
For self-hosted GitLab instances, the position is even cleaner: GitLab Inc. has no access to your repositories at all. This is not a policy commitment — it is an architectural fact.
What the migration actually involved
We moved all client project repositories to GitLab. Internal tooling went to a self-hosted GitLab CE instance on our own infrastructure. The migration took approximately three working days, the majority of which was rewriting CI pipeline files from GitHub Actions syntax to GitLab CI YAML.
Functional parity is close:
- GitLab CI/CD replaces GitHub Actions — feature-equivalent for our build, test and deploy workflows
- GitLab Container Registry replaces GitHub Container Registry
- GitLab Issues and Milestones replace GitHub Issues and Projects
- Vercel deploy integration works unchanged via webhooks and the Vercel CLI
Is this an overreaction?
Perhaps. GitHub has not published training set provenance data, so it is impossible to know whether any specific code was actually used. But "may be used" in a privacy policy means you have no assurance it was not. For personal projects, the risk calculus is yours to make. For client code under NDA — or code that belongs to a client's infrastructure — we do not work with assurances. We work with controls. GitLab gives us controls where GitHub gives us a hope.

