Light bulb Limited Spots Available: Secure Your Lifetime Subscription on Gumroad!

Oct 04, 2026 · 8 min read

543,699 Exposed GitHub Credentials Still Work After Years

Truffle Security scanned The Stack v3, a snapshot of 224 million public GitHub repositories assembled to train AI models, then tested every candidate key against its issuer on July 27 and 28, 2026. Over half a million still logged in. The median one had been public for 784 days.

A secret committed to a public repo is not a mistake you can quietly undo. It gets forked, mirrored, and eventually swept into the datasets that teach language models to write code. Truffle Security went looking in one of those datasets and published the results on September 29, 2026 under a blunt headline: GitHub Repos Exposed 543,699 Credentials. Nobody Revoked Them.

The size of the pile is not the story. Which keys survived is, and that split has little to do with GitHub's push protection and everything to do with whether the issuer bothers to kill a leaked key.

Key Takeaways

  • Truffle Security found 543,699 unique credentials that still authenticated in July 2026, drawn from 1,103,438 separate exposures across The Stack v3's 224,553,295 repositories and 58,467,468,698 files.
  • The Stack v3's crawl closed on August 7, 2025, so the keys were verified live eleven months later, and the oldest dated back to a file last touched on June 13, 2009.
  • 199,843 of the live credentials were committed after GitHub turned push protection on by default on February 29, 2024, and 51.8 percent of all live credentials are connection strings, Google API keys or private keys that default push protection does not block.
  • Survival depended on revocation: 1 of 101,886 committed npm tokens still worked, against 11,465 of 12,985 Postgres connection strings and 69,041 of 126,963 Google Cloud service account credentials.
  • SendGrid, the only email service named in the data, had 22,800 keys committed and 9,189 still live, even though SendGrid is a GitHub secret scanning partner.
A developer's monitor at night showing blurred lines of code, with a brass key resting on the keyboard in a dim office

What Did Truffle Security Find in The Stack v3?

Truffle found 543,699 live, unique credentials in The Stack v3, a corpus its researchers describe as "a snapshot of public code assembled to train AI models." They walked all 4,096 metadata shards, matched candidate secrets by pattern, and then asked each issuing service whether the credential still worked. Only the ones that answered yes made the count.

The Stack v3 keeps only each repository's default branch as it stood at crawl time, with no commit history. So Truffle dated each leak by the last modification time of the file holding it, taking the earliest date across every copy. A file edited after the key landed reads as newer, which makes the ages conservative.

Even so, the ages are ugly:

  • 784 days median age at verification, with the 90th percentile at 6.3 years.
  • June 13, 2009 for the oldest: database credentials in an Erlang web server config, still valid 16.1 years later.
  • 2,636 live credentials sit in files last modified before 2015.

Density keeps rising too. Truffle measured 3.72 live credentials per million files in 2014, 9.54 in 2022, 11.09 in 2023 and 11.62 in 2025, the highest in the dataset. Because the method only sees default branches, anything force pushed away or reverted before August 2025 is invisible. Truffle's own conclusion: "The real population is larger."

Why Does It Matter That This Is AI Training Data?

It matters because training corpora copy secrets wholesale and keep them long after the source repo is cleaned up. Once a key is in a dataset like The Stack v3, deleting it from GitHub does not delete it from every downloaded copy of the corpus.

This is not Truffle's first scan of AI training data. In February 2025 it scanned the December 2024 Common Crawl archive, 400TB of web data from 2.67 billion pages, and found 11,908 live secrets. That post made the modeling point directly: "LLMs can't distinguish between valid and invalid secrets during training, so both contribute equally to providing insecure code examples." Later, a scan of 7.6 petabytes of Hugging Face training data turned up 221,303 working credentials. The Stack v3 result is more than double that.

Put the three side by side and the trend line is the story. In 19 months, across three AI datasets, Truffle's live count went from about 12,000 to about 221,000 to 543,699. Each of those keys is also a code sample showing a model that a password belongs in a config file.

Why Didn't GitHub Push Protection Stop Them?

Push protection only blocks secret types it can recognize with confidence, and most of what survives in this corpus falls outside that list. GitHub made secret scanning alerts free for public repositories on February 28, 2023, and turned push protection on by default for public repos on February 29, 2024. Truffle split the live keys along those dates: 245,959 predate free alerts, 97,897 arrived while protection was opt in, and 199,843 landed after it became the default.

Where it applies, the block works. Across the twelve months either side of the rollout, the leak rate for protected secret types fell 53 percent while unprotected types fell 7 percent. Slack tokens dropped 64 percent, GitHub tokens and AWS access keys 59 percent each, SendGrid 56 percent.

The gap is in what GitHub chooses not to block. Per GitHub's supported secret scanning patterns, Google API keys are not push protected. Database connection strings and private keys count as generic patterns, which stay off unless someone enables them. Together those shapes make up 51.8 percent of every live credential Truffle found.

Google keys show why. A Gemini key and a Maps key share the same AIzaSy prefix, but only one of them is a billable secret. One pattern cannot tell them apart, so GitHub blocks neither. Truffle counted 31,374 live Gemini keys with a median leak date of February 2025, every one of them younger than push protection.

Which Leaked Keys Are Still Live?

The keys still live are the ones no provider kills automatically. Truffle's survival figures, from its primary research post, run from 0.001 percent to 88 percent inside one corpus:

Credential type Committed Still live Survival
npm tokens 101,886 1 0.001%
GitHub tokens 73,048 260 0.36%
Stripe keys (includes test mode) 124,132 4,493 4%
AWS access keys 82,411 6,819 8%
SendGrid keys 22,800 9,189 40%
Google Cloud service accounts 126,963 69,041 54%
MySQL connection strings 2,421 1,806 75%
Postgres connection strings 12,985 11,465 88%

Truffle's verdict fits in two sentences: "GitHub and npm revoke their own tokens and almost nothing survives. Nobody revokes a Postgres URL, and almost everything does." MongoDB is left out: its detector only reports URIs it managed to connect to, so all 51,067 found were live and no survival rate is possible.

The GitHub secret scanning partner program forwards a leaked token to its issuer, but Truffle notes it "does not require partners to revoke anything." That is the gap. We saw the same pattern in August, when a separate Truffle study found 9,300 leaked AWS keys that still worked, and in the Stripe vendor leak of 1,033 live API keys.

What This Means for Your Inbox

For email, the number that matters is 9,189: live SendGrid API keys sitting in public code. SendGrid is the only email provider Truffle's post names. It does not break out Mailgun, Postmark, Amazon SES or raw SMTP passwords, so any of those in the corpus are either folded into other categories or not counted.

Here is the counterintuitive part. SendGrid has been a GitHub secret scanning partner since June 20, 2022, its keys are push protected, and its exposed API key support page says it "may automatically delete your exposed API key." Yet 40 percent of committed SendGrid keys still worked, a survival rate five times AWS's and roughly 110 times GitHub's own tokens. For SendGrid keys leaked in 2025 alone, 13 percent were still live. "May" is doing a lot of work in that sentence.

SendGrid spells out the risk on the same page: exposed keys "can be used to send malicious email, access personal information, acquire the information of your recipients." A key with send permission mails from the owner's account, and if that account has domain authentication set up, the SPF and DKIM records vouch for it. The phishing email you receive does not come from a lookalike domain. It comes from the real one. We documented the same play with leaked cloud keys in attackers mining GitHub for Amazon SES credentials.

A passing authentication check tells you which server sent a message, not who held the key. For developers, a SendGrid key in a repo is a key to your customers' inboxes.

How Do You Find and Kill Leaked Secrets in Your Repos?

Rotate first and clean up second, because a committed key is compromised the moment it lands. Truffle's own advice: "Treat a committed credential as burned the moment it lands, whether or not anything flagged it." In practice:

  1. Rotate before you rewrite history. Revoke the old key at the provider and issue a new one. Scrubbing Git history does nothing for the forks, mirrors and dataset copies already made.
  2. Scan your full history, not just new pushes. Anything committed before February 2024 was never in scope for default push protection. The open source TruffleHog checks findings against the issuer: trufflehog git file://. --results=verified,unknown. Gitleaks runs offline with gitleaks git -v . and also ships as a pre commit hook.
  3. Turn on generic patterns. In your repo, go to Settings, then Advanced Security, and under Secret Protection enable Generic patterns, per GitHub's guide. That is the bucket connection strings and private keys fall into.
  4. Prefer credentials that expire. Truffle found that "most of what we found would have aged out harmlessly if it had ever been given a lifetime." Use short lived tokens and workload identity over static keys wherever the provider offers them.
  5. Scope email keys tightly. SendGrid's API key documentation offers Custom Access keys in place of Full Access. A key that can only send cannot export your contact lists.
  6. Keep secrets out of the repo entirely. Load them from a secret manager at runtime, and add .env to .gitignore before the first commit.
  7. Check whether your provider revokes. If it does not, your rotation process is the only kill switch.

Push protection locks the front door. The half million keys already inside only die when someone rotates them.

Stop Email Tracking in Gmail

Spy pixels track when you open emails, where you are, and what device you use. Gblock blocks them automatically.

Try Gblock Free for 30 Days

No credit card required. Works with Chrome, Edge, Brave, and Arc.