AI Blitz on Bitcoin Code: 4,962 Security Findings Across 390 Projects in ~30 Hours

A volunteer crew used AI agents to scan 390 Bitcoin projects, surfacing 4,962 issues—720 high or critical—in about 30 hours. Maintainers now face triage at machine speed.

Bitcoin
Cryptocurrency
Regulations
Economy
Because Bitcoin
Because Bitcoin

Because Bitcoin

August 6, 2026

Bitcoin’s codebase just met automated scale. A volunteer collective, the Bitcoin Red Team, pointed AI agents at the ecosystem and recorded 4,962 security findings across 390 projects in roughly 30 hours—an average rate of 166 issues per hour. In the first situation report published Wednesday by Cashu creator calle, 85 issues were labeled critical and 635 high, putting 14.5% of the corpus in the top two tiers and averaging 1.85 serious findings per project.

Here’s the operating model: 16 contributors working around the clock—17 logged in the report, with 14 human reviewers and three automated agents—each prompting their own toolchains. That diversity of prompts and workflows, calle said, is producing broader coverage than a single scanning method. While much of the work is still “hand-holding the AI,” 91% of findings arrived through automated intake, and about 21% have been dynamically reproduced with proof-of-concept code.

The hit rate wasn’t uniform. Privacy and coinjoin tools carried the highest share of high-or-critical outcomes at 24%, followed by swaps and exchanges at 21% and payments or merchant tools at 17%. Cryptographic libraries and SDKs yielded the largest raw count—1,101 findings—but only 10% reached high or critical, a pattern that often suggests noisy static flags mixed with genuine edge-case risks.

My focus: throughput versus trust. AI gives defense massive surface coverage, but credibility hinges on validation speed. Only 19 projects—under 5% of those reviewed—have had findings disclosed upstream so far, and eight reports have already been retired as false positives. With 91% of items sourced from automated scans and just 21% re-run to POC, maintainers are staring at a classic alert-fatigue trap. Teams that feel buried often under-triage, not because they don’t care, but because context switching and reproductions kill hours they don’t have.

There’s a better pattern emerging: - Prioritize exploitability over elegance. Bucket by blast radius, ease of weaponization, and user exposure—not by rule severity alone. - Standardize intake. Ship SARIF or similar machine-readable outputs, deduplicate across agents, and attach minimal POCs or payloads for anything labeled high or critical. - Time-bound coordination. Rapid disclosure can be fair when “anyone running the same tools will find the same bugs,” as calle argued, but agree on short embargo windows where exploitability is clear and patching is feasible. - Measure quality. Track false positives explicitly, raise the bar for “critical,” and use ensemble LLM reviews to challenge each high-sev before filing.

Why this matters: attackers are already here. The Coldcard incident shows what lag looks like. A March 2021 firmware build reportedly fell back to software randomness instead of the hardware RNG, leaving wallet seeds guessable. Users lost an estimated $130 million. The flaw sat in public code for more than five years before an adversary likely used AI to surface it. As Ledger CTO Charles Guillemet put it this week, vulnerabilities are now identified “at machine speed,” and “open source and reviewed are not the same thing.”

Seen through that lens, the Bitcoin Red Team is directionally right: flood the zone with competent automation, then compress validation. But the trust contract with maintainers matters. If contributors keep volume high while sharpening reproducibility, deduplication, and exploit-first ranking, we get the upside of AI without burning the humans who must ship the fixes. Defense won’t win by being quieter; it wins by being faster—and clearer—than the offense.