AI-assisted security campaign focused on the Bitcoin ecosystem, Bitcoin Red Team, said it generated 6,700 findings across 425 projects in its first 55 hours. The campaign labeled 1,029 of them high or critical.
The Aug. 6 update measures how much material entered a security triage pipeline, and its effect on software security remains unreported.
The retrieved thread omitted audit-ready definitions and denominators for the severity counts, as well as case-level outcomes, an aggregate false-positive rate, and a fix rate.
Those missing fields prevent a calculation of how many alerts became confirmed vulnerabilities, how many maintainers rejected or downgraded, and how many led to patches.
The first 55 hours still reveal a consequential capability, noting how AI systems can fill an ecosystem-scale review pipeline quickly. Expert prompting, reproduction, disclosure, and maintainer response remained necessary at every later stage.
The campaign published two snapshots as its roster and workload expanded:
| Elapsed time | Projects | Total findings | Reported severity | Participants |
|---|---|---|---|---|
| 27.5 hours | 390 | 4,962 | 85 critical; 635 high | 16 |
| 55 hours | 425 | 6,700 | 1,029 high or critical | 24 reported, including three bots |
The 27.5-hour update covered 390 projects and 4,962 findings. By the 55-hour mark, the project count had risen by 35 and the finding count by 1,738. The later thread put high-or-critical findings at 15.4% of the total and clarified that three of the 24 reported participants were bots.
The earlier post separated critical and high findings, while the later one combined them, with both sets of figures reflecting campaign assessments. Maintainer-confirmed exploitability and remediation outcomes require separate evidence.
Bitcoin Red Team scanned 425 projects and reported 6,700 findings, including 1,029 high or critical issues, while public validation rates remain unpublished.
Rob Hamilton described Kimi K3 as handling the heavy analysis, with GPT Sol, Fable/Opus, and GLM 5.2 supporting the documentation. He said OpenAI's Cyber Harness covered selected components he considered load-bearing.
A day later, Hamilton wrote that subject-matter experts could change an assessment with one or two sentences of context or a small block of code. In examples he described, that input pushed middling concerns into high or critical territory. He also identified operations, disclosure handoff, and triage as bottlenecks.
In Hamilton's account, models searched broadly while specialists shaped prompts, interpreted output, attempted reproduction, and decided which reports were ready for disclosure. That division of labor makes the campaign a human-AI review system.
The developer known as Calle said most critical reports were quickly verified by project owners. The post supplied no denominator, verified-report count, rejection count, or patch status, leaving the breadth and outcome of that verification unresolved.
In the 55-hour update, Bitcoin Red Team reported that 19.5% of scanned projects had a SECURITY.md file and 13.1% had an email there. The retrieved thread omitted the project corpus, denominator interpretation, and measurement method, so the percentages only describe the campaign's scan.
On Aug. 3, Hamilton said the effort had spent over $10,000 scanning over 100 repositories and had immediately disclosed critical findings when a proof of concept demonstrated exploitability. On Aug. 4, he reported about $20,000 in spending, more than a dozen disclosures and 150 repositories scanned.
Scanning continued to expand, while the campaign described outreach, handoff and triage as active operational constraints. The published snapshots offer no comparable disclosure denominator at 55 hours, so they cannot establish the relative speed of scanning and resolution.
Hamilton later identified the separate Coldcard incident as a catalyst for the wider campaign. The campaign record attributes no discovery of the Coldcard flaw to this sprint.
A useful public accounting would separate findings that were reproduced, acknowledged, downgraded, rejected, and fixed, with definitions and denominators for each rate. That breakdown would show how much of the campaign's volume became actionable security work.
A public critic, JW Weatherman, argued that the campaign could not triage its output. His post identified no campaign-linked issue, patch, or advisory, so it supplies criticism without a measurable failure rate. The campaign's missing disposition data leaves the underlying question open.
For now, 6,700 represents campaign-labeled findings and triage candidates. The sprint demonstrated the speed of machine-assisted review. Its lasting security value depends on the share that experts can validate, disclose, and convert into fixes.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.





























