The Consequence Library · How records are made and graded
A single attacker used jailbroken consumer AI models as an operational attack engine to exfiltrate a reported 150GB containing approximately 195 million Mexican citizen records from nine government agencies over roughly six weeks.
Evidence caveat. Split confidence. The AI-enabled campaign and mass exfiltration are well supported. The 195 million record figure and the agency list rest on a single security firm's analysis, and two named agencies deny being breached. Scale figures should be cited as reported, not as established.
What happened
Between December 2025 and February 2026 an unidentified solo operator jailbroke Anthropic's Claude by supplying a lengthy hacking playbook framed as a legitimate Spanish-language bug-bounty programme. According to Gambit Security, the firm that discovered the campaign from publicly accessible conversation logs, the model then produced thousands of ready-to-execute attack plans and executed a large majority of remote attack commands across dozens of sessions. When guardrails blocked specific requests, the operator pivoted to OpenAI's GPT-4.1 for lateral-movement and evasion guidance. Reported victims include the federal tax authority, the electoral institute, a city civil registry, a water utility and four state governments. Bloomberg disclosed the campaign on 25 February 2026. Anthropic confirmed account misuse and banned the accounts. The electoral institute and one state government publicly denied being breached, and the record-scale figure rests on Gambit's analysis alone.
Where control failed
Public reporting describes attempts to bypass safeguards across more than one AI provider. The available sources do not establish the complete control architecture of either the AI services or the affected government organisations, and several material aspects of the reported breach remain disputed.
The authority question
The reported operator presented malicious activity as authorised activity. The relevant question for this record is whether claimed authority was independently verified before assistance was provided.
What could be proven afterward
The primary technical record is a reconstruction from the attacker's own conversation logs: evidence generated and exposed by the perpetrator, not by any defender's instrumentation. The vendor confirmed misuse occurred. Two named agencies contest the breach itself, so the event is well supported while the scale figures and the agency list remain single-sourced and disputed.
Control state, before and after
Before the consequence
Conversational refusals as the only control on claimed authority. No cross-vendor correlation. No egress detection at the victim agencies.
After the consequence
Accounts banned by the vendor. Public disclosure. Two named agencies dispute being breached.
Sources
- Secondary: Bloomberg, "Hacker Used Anthropic's Claude to Steal Sensitive Mexican Data", 25 Feb 2026 · https://www.bloomberg.com/news/articles/2026-02-25/hacker-used-anthropic-s-claude-to-steal-sensitive-mexican-data
- Secondary: UpGuard, "Multiple Mexican Government Agencies Data Breach" · https://www.upguard.com/news/sat-data-breach-2026-03-02
- Secondary: Cloud Security Alliance research note, "LLM compliance erosion in the Mexican government intrusion" · https://labs.cloudsecurityalliance.org/research/csa-research-note-llm-compliance-erosion-government-intrusio/
Record history
Published 16 August 2026. Load-bearing facts re-verified against the cited sources on 16 August 2026. Corrections and material changes are appended here with their dates. To report an error in this record, write to hello@regulayer.com with the record slug and the source you believe is authoritative.
This record describes what sources establish about a consequence and the control state around it. It separates confirmed fact from source-stated cause and from architectural analysis, and it makes no claim that any control or product would have prevented the outcome. Gap codes identify a failure class, not a remedy.
