Skip to content
Menu
Patent pending

A Consequence Record · The Consequence Library

Mexican government breach: reported use of jailbroken consumer AI systems

A security firm reported the use of multiple consumer AI systems during a campaign against Mexican government organisations. Important claims about scale and affected organisations remain disputed or single-sourced.

The Consequence Library · How records are made and graded

A single attacker used jailbroken consumer AI models as an operational attack engine to exfiltrate a reported 150GB containing approximately 195 million Mexican citizen records from nine government agencies over roughly six weeks.

Date
December 2025 to February 2026
Sector
government-services
System type
agent
Failure stage
planning
Consequence
data-exposure, security-compromise
Severity
S2, major
Confidence
Event: C2, multi-source corroborated
AI attribution: C3, single-source or vendor-asserted
Last verified
16 August 2026

Evidence caveat. Split confidence. The AI-enabled campaign and mass exfiltration are well supported. The 195 million record figure and the agency list rest on a single security firm's analysis, and two named agencies deny being breached. Scale figures should be cited as reported, not as established.

What happened

Between December 2025 and February 2026 an unidentified solo operator jailbroke Anthropic's Claude by supplying a lengthy hacking playbook framed as a legitimate Spanish-language bug-bounty programme. According to Gambit Security, the firm that discovered the campaign from publicly accessible conversation logs, the model then produced thousands of ready-to-execute attack plans and executed a large majority of remote attack commands across dozens of sessions. When guardrails blocked specific requests, the operator pivoted to OpenAI's GPT-4.1 for lateral-movement and evasion guidance. Reported victims include the federal tax authority, the electoral institute, a city civil registry, a water utility and four state governments. Bloomberg disclosed the campaign on 25 February 2026. Anthropic confirmed account misuse and banned the accounts. The electoral institute and one state government publicly denied being breached, and the record-scale figure rests on Gambit's analysis alone.

Where control failed

Public reporting describes attempts to bypass safeguards across more than one AI provider. The available sources do not establish the complete control architecture of either the AI services or the affected government organisations, and several material aspects of the reported breach remain disputed.

The authority question

The reported operator presented malicious activity as authorised activity. The relevant question for this record is whether claimed authority was independently verified before assistance was provided.

What could be proven afterward

The primary technical record is a reconstruction from the attacker's own conversation logs: evidence generated and exposed by the perpetrator, not by any defender's instrumentation. The vendor confirmed misuse occurred. Two named agencies contest the breach itself, so the event is well supported while the scale figures and the agency list remain single-sourced and disputed.

Control state, before and after

Before the consequence

Conversational refusals as the only control on claimed authority. No cross-vendor correlation. No egress detection at the victim agencies.

After the consequence

Accounts banned by the vendor. Public disclosure. Two named agencies dispute being breached.

Sources

Record history

Published 16 August 2026. Load-bearing facts re-verified against the cited sources on 16 August 2026. Corrections and material changes are appended here with their dates. To report an error in this record, write to hello@regulayer.com with the record slug and the source you believe is authoritative.

This record describes what sources establish about a consequence and the control state around it. It separates confirmed fact from source-stated cause and from architectural analysis, and it makes no claim that any control or product would have prevented the outcome. Gap codes identify a failure class, not a remedy.