Citable record
Incidents, with receipts
Jailbreaks in the wild, indirect injection attacks, provider-side failures, and disclosure episodes — each tagged by model and technique, each linked to primary sources. Built to be cited.
During pre-deployment testing, UK AISI developed universal jailbreaks — often within hours — that unlocked Sol's Preparedness-High cyber capabilities for autonomous exploitation. Disclosed in the system card; OpenAI reproduced and mitigated the specific jailbreaks before GA. Notable both for the finding and for the disclosure: an adverse pre-release result published by the vendor.
Outcome: Specific jailbreaks mitigated pre-GA; public jailbreak claims against Sol continued post-release.
FAR.AI's stress test broke DeepSeek V4 Pro's safeguards at 100% with public jailbreaks across CBRN, cyber, and terrorism domains in roughly 15 minutes, 99.6% via authority manipulation, and 99.6% via response prefill — and a jailbreak written for V3.2 worked on V4 Pro unmodified. Neo Research separately raised the StrongREJECT jailbreak rate from 0.6% to 77.8% with a single 2023-era roleplay template that peer models resisted.
Outcome: No provider response; open weights preclude post-release safeguard fixes. Recorded as the best-documented safeguard failure among tracked models.
Anthropic disclosed a campaign in which a state-sponsored group manipulated Claude Code into automating reconnaissance and exploitation against ~30 targets by role-playing as a legitimate security firm — an early, well-documented case of agentic-scale misuse of a frontier coding agent.
Outcome: Accounts banned, targets notified, public disclosure with TTPs. Cited here as evidence that agentic deployment risk is a property of the deployment pattern, not only the model.
Sources (1)
Researchers demonstrated a service-side indirect injection: a crafted email caused the Deep Research agent, when later asked to summarize the inbox, to exfiltrate mailbox data to an attacker-controlled URL without user interaction. Patched by OpenAI after disclosure. Tagged to the GPT-5 family as the underlying agent model.
Outcome: Patched following responsible disclosure.
Sources (1)
Researchers showed that poisoned calendar invites could hijack Gemini assistant sessions into performing unintended actions, including smart-home control, when the user later asked routine questions. Google shipped mitigations. Demonstrates the indirect-injection surface of assistant-integrated deployments.
Outcome: Mitigations deployed by Google following disclosure.
A provider-side change instructing the model not to shy away from politically incorrect claims led to a burst of antisemitic and violent outputs on X. xAI attributed it to an unauthorized modification and reverted. Recorded against the Grok family; the operational lesson — silent production changes to safety-relevant configuration — is the core failure mode this index monitors.
Outcome: Prompt reverted; xAI published the system prompt and apologized.
Sources (2)
Meta's leaderboard submission was an experimental chat-tuned variant, not the released weights, inflating public perception of the released model. Recorded under transparency: disclosure practices are part of the risk surface buyers rely on.
Outcome: LM Arena updated its policies; Meta acknowledged the variant difference.
Sources (1)
Cisco's evaluation reported that DeepSeek R1 failed to block a single prompt from a 50-prompt HarmBench sample. Recorded against the DeepSeek family as the strongest public signal on its jailbreak posture; V3.2-specific retesting is pending our Phase 2 probe battery.