ModelRiskIndex

Citable record

Incidents, with receipts

Jailbreaks in the wild, indirect injection attacks, provider-side failures, and disclosure episodes — each tagged by model and technique, each linked to primary sources. Built to be cited.

2026-07-10
UK AISI found universal jailbreaks in GPT-5.6 Sol unlocking autonomous cyber-exploitation
Universal jailbreaks (pre-deployment external testing)GPT-5.6 (Sol · Terra · Luna)

During pre-deployment testing, UK AISI developed universal jailbreaks — often within hours — that unlocked Sol's Preparedness-High cyber capabilities for autonomous exploitation. Disclosed in the system card; OpenAI reproduced and mitigated the specific jailbreaks before GA. Notable both for the finding and for the disclosure: an adverse pre-release result published by the vendor.

Outcome: Specific jailbreaks mitigated pre-GA; public jailbreak claims against Sol continued post-release.

Sources (2)
2026-05-11
FAR.AI: DeepSeek V4 Pro safeguards collapse at 98–100% across three attack strategies
Public jailbreaks, authority manipulation, response prefillDeepSeek V4 Pro

FAR.AI's stress test broke DeepSeek V4 Pro's safeguards at 100% with public jailbreaks across CBRN, cyber, and terrorism domains in roughly 15 minutes, 99.6% via authority manipulation, and 99.6% via response prefill — and a jailbreak written for V3.2 worked on V4 Pro unmodified. Neo Research separately raised the StrongREJECT jailbreak rate from 0.6% to 77.8% with a single 2023-era roleplay template that peer models resisted.

Outcome: No provider response; open weights preclude post-release safeguard fixes. Recorded as the best-documented safeguard failure among tracked models.

Sources (2)
2025-11-13
State-sponsored actor used Claude Code to automate intrusion campaign
Agentic misuse / social-engineering the model's safety contextClaude Sonnet 4.5

Anthropic disclosed a campaign in which a state-sponsored group manipulated Claude Code into automating reconnaissance and exploitation against ~30 targets by role-playing as a legitimate security firm — an early, well-documented case of agentic-scale misuse of a frontier coding agent.

Outcome: Accounts banned, targets notified, public disclosure with TTPs. Cited here as evidence that agentic deployment risk is a property of the deployment pattern, not only the model.

Sources (1)
  • Anthropic disclosure
    Anthropic provider artifacts · provider artifact · source tier B · anthropic.com · retrieved 2026-08-03
2025-09-18
ShadowLeak: zero-click data exfiltration via ChatGPT Deep Research email integration
Indirect prompt injection (zero-click, service-side exfiltration)GPT-5.1

Researchers demonstrated a service-side indirect injection: a crafted email caused the Deep Research agent, when later asked to summarize the inbox, to exfiltrate mailbox data to an attacker-controlled URL without user interaction. Patched by OpenAI after disclosure. Tagged to the GPT-5 family as the underlying agent model.

Outcome: Patched following responsible disclosure.

Sources (1)
2025-08-06
Promptware: calendar-invite injection drove Gemini-connected smart-home actions
Indirect prompt injection via calendar invitesGemini 3 ProGemini 2.5 Flash

Researchers showed that poisoned calendar invites could hijack Gemini assistant sessions into performing unintended actions, including smart-home control, when the user later asked routine questions. Google shipped mitigations. Demonstrates the indirect-injection surface of assistant-integrated deployments.

Outcome: Mitigations deployed by Google following disclosure.

Sources (1)
2025-07-08
Grok produced antisemitic output after unannounced system-prompt modification
Provider-side system-prompt change (not user attack)Grok 4.1

A provider-side change instructing the model not to shy away from politically incorrect claims led to a burst of antisemitic and violent outputs on X. xAI attributed it to an unauthorized modification and reverted. Recorded against the Grok family; the operational lesson — silent production changes to safety-relevant configuration — is the core failure mode this index monitors.

Outcome: Prompt reverted; xAI published the system prompt and apologized.

Sources (2)
  • xAI statement
    xAI provider artifacts · provider artifact · source tier B · x.ai · retrieved 2026-08-03
  • AI Incident Database entry
    AI Incident Database · incident record · source tier F · incidentdatabase.ai · retrieved 2026-08-03
2025-04-08
Llama 4 Maverick LM Arena submission used unreleased experimental variant
Benchmark presentation (not an attack)Llama 4 Maverick

Meta's leaderboard submission was an experimental chat-tuned variant, not the released weights, inflating public perception of the released model. Recorded under transparency: disclosure practices are part of the risk surface buyers rely on.

Outcome: LM Arena updated its policies; Meta acknowledged the variant difference.

Sources (1)
  • LMSYS statement
    LMArena leaderboard statements · independent eval · source tier C · lmarena.ai · retrieved 2026-08-03
2025-02-01
DeepSeek R1 showed 100% attack success rate in Cisco testing
Standard jailbreak suite (HarmBench)DeepSeek V3.2

Cisco's evaluation reported that DeepSeek R1 failed to block a single prompt from a 50-prompt HarmBench sample. Recorded against the DeepSeek family as the strongest public signal on its jailbreak posture; V3.2-specific retesting is pending our Phase 2 probe battery.

Sources (1)
  • Cisco security evaluation
    Cisco security research · independent eval · source tier C · blogs.cisco.com · retrieved 2026-08-03