ModelRiskIndex

Rankings / Google

Gemini 3 Pro

Tier 270/100retired

gemini-3-pro-preview via generativelanguage.googleapis.com (retired 2026-03-09; ID now aliases gemini-3.1-pro-preview)

retired. gemini-3-pro-preview was shut down 2026-03-09 — roughly four months after launch, never reaching a GA/stable ID — and the model ID was silently aliased to gemini-3.1-pro-preview. Requests to the old ID now reach a different model. Successor: Gemini 3.1 Pro (gemini-3.1-pro-preview, released 2026-02-19).

Re-verified against live sources 2026-08-05. This assessment covers a retired model kept published as history: its ID now silently serves Gemini 3.1 Pro Preview, tracked separately as gemini-3-1-pro. Grades assess the Gemini API / Vertex deployment; consumer Gemini Apps have materially different data-handling defaults.

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Customer data is not used for training by default, or a documented opt-out exists.
Tier 2 requirements
  • External pre-deployment testing. Independent external parties tested the model before deployment, and this is disclosed.External tester names (UK AISI, Apollo Research, Vaultis, Dreadnode) come from Google's release materials and secondary reporting; the archived model card PDF does not name them, so this requirement is satisfied on weaker evidence than the rest of the entry and needs a better primary citation.
  • Third-party certification. The operating organization holds verifiable third-party certification (e.g. SOC 2, ISO/IEC 42001).
  • Versioning with changelogs. Model versions are explicitly identified and changes are changelogged.Changelog and lifecycle pages are real, but the 3-pro-preview shutdown-and-alias episode shows preview IDs offer no version stability.
  • Stated deprecation policy. A deprecation policy with notice windows is published.
Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancepartial

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Paid API tier is not used for product improvement (abuse logging only, 55 days); unpaid tier content is used for improvement including human review (EEA/CH/UK get paid-tier terms); consumer Gemini Apps default to data use unless disabled. A zero-data-retention option now exists by approval.

Receipts (3)
  • Gemini API terms of service
    Google / DeepMind provider artifacts · provider artifact · source tier B · ai.google.dev · retrieved 2026-08-05
    To help with quality and improve our products, human reviewers may read, annotate, and process your API input and output.
  • Gemini API zero-data-retention documentation
    Google / DeepMind provider artifacts · provider artifact · source tier B · ai.google.dev · retrieved 2026-08-05
    When your request for ZDR for a particular project is approved, all user content (prompts and responses) and identifiable metadata (such as IP addresses and Google Account IDs) are cleared prior to logging.
  • Gemini Apps privacy hub (consumer defaults)
    Google / DeepMind provider artifacts · provider artifact · source tier B · support.google.com · retrieved 2026-08-05

Operational stabilitypartial

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Vertex publishes stable-version lifecycles with retirement dates, but the preview churn is severe: this model was shut down four months after launch with its ID silently repointed to a different model — precisely the failure mode this index monitors.

Receipts (2)

Adversarial resistancepartial

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistancepartial

Model card reports internal safety evaluations and red-teaming, but F5's monthly CASI leaderboard shows the Gemini 3.x line mid-to-low tier and strikingly volatile — a 55.7-point CASI swing between Feb and May 2026, with one month at 19.14 — against a top-10 threshold around 86.

Prompt injection (agentic)partial

Google publishes genuine layered defenses — model hardening against indirect injection, an agentic security architecture — yet researcher demonstrations against Gemini-connected surfaces keep landing, most recently a voice-assistant injection exploit.

Receipts (6)
  • Gemini 3 Pro model card (May 2026 revision)
    Google / DeepMind provider artifacts · provider artifact · source tier B · storage.googleapis.com · retrieved 2026-08-05
    A range of evaluations and red teaming activities were conducted to help improve the model and inform decision-making.
  • F5 Labs CASI/ARS leaderboards
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-05
    Gemini 3 performed strongly on raw capability but weakly on security, scoring only around 50 , up from previous versions’ low-40s but still far behind OpenAI and Anthropic.
  • F5 Labs: The Small Model Cliff (Gemini volatility analysis)
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-05
    Gemini 3 Pro preview moved 55.7 points across the period, ending close to the volatile cohort.
  • Advancing Gemini's security safeguards (DeepMind)
    Google / DeepMind provider artifacts · provider artifact · source tier B · deepmind.google · retrieved 2026-08-05
    This model hardening has significantly boosted Gemini’s ability to identify and ignore injected instructions, lowering its attack success rate.
  • Lessons from Defending Gemini Against Indirect Prompt Injections (white paper)
    Google / DeepMind provider artifacts · provider artifact · source tier B · storage.googleapis.com · retrieved 2026-08-05
    This is another reason we advocate for adversarial training as a necessary but not sufficient protection mechanism, and encourage the research community to focus on defense in depth approaches that provide mitigations at both the model and system level.
  • SafeBreach: Gemini voice-assistant prompt-injection exploit
    SafeBreach security research · reporting · source tier F · safebreach.com · retrieved 2026-08-05
    SafeBreach Labs researchers discovered a new security vulnerability that allows attackers to exploit Google Gemini through notification-based indirect prompt injections from messaging apps like WhatsApp, Slack, and SMS.

Transparencystrong

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Published model card with Frontier Safety Framework evaluation, and disclosed external pre-deployment testing for Gemini 3: UK AISI early access plus independent assessments by Apollo Research, Vaultis, and Dreadnode.

Receipts (2)

Compliance posturestrong

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

Google Cloud holds accredited ISO/IEC 42001:2023 certification with Vertex generative AI in scope, alongside SOC 1/2/3, ISO 27001/27017/27018, and HIPAA support.

Receipts (2)
  • Google Cloud ISO/IEC 42001 certification
    Google / DeepMind provider artifacts · provider artifact · source tier B · cloud.google.com · retrieved 2026-08-05
    Google Cloud Platform, Google Workspace, and Gemini (App) are certified as ISO/IEC 42001:2023 compliant.
  • Google Cloud compliance resource center
    Google / DeepMind provider artifacts · provider artifact · source tier B · cloud.google.com · retrieved 2026-08-05

Governance & evidence

Where your data goes

  • US
  • EU

Regional pinning available — customers can pin processing to a chosen region.

Vertex regional/jurisdictional endpoints (US, EU; expanding) with in-region ML processing; the global endpoint carries no residency guarantee.

Enterprise vs consumer terms

Enterprise vs consumer gap: wide. Paid API tier is abuse-logging only, while unpaid-tier content is used for improvement including human review and consumer Gemini Apps default to data use unless disabled.CONSUMERENTERPRISE / APIWORSE TERMS →
Wide gap

Paid API tier is abuse-logging only, while unpaid-tier content is used for improvement including human review and consumer Gemini Apps default to data use unless disabled.

Change cadence

within cadence4 tracked changes · typical interval ~111d 0d since last change (2026-08-05) Low confidence: only 4 events0d

4 tracked changes. The typical (median) interval between them is ~111 days (low confidence: a median of only 3 intervals). The last change was on 2026-08-05, 0 days before the as-of date (2026-08-05) — about 0.0× the typical interval. That is within its historical cadence.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

15 evidence refs3 distinct sources1 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×2 references
  • BGoogle / DeepMind provider artifactsprovider artifact×12 references
  • FSafeBreach security researchreporting×1 reference

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
Yes
ISO/IEC 42001
Yes
HIPAA eligible
Yes
Retention window
Paid tier: 55-day abuse-monitoring logs; ZDR by approval (sanitized logs); unpaid tier: content used for improvement incl. human review
Data residency
Vertex regional/jurisdictional endpoints (US, EU; expanding) with in-region ML processing; global endpoint carries no residency guarantee
EU AI Act
Signed the EU GPAI Code of Practice (2025-07-30).
Deprecation policy
Published Vertex model lifecycle with retirement dates; preview-tier IDs excluded in practice
Available via
Gemini API · Google Vertex AI

Incident history

2025-08-06
Promptware: calendar-invite injection drove Gemini-connected smart-home actions
Indirect prompt injection via calendar invites

Researchers showed that poisoned calendar invites could hijack Gemini assistant sessions into performing unintended actions, including smart-home control, when the user later asked routine questions. Google shipped mitigations. Demonstrates the indirect-injection surface of assistant-integrated deployments.

Outcome: Mitigations deployed by Google following disclosure.

Sources (1)

Change timeline

2026-08-05
Correction: two of our own claims did not survive verification
methodologywarning

Binding claims to verbatim source text caught two errors in our published data. (1) We stated NVIDIA's hosted API trial carries a no-training clause; the Trial Terms in fact reserve use of submitted content 'to improve NVIDIA products and services, including AI models' — the opposite. The data-governance summary and retention field are corrected and the clause is now quoted directly. (2) Our Gemini transparency entries credited the model card with naming external testers (UK AISI, Apollo, Vaultis, Dreadnode); the archived card names none of them, so that requirement now rests on Google's release materials and is flagged as needing a better primary citation. Published rather than silently fixed, per the methodology.

Evidence (1)
  • NVIDIA API Trial Terms of Service (§3.3)
    NVIDIA Nemotron provider artifacts · provider artifact · source tier B · assets.ngc.nvidia.com · retrieved 2026-08-05
    3.3 NVIDIA will collect the following data, without identifying specific users, to operate and improve the API Services and other products and services: (i) session metrics (e.g., the amount of processing power consumed, type of request made); (ii) error logs and execution logs relating to your session (e.g., whether your request was executed successfully); (iii) your feedback and ranking of specific API Services; and (iv) User Content and Generated Content to improve NVIDIA products and services, including AI models.
2026-08-05
Verification pass: top-3 entries re-verified against live sources; usage shares corrected
methodologynotice

Every citation for GPT-5.1, Claude Sonnet 4.5, and Gemini 3 Pro was fetched and checked. Grades held; corrections were citations, migrated documentation domains (claude.com, developers.openai.com), and facts: ISO/IEC 42001 confirmed for OpenAI (previously unverified), lifecycle markers added (all three models are now legacy or retired), and seeded usage-share estimates replaced with measured OpenRouter daily data — under which every tracked model now sits below 1% share.

Evidence (2)
2026-03-09
gemini-3-pro-preview shut down; model ID silently aliased to Gemini 3.1 Pro
versionwarning

Roughly four months after launch and without reaching a stable GA ID, Google shut down gemini-3-pro-preview and repointed the ID to gemini-3.1-pro-preview. Requests to the old ID now reach a different model — the canonical silent-version-swap failure mode this index tracks.

Evidence (1)
  • Gemini API changelog
    Google / DeepMind provider artifacts · provider artifact · source tier B · ai.google.dev · retrieved 2026-08-05
2025-11-18
Gemini 3 Pro launched with model card
versioninfo

Flagship release with published model card and Frontier Safety Framework coverage.

Evidence (1)
  • Gemini 3 Pro model card
    Google / DeepMind provider artifacts · provider artifact · source tier B · storage.googleapis.com · retrieved 2026-08-03

Compare this model: Gemini 3 Pro + open compare view →