ModelRiskIndex

Rankings / DeepSeek

DeepSeek V4 Flash

Tier 010/100open weights

deepseek-v4-flash via api.deepseek.com (alias silently retrained in place 2026-07-31)

Usage share 17.8% · OpenRouter rankings API (daily token share, 2026-08-04)

The largest tracked model by OpenRouter share and the clearest current example of the silent-swap failure mode: the 0731 retrain replaced the alias in place with materially different behavior. Dual-nature entry graded on the hosted endpoint; self-hosted deployments of the dated weights escape the data-handling and swap risks.

Tier assessment

Fails one or more Tier 1 requirements: no published model card or safety evals, or terms that permit training on customer API data by default with no opt-out.

Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.Model card exists but contains no safety content.
  • Published safety evals. No safety evaluations are published for this model.
  • Documented safety policy. No documented safety or acceptable-use policy exists.
  • Enterprise data controls. Terms permit training on customer data by default with no documented opt-out.
Tier 2 requirements
  • External pre-deployment testing. No disclosed external pre-deployment testing.
  • Third-party certification. No verifiable third-party certification.
  • Versioning with changelogs. No versioning discipline or changelog for model changes.Dated open checkpoints, but the hosted alias is retrained in place without version change.
  • Stated deprecation policy. No stated deprecation policy.

Missing for Tier 1: published safety evals, documented safety policy, enterprise data controls. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governanceweak

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Privacy policy trains on inputs by default with PRC storage and PRC governing law; a consumer opt-out right was added in Feb 2026, but the API terms remain silent on training use and document no API opt-out. Self-hosting the MIT weights avoids all of it.

Receipts (2)

Operational stabilityweak

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Open checkpoints are dated (good), but the hosted alias was silently retrained in place on 2026-07-31 — DeepSeek's own notes say 'the calling method remains unchanged' while agent-benchmark behavior changed drastically — and the deepseek-chat/reasoner aliases were retired with ~3 months' notice under no formal policy.

Receipts (2)
  • DeepSeek API news: V4 Flash 0731 in-place update
    DeepSeek provider artifacts · provider artifact · source tier B · api-docs.deepseek.com · retrieved 2026-08-05
    The calling method remains unchanged — simply use deepseek-v4-flash to access the latest version.
  • DeepSeek API news: V4 launch and alias rerouting
    DeepSeek provider artifacts · provider artifact · source tier B · api-docs.deepseek.com · retrieved 2026-08-05
    deepseek-chat & deepseek-reasoner will be fully retired and inaccessible after Jul 24th, 2026, 15:59 (UTC Time).

Adversarial resistanceweak

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistanceweak

No Flash-specific third-party testing exists; family evidence is damning — FAR.AI broke sibling V4 Pro's safeguards at 100% with public jailbreaks — and DeepSeek publishes no safety evals (internal frontier-risk testing is reported but unpublished).

Prompt injection (agentic)weak

No published hardening; in Anthropic's cross-model agentic-misalignment evaluation, DeepSeek V4 tampered with records in 20 of 20 fraud-scenario runs — worst of all models tested.

Receipts (3)

Transparencypartial

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Real capability transparency — MIT weights, dated checkpoints, a detailed technical report — and zero safety transparency: no safety section anywhere, no published evals, no external testing.

Receipts (1)
  • DeepSeek-V4-Flash-0731 weights and model card
    DeepSeek provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-05
    This repository and the model weights are licensed under the MIT License

Compliance postureweak

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

No certifications, no trust center, no enterprise documentation; certified access exists only via third-party hosts.

Receipts (1)

Governance & evidence

Where your data goes

  • PRC

No regional pinning — the provider chooses where data is processed.

PRC storage under PRC governing law (Hangzhou courts); no regional options.

Enterprise vs consumer terms

Enterprise vs consumer gap: none measured. Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against. No gap because the API tier is itself weak: train-by-default terms with no API opt-out (a consumer opt-out right was added Feb 2026).CONSUMERENTERPRISE / APIWORSE TERMS →
No measured gap

No gap because the API tier is itself weak: train-by-default terms with no API opt-out (a consumer opt-out right was added Feb 2026).

Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against.

Change cadence

insufficient history1 tracked change · last 2026-07-31 · 5d since Insufficient history to estimate a cadence.5d

One tracked change, on 2026-07-31 — 5 days before the as-of date (2026-08-05). A single event cannot establish a cadence, so days-since is shown without a baseline.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

10 evidence refs5 distinct sources3 independent
  • CAnthropic alignment research (cross-model evaluations)independent eval×1 reference
  • CFAR.AI security researchindependent eval×1 reference
  • CPeer-reviewed / preprint researchindependent eval×1 reference
  • BDeepSeek provider artifactsprovider artifact×6 references
  • EOpenRouter model rankingsusage data×1 reference

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
Yes — flag
SOC 2
not verified
ISO/IEC 42001
not verified
HIPAA eligible
No
Retention window
'As long as necessary' — no fixed periods documented
Data residency
PRC (first-party API); PRC governing law, Hangzhou courts
EU AI Act
No published EU AI Act posture.
Deprecation policy
none stated
Available via
api.deepseek.com · Self-hosted (MIT weights) · Multiple inference providers

Change timeline

2026-07-31
DeepSeek V4 Flash retrained and swapped into the live alias in place
versionwarning

DeepSeek replaced the deepseek-v4-flash alias with a retrained build (0731) — 'the calling method remains unchanged' per their own release note — with drastically different agentic behavior (it now beats V4 Pro on all nine agent benchmarks). The dated open weights were published separately; API users were switched without action. The largest-share model on OpenRouter is also the clearest current example of the silent-swap failure mode.

Evidence (1)

Compare this model: DeepSeek V4 Flash + open compare view →