ModelRiskIndex

Rankings / Alibaba (Qwen)

Qwen3 235B-A22B

Tier 150/100draft — pending re-verificationopen weights

Qwen3-235B-A22B-Instruct (open weights, self-hosted reference)

Usage share 0.13% · OpenRouter rankings API (daily token share, 2026-08-04)

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.Basic safety benchmarks in the technical report; not a dedicated safety eval suite.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.
  • Enterprise data controls. Not applicable to this artifact.
Tier 2 requirements
  • External pre-deployment testing. No disclosed external pre-deployment testing.
  • Third-party certification. No verifiable third-party certification.No certification pathway currently exists for open weights — an open methodology question, not a gap unique to this provider.
  • Versioning with changelogs. Model versions are explicitly identified and changes are changelogged.
  • Stated deprecation policy. A deprecation policy with notice windows is published.Weights remain available once released.

Missing for Tier 2: external pre-deployment testing, third-party certification. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancen/a — deployer

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Deployer property: in the self-hosted reference deployment, data never leaves deployer infrastructure. Alibaba Cloud's hosted Model Studio has separate terms not graded here.

Receipts (1)

Operational stabilitystrong

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Immutable published checkpoints with explicit versioned releases; no silent-change surface in the self-hosted reference deployment.

Receipts (1)

Adversarial resistanceweak

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistanceweak

Independent testing shows high attack success on the bare weights; safety behavior is light and easily bypassed, consistent with most open-weight releases.

Prompt injection (agentic)weak

No injection hardening or published agent-scenario safety results.

Receipts (2)
  • F5 Labs CASI leaderboard
    F5 Labs CASI/ARS leaderboard · independent eval · source tier C · f5.com · retrieved 2026-08-03
  • Qwen3 technical report
    Alibaba Qwen provider artifacts · provider artifact · source tier B · qwenlm.github.io · retrieved 2026-08-03

Transparencypartial

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Detailed technical report and model card, Apache 2.0 license; no safety evals of substance and no external testing.

Receipts (1)
  • Qwen3 technical report
    Alibaba Qwen provider artifacts · provider artifact · source tier B · qwenlm.github.io · retrieved 2026-08-03
    While Qwen2.5 was pre-trained on 18 trillion tokens, Qwen3 uses nearly twice that amount, with approximately 36 trillion tokens covering 119 languages and dialects.

Compliance posturen/a — deployer

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

Deployer property: no certification attaches to the weights; deployer inherits all compliance obligations.

Receipts (1)

Governance & evidence

Where your data goes

  • SELF

No regional pinning — the provider chooses where data is processed.

Self-hosted reference deployment; Alibaba Cloud's hosted Model Studio has separate terms not graded here.

Enterprise vs consumer terms

Enterprise vs consumer gap: none measured. Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against. No consumer tier exists for the self-hosted reference deployment; data handling is deployer-controlled.CONSUMERENTERPRISE / APIWORSE TERMS →
No measured gap

No consumer tier exists for the self-hosted reference deployment; data handling is deployer-controlled.

Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against.

Change cadence

insufficient history1 tracked change · last 2026-08-04 · 1d since Insufficient history to estimate a cadence.1d

One tracked change, on 2026-08-04 — 1 days before the as-of date (2026-08-05). A single event cannot establish a cadence, so days-since is shown without a baseline.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

7 evidence refs3 distinct sources1 independent
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • BAlibaba Qwen provider artifactsprovider artifact×5 references
  • EOpenRouter model rankingsusage data×1 reference

retrieved 2026-08-03 — 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
not verified
ISO/IEC 42001
not verified
HIPAA eligible
not verified
Retention window
Deployer-controlled
Data residency
Deployer-controlled
EU AI Act
Open-weight GPAI provisions apply; deployer obligations dominate.
Deprecation policy
Weights remain available once released
Available via
Self-hosted · Alibaba Cloud Model Studio · Multiple inference providers

Change timeline

2026-08-04
Methodology v0.2: computed tiers and first-class N/A grades
methodologynotice

Tiers are now derived from a per-requirement checklist rather than assigned, and not-applicable / under-review became first-class grade states excluded from the composite score. Deployer-property vectors on open-weight models (data handling, compliance) moved from graded to not-applicable, changing the composite scores of the tagged models.

Evidence (1)

Compare this model: Qwen3 235B-A22B + open compare view →