ModelRiskIndex

Rankings / NVIDIA

NVIDIA Nemotron 3 Ultra

Tier 133/100open weights

NVIDIA-Nemotron-3-Ultra-550B-A55B (open weights, OpenMDW-1.1; self-hosted NIM reference)

Usage share 4.1% · OpenRouter rankings API (daily token share, 2026-08-04)

The static-versus-adaptive gap (99.4% → 8.3% injection resistance) is this entry's headline and a general caution about benchmark-based grades — exactly why the methodology weighs adaptive results over static suites. Graded on the self-hosted reference deployment.

Tier assessment
Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. Safety evaluations for this model are published.Model Card++ safety subcards and technical report; quantitative depth below closed-lab system cards.
  • Documented safety policy. A documented safety, usage, or acceptable-use policy governs the model.Trustworthy-AI subcards plus an explicit guardrail-circumvention warning in the license terms.
  • Enterprise data controls. Not applicable to this artifact.
Tier 2 requirements
  • External pre-deployment testing. No disclosed external pre-deployment testing.NR Labs' pre-release disclosure on the Nano sibling informed hardening, and F5/Ubicloud tested post-release — but no disclosed pre-deployment testing program exists.
  • Third-party certification. No verifiable third-party certification.No certification pathway currently exists for open weights.
  • Versioning with changelogs. Model versions are explicitly identified and changes are changelogged.Immutable dated checkpoints; no changelog.
  • Stated deprecation policy. No stated deprecation policy.

Missing for Tier 2: external pre-deployment testing, third-party certification, stated deprecation policy. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governancen/a — deployer

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Deployer property in the self-hosted NIM reference deployment, which is the intended production path. Note the hosted build.nvidia.com trial is the opposite of a safe default: its Trial Terms reserve the right to use submitted content to improve NVIDIA products including AI models, so the trial endpoint should not be treated as a no-train surface.

Receipts (1)
  • NVIDIA API Trial Terms of Service (§3.3 data collection)
    NVIDIA Nemotron provider artifacts · provider artifact · source tier B · assets.ngc.nvidia.com · retrieved 2026-08-05
    3.3 NVIDIA will collect the following data, without identifying specific users, to operate and improve the API Services and other products and services: (i) session metrics (e.g., the amount of processing power consumed, type of request made); (ii) error logs and execution logs relating to your session (e.g., whether your request was executed successfully); (iii) your feedback and ranking of specific API Services; and (iv) User Content and Generated Content to improve NVIDIA products and services, including AI models.

Operational stabilitypartial

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Immutable dated HF checkpoints and a v1.0-GA versioned card; but NVIDIA's own naming is undated, there is no changelog, and no deprecation policy is stated.

Receipts (1)

Adversarial resistanceweak

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistancepartial

A genuinely split verdict: debuted at CASI 83.68 — #4 overall and the highest security score ever recorded for open weights — with 99.5% jailbreak robustness on static benchmarks; but Ubicloud's adaptive red-teaming collapsed it to a 50.2/100 adaptive index with 33% harmful-refusal resistance, including refusals whose reasoning traces still worked through the disallowed content.

Prompt injection (agentic)weak

Training includes synthetic indirect-injection data, yet Ubicloud measured 8.3% injection resistance under adaptive attack (99.4% on static benchmarks — the gap is the lesson). NVIDIA's real defenses are external add-ons: the Nemotron 3.5 Content Safety guard model and NeMo Guardrails, both trivially omitted by deployers.

Receipts (4)

Transparencypartial

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

The best open-weight disclosure tracked: Model Card++ with safety/bias/privacy subcards, a technical report, and training-data disclosure (~20T tokens) — held back from strong because quantitative safety results are thinner than closed-lab system cards and external testing was not a disclosed pre-deployment program.

Receipts (2)
  • Nemotron 3 Ultra model card (Model Card++ with safety subcards)
    NVIDIA Nemotron provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-05
    For more detailed information on ethical considerations for this model, please see the Model Card++ Bias , and Privacy Subcards.
  • Nemotron 3 Ultra technical report
    NVIDIA Nemotron provider artifacts · provider artifact · source tier B · research.nvidia.com · retrieved 2026-08-05
    We pretrained our base model in NVFP4 with 20 trillion text tokens using a Warmup-Stable-Decay learning rate schedule.

Compliance posturen/a — deployer

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

Deployer property: no certification attaches to open weights, and none is published for the hosted trial endpoint.

Receipts (1)

Governance & evidence

Where your data goes

  • SELF

No regional pinning — the provider chooses where data is processed.

Graded on the self-hosted NIM reference deployment; the build.nvidia.com trial endpoint carries separate trial terms.

Enterprise vs consumer terms

Enterprise vs consumer gap: none measured. Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against. No consumer tier exists; production use is self-hosted by design.CONSUMERENTERPRISE / APIWORSE TERMS →
No measured gap

No consumer tier exists; production use is self-hosted by design.

Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against.

Change cadence

insufficient history1 tracked change · last 2026-08-05 · 0d since Insufficient history to estimate a cadence.0d

One tracked change, on 2026-08-05 — 0 days before the as-of date (2026-08-05). A single event cannot establish a cadence, so days-since is shown without a baseline.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

10 evidence refs4 distinct sources2 independent
  • CUbicloud AI safety studiesindependent eval×2 references
  • CF5 Labs CASI/ARS leaderboardindependent eval×1 reference
  • BNVIDIA Nemotron provider artifactsprovider artifact×6 references
  • EOpenRouter model rankingsusage data×1 reference

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
No
SOC 2
not verified
ISO/IEC 42001
not verified
HIPAA eligible
not verified
Retention window
Deployer-controlled (self-hosted NIM). Hosted trial: Trial Terms §3.3 reserve use of submitted content to improve NVIDIA products including AI models
Data residency
Deployer-controlled
EU AI Act
Open-weight GPAI provisions apply; deployer obligations dominate.
Deprecation policy
None stated; weights remain available once released
Available via
Self-hosted (NIM) · build.nvidia.com (trial) · OpenRouter and other hosts · Perplexity Pro

Change timeline

2026-08-05
Correction: two of our own claims did not survive verification
methodologywarning

Binding claims to verbatim source text caught two errors in our published data. (1) We stated NVIDIA's hosted API trial carries a no-training clause; the Trial Terms in fact reserve use of submitted content 'to improve NVIDIA products and services, including AI models' — the opposite. The data-governance summary and retention field are corrected and the clause is now quoted directly. (2) Our Gemini transparency entries credited the model card with naming external testers (UK AISI, Apollo, Vaultis, Dreadnode); the archived card names none of them, so that requirement now rests on Google's release materials and is flagged as needing a better primary citation. Published rather than silently fixed, per the methodology.

Evidence (1)
  • NVIDIA API Trial Terms of Service (§3.3)
    NVIDIA Nemotron provider artifacts · provider artifact · source tier B · assets.ngc.nvidia.com · retrieved 2026-08-05
    3.3 NVIDIA will collect the following data, without identifying specific users, to operate and improve the API Services and other products and services: (i) session metrics (e.g., the amount of processing power consumed, type of request made); (ii) error logs and execution logs relating to your session (e.g., whether your request was executed successfully); (iii) your feedback and ranking of specific API Services; and (iv) User Content and Generated Content to improve NVIDIA products and services, including AI models.

Compare this model: NVIDIA Nemotron 3 Ultra + open compare view →