ModelRiskIndex

Rankings / Moonshot AI

Kimi K3

Tier 010/100open weights

kimi-k3 via platform.kimi.ai (weights published 2026-07-26, custom Kimi K3 License)

Usage share 2.4% · OpenRouter rankings API (daily token share, 2026-08-04)

Graded on the hosted endpoint, unlike the K2 entry's self-hosted reference — Moonshot shifted to a hosted-first commercial launch (API two weeks before weights) and a custom license with revenue and attribution triggers replacing K2's modified-MIT. That shift moves the graded surface, and the tier, with it: K2 reached Tier 1 on the self-hosted convention; K3's hosted terms fail Tier 1 on data controls and safety disclosure.

Tier assessment

Fails one or more Tier 1 requirements: no published model card or safety evals, or terms that permit training on customer API data by default with no opt-out.

Tier 1 requirements
  • Published model card. A model card or equivalent technical documentation is published for this model.
  • Published safety evals. No safety evaluations are published for this model.
  • Documented safety policy. No documented safety or acceptable-use policy exists.
  • Enterprise data controls. Terms permit training on customer data by default with no documented opt-out.
Tier 2 requirements
  • External pre-deployment testing. No disclosed external pre-deployment testing.
  • Third-party certification. No verifiable third-party certification.
  • Versioning with changelogs. No versioning discipline or changelog for model changes.
  • Stated deprecation policy. No stated deprecation policy.

Missing for Tier 1: published safety evals, documented safety policy, enterprise data controls. The tier is computed from this checklist — satisfying these requirements moves the badge, automatically.

Risk analysisfive vectors · click a wedge for its evidence

Risk vectors — the receipts

Data governanceweak

What happens to your data: training-on-customer-data defaults, retention windows, residency options, and the enterprise-versus-consumer terms gap. A legal property, not a capability — it survives every model generation.

Hosted-platform privacy policy permits training on prompts, files, and outputs with no documented in-product opt-out; PRC processing; retention unspecified; enterprise restrictions by negotiation only. The published weights offer a deployer-controlled alternative.

Receipts (1)
  • Moonshot/Kimi privacy policy analysis (training use, no opt-out)
    Independent security analysis blogs · reporting · source tier F · cometapi.com · retrieved 2026-08-05
    Moonshot AI's privacy policy says user prompts and uploaded content may be used to improve and train its models, personal information may be shared with service providers and affiliates, and AI output may be inaccurate.
    Secondary analysis; first-party policy page pending direct capture by the snapshot pipeline.

Operational stabilityweak

Whether it changes without warning: versioning discipline, changelog quality, deprecation policy, and observed silent changes. The signal no one else tracks.

Dateless kimi-k3 API ID with no snapshot scheme, no changelog, and no published deprecation policy — weaker versioning than the K2 open-checkpoint era.

Receipts (1)

Adversarial resistanceweak

Whether an attacker can make it misbehave — direct jailbreaks against the model's own policies and indirect prompt injection in agentic tool use. Graded to the weaker of the two, because an attacker takes the easier path.

Jailbreak resistanceweak

Pliny-style jailbreaks circulated within days of launch; third-party analysis reports no external safety classifiers with remaining safeguards falling to persona and rephrasing attacks; no CASI listing yet and no developer safety evals — a regression in evidence terms from K2, which at least had eval history.

Prompt injection (agentic)weak

Heavily agentic positioning (Kimi Code, tool use, 1M context) with no published injection results for K3; academic evaluation of predecessor K2.5 found meaningful risks in tool-enabled agentic settings.

Receipts (3)
  • Penligent analysis: Kimi K3 jailbreak surface
    Independent security analysis blogs · reporting · source tier F · penligent.ai · retrieved 2026-08-05
    At publication time, Moonshot AI's official K3 materials do not describe a universal K3-specific jailbreak, a CVE assigned to the K3 model, or a dedicated public prompt-injection evaluation for the model.
    Third-party security analysis; success-rate specifics not independently confirmed.
  • Kimi K3 model card (no safety content)
    Moonshot AI provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-05
  • Academic safety evaluation of Kimi K2.5 (agentic risks; family-level evidence)
    Peer-reviewed / preprint research · independent eval · source tier C · arxiv.org · retrieved 2026-08-05
    Across diverse benchmarks covering CBRNE, cybersecurity, misalignment, bias, and harmfulness aspects, we identify patterns suggesting that Kimi K2.5 poses meaningful risks in tool-enabled agentic settings, and can be used to enable serious harm.

Transparencypartial

Whether you can see how it was built and tested: model cards, published safety evals, external pre-deployment testing, and disclosure of changes. The mechanism behind the tier ladder.

Weights published (1.5TB, one day ahead of pledge) with a 47-page technical report — but the report is architecture and benchmarks only: no safety or red-teaming section, training data and cutoff undisclosed, no external testing.

Receipts (2)
  • Kimi K3 weights
    Moonshot AI provider artifacts · provider artifact · source tier B · huggingface.co · retrieved 2026-08-05
    We release the full Kimi K3 model weights under the Kimi K3 License, making frontier intelligence openly available for research, deployment, and further innovation.
  • Kimi K3 technical report (repository)
    Moonshot AI provider artifacts · provider artifact · source tier B · github.com · retrieved 2026-08-05

Compliance postureweak

Whether it is certified and compliant: SOC 2, ISO/IEC 42001, HIPAA eligibility, EU AI Act readiness, and audit availability.

No SOC 2, ISO 27001, or ISO 42001 certifications found for the hosted platform.

Receipts (1)

Governance & evidence

Where your data goes

  • PRC

No regional pinning — the provider chooses where data is processed.

Hosted platform processes data in the PRC; no regional options.

Enterprise vs consumer terms

Enterprise vs consumer gap: none measured. Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against. No gap because there is no better tier: the hosted platform itself permits training on prompts, files, and outputs with no documented opt-out.CONSUMERENTERPRISE / APIWORSE TERMS →
No measured gap

No gap because there is no better tier: the hosted platform itself permits training on prompts, files, and outputs with no documented opt-out.

Level tiers can mean both are clean, or that the API tier is itself weak with nothing better to compare against.

Change cadence

no tracked changesNo tracked changes for Kimi K3. Absence of detection is not evidence of stability.none

No tracked changes for this model. That is absence of detection, not evidence of stability — it may mean the model is unwatched, not that it is unchanging.

Score volatility

No dated score readings recorded for this model yet. Readings are only entered where multiple real, dated third-party values exist — never interpolated.

Receipts — what backs this assessment

9 evidence refs4 distinct sources1 independent
  • CPeer-reviewed / preprint researchindependent eval×1 reference
  • BMoonshot AI provider artifactsprovider artifact×5 references
  • EOpenRouter model rankingsusage data×1 reference
  • FIndependent security analysis blogsreporting×2 references

retrieved 2026-08-05

Compliance & deployment

Trains on customer data by default
Yes — flag
SOC 2
not verified
ISO/IEC 42001
not verified
HIPAA eligible
not verified
Retention window
Not documented
Data residency
PRC (hosted platform)
EU AI Act
No published EU AI Act posture.
Deprecation policy
none stated
Available via
platform.kimi.ai · Self-hosted (Kimi K3 License weights) · Multiple inference providers

Change timeline

No tracked changes yet for this model.

Compare this model: Kimi K3 + open compare view →