Breaking: Moonshot AI Releases Kimi K3—The World’s First 2.8 Trillion Parameter Open-Source Model

2026-07-17
BREAKING // KIMI K3 RELEASED // 2.8 TRILLION PARAMETERS // 896 EXPERTS, 16 ACTIVE // OPEN-SOURCE WEIGHTS JULY 27 // BREAKING // KIMI K3 RELEASED // 2.8 TRILLION PARAMETERS // 896 EXPERTS, 16 ACTIVE // OPEN-SOURCE WEIGHTS JULY 27 // 
WIRE REPORT · MOONSHOT AI

Breaking: Moonshot AI Releases Kimi K3 — The World's First 2.8 Trillion Parameter Open-Source Model

The open-source AI landscape just experienced a seismic shift. Moonshot AI has officially announced Kimi K3, fracturing previous open-source limitations by introducing a staggering 2.8 trillion parameter Mixture of Experts (MoE) model.

Designed from the ground up for long-horizon software engineering, end-to-end knowledge work, and deep multi-step reasoning, Kimi K3 represents a critical milestone. It marks the first time an open-source architecture has comfortably crossed the 2-trillion parameter threshold, challenging proprietary titans on their own turf.

Here is a comprehensive look at the architecture, benchmarks, real-world utility, and developer impact of this newly minted AI powerhouse.

2.8T
Parameters
896
Total Experts
16
Active / Token
1M
Context Window
Exhibit 01 Kimi K3 overview
01

Architectural Brilliance: Stable LatentMoE & Ultra-Sparsity

At the heart of Kimi K3's massive 2.8T scale lies a masterclass in computing efficiency. Utilizing the revolutionary Stable LatentMoE framework, Moonshot AI has radically expanded the model's sparsity.

The model features an unprecedented 896 total experts, but dynamically routes and activates just 16 experts per token.

[Input Token] [Stable LatentMoE Router] (Activates 16 out of 896 Experts) [Optimized Output]

This structural breakthrough yields major advantages:

  • Massive Scaling, Minimal Compute: while holding 2.8 trillion parameters of foundational knowledge, the active inference cost mimics a vastly smaller model.
  • Native 1M Token Context Window: Kimi K3 handles a 1 million token context window natively, supported by advanced automatic context caching to slash API latency and overhead costs.
  • Always-On Reasoning Mode: unlike its predecessors, Kimi K3 is natively configured with an immutable, deep-thinking reasoning mode.
Dev Note

⚠️ Developer Note: Kimi K3 deprecates the older thinking parameter from K2.x. Developers must now use the reasoning_effort configuration (currently defaulting to max) to orchestrate its cognitive depth.

02

Global Benchmarks: Where Does Kimi K3 Stand?

Moonshot AI's internal evaluations and early benchmark tracking position Kimi K3 at the absolute forefront of global AI intelligence. In comprehensive evaluation suites, Kimi K3's overall cognitive capability ranks just behind proprietary flagships like Claude Fable 5 and GPT-5.6 Sol, outperforming all existing open-source counterparts.

Exhibit 02 Kimi K3 benchmark chart

The Core Strengths Matrix

Kimi K3 excels across three core pillars optimized heavily for enterprise agent workflows:

Capability Pillar Technical Implementation Practical Enterprise Focus
Long-Horizon Coding Multi-step task execution, native terminal tool interactions, large codebase ingestion. Autonomously fixing bugs, refactoring legacy repositories, software automation agents.
Visual Reasoning & SE Blends software engineering workflows with visual UI/UX understanding. Frontend optimization, game development debugging, and CAD layout analysis using screenshots.
Knowledge Work Deep research, automated document synthesis, multi-hop logical deductions. High-stakes legal review, financial compliance audits, and multi-document synthesis.
Exhibit 03 Kimi K3 architecture diagram
03

Deep-Dive: Enterprise Use Cases & API Economics

Kimi K3 isn't just an engineering flex; it is highly optimized for actual production economics. Moonshot AI has introduced a highly competitive, non-tiered API structure that leverages automatic context caching.

Kimi K3 API Pricing Structure

Rather than punishing developers for large context windows with tiered penalties, Kimi K3 bills usage at flat volumetric rates:

Input Tokens (Standard)¥20.00 / MTok
Context Cache Hits (90% discount for repeated context structures!)¥2.00 / MTok
Output Tokens¥100.00 / MTok
Exhibit 04 Kimi K3 API pricing chart

Advanced Agent Toolkit

To support next-generation agentic workflows, Kimi K3 also rolls out powerful API enhancements:

  • Dynamic Tool Loading & Constraints: advanced tool_choice parameters allow developers to strictly enforce or dynamically adjust which local tools the model can invoke mid-stream.
  • High-Fidelity Structured Outputs: native JSON Mode and strict response_format via JSON Schema ensure that the model formats its 2.8T reasoning into completely predictable data structures for backend ingestion.
04

Open Source Timeline & Ecosystem Impact

The open-source community won't have to wait long to get their hands on the weights. Moonshot AI is currently collaborating directly with inference partners, chip manufacturers, and open-source maintainers to align optimization details.

Full Weights Public Release
July 27, 2026
Alongside a full Technical Report

This open release is poised to completely democratize ultra-large-scale MoE modeling. Historically, running or fine-tuning anything near a 3-trillion parameter scale was restricted to tech oligarchs. By optimizing the Stable LatentMoE structure, Moonshot AI provides the global developer ecosystem with a template on how to build, run, and scale massive-parameter models without requiring a localized supercomputer array.

The Verdict: A Bold New Era for Open-Source AI

Kimi K3 proves that open-source AI is no longer playing catch-up; it is actively setting the tempo. By marrying an incredible 2.8T parameter foundation with hyper-sparse 16-expert activation, Moonshot AI has built a model that is intensely smart yet commercially practical.

For tech companies looking to deploy complex, autonomous software agents or deep document-crunching pipelines without proprietary vendor lock-in, the countdown to July 27 has officially begun.

What are your thoughts on Kimi K3's architecture? Will the 90% context caching discount change how you build AI agents? Let us know in the comments below, and don't forget to share this breakdown with your engineering team!

Wire Report · Moonshot AI · Kimi K3

300+ AI Models for
OpenClaw & AI Agents

Save 20% on Costs