← all stories

Mercury 2 5

Facts

Announced
Mercury 2.5 · 2026-09-09
Noted
a diffusion-based language model · 2026-09-09
Noted
refines multiple tokens in parallel · 2026-09-09
Noted
outputs 1107 tokens per second · 2026-09-09
Noted
up from Mercury 2's 1009 · 2026-09-09
Noted
context length doubled to 260,000 tokens · 2026-09-09
Noted
rates its quality on par with GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5 · 2026-09-09
Noted
Output pricing stays at $0.75 per million tokens · 2026-09-09
Noted
input pricing drops to $0.20 · 2026-09-09
Noted
80 percent launch discount on OpenRouter · 2026-09-09

Structured graph also available as JSON at /public/entities/mercury-2-5. CC BY 4.0.

All coverage

2h ago

Inception Ships Mercury 2.5 Diffusion LLM at 1107 Tokens Per Second

Inception announced Mercury 2.5, a diffusion-based language model that refines multiple tokens in parallel instead of generating text left to right. It outputs 1107 tokens per second, up from Mercury 2's 1009, with context length doubled to 260,000 tokens. Inception rates its quality on par with GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5. Output pricing stays at $0.75 per million tokens while input pricing drops to $0.20, with an 80 percent launch discount on OpenRouter.