

Mercury 2.5
The diffusion LLM that generates at 1,100+ tokens/sec
What is Mercury 2.5?
Mercury 2.5 is Inception's most capable diffusion language model yet, generating at over 1,100 tokens per second by producing text in parallel instead of one token at a time. It brings a 40% intelligence improvement over Mercury 2, 260K context, tunable reasoning, parallel tool calls, and structured JSON — built for latency-sensitive search, voice, coding, and agent workloads.
Screenshots
?
No comments yet. Be the first!





