mercury-2.5-preview · Inception
Mercury 2.5 is the latest diffusion-based large language model (dLLM) released by Inception. It is the fastest inference LLM; unlike the sequential token-by-token generation approach, Mercury 2.5 can generate and optimize multiple tokens in parallel, achieving a generation speed of 1,107 tokens per second on standard GPUs. Compared to Mercury 2, its intelligence has increased by more than 10 percentage points, and its quality rivals leading cost-optimized frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
80% off
Mercury 2.5 is the latest diffusion-based large language model (dLLM) released by Inception. It is the fastest inference LLM; unlike the sequential token-by-token generation approach, Mercury 2.5 can generate and optimize multiple tokens in parallel, achieving a generation speed of 1,107 tokens per second on standard GPUs. Compared to Mercury 2, its intelligence has increased by more than 10 percentage points, and its quality rivals leading cost-optimized frontier models such as GPT-5.6 Luna (Low), Gemini 3.5 Flash-Lite, and Claude Haiku 4.5.
Mercury 2.5 Preview has a 260,000 token context window.
On AIHubMix, Mercury 2.5 Preview costs $0.04 per million input tokens and $0.15 per million output tokens. Cached input reads are billed at $0.004 per million tokens. These are promotional rates — 80% off.
Mercury 2.5 Preview accepts text input.
Mercury 2.5 Preview supports thinking and tool calling. Per-protocol parameter support is listed in the capability table on this page.
Mercury 2.5 Preview is available through the AIHubMix unified API. The API is OpenAI-compatible: point your OpenAI SDK at https://aihubmix.com/v1, use your AIHubMix API key, and set the model name to mercury-2.5-preview — no other code changes needed.
Mercury 2.5 Preview is developed by Inception. AIHubMix aggregates it alongside models from other providers behind one API and one bill.
Mercury 2.5 Preview was released on September 2, 2026 by Inception.
Use Mercury 2.5 Preview via the AIHubMix unified API — one interface for every major LLM.