Companies

Inception Labs

inceptionlabs.ai

Inception builds diffusion-based large language models delivering faster, cheaper, controllable AI for developers and enterprises.

HQPalo Alto, California, United States
Employees11-50
Funding$50M
Valuation$9.74M
19 active roles
Profile 6mo agoJobs checked 19h ago
AI / MLFoundation Model ProviderB2B SaaSSeries A$50M-$200M

About

Inception Labs is an AI research and product company building Mercury, a commercially available family of diffusion-based LLMs for production applications. It targets developers and enterprises, including Fortune 500 companies, and differentiates through parallel token generation, high speed, lower costs, controllable outputs, multimodal support, and OpenAI-compatible deployment.

Market

Inception Labs competes in the generative-AI and LLM infrastructure market, positioning diffusion-based LLMs as a faster and more efficient alternative to conventional autoregressive models. It differentiates through parallel token processing, multimodal inputs, fine-grained schema/output control, and lower-cost real-time deployment for enterprises and developers.

Target Customers

Inception Labs targets enterprises and software developers building production-grade, real-time AI applications, particularly teams that need fast, multimodal generation. Likely buyers include enterprise AI, platform, and product-engineering teams, as well as developers integrating models through an API.

At a Glance

Problem

Inception Labs targets the central production bottleneck in generative AI: traditional autoregressive language models generate one token at a time, making high-quality responses slow and expensive—especially when reasoning requires longer chains, retries, or multiple agent steps. The company says its approach delivers sub-300-millisecond time to first token, five-to-seven-times higher throughput, and roughly 70% lower cost per task. This matters because latency-sensitive applications have often been forced to use smaller, less capable models simply to stay within speed and cost budgets.

The clearest killer use case is in-flow software development: rapid code generation, autocomplete, editing, and agentic coding workflows where waiting for a response interrupts the developer. Mercury Coder was positioned as running at more than 1,000 tokens per second while matching or exceeding the quality of speed-optimized alternatives. The same low-latency economics also support real-time voice, customer support, search, and business-process agents.

Product / Service

Inception builds diffusion language models, or dLLMs, rather than conventional autoregressive LLMs. Instead of producing tokens sequentially, Mercury generates multiple tokens in parallel through a coarse-to-fine process, improving speed and GPU efficiency while retaining control over schemas and structured outputs. Its current family includes Mercury 2, a reasoning model with a 128K context window, tool use, and structured output, and Mercury Edit 2, a smaller coding-focused model for autocomplete and next-edit workflows.

The company sells access through a hosted, usage-priced API, with free, developer, and enterprise plans, and also offers enterprise and on-premise deployment. The API is OpenAI-compatible, so existing clients and application stacks can be reused with limited integration work; models are also distributed through platforms including AWS Bedrock, Azure AI Foundry, OpenRouter, and Models.dev. Mercury 2 and Mercury Edit 2 are listed at $0.25 per million input tokens and $0.75 per million output tokens, giving developers a relatively inexpensive way to add fast reasoning, coding, voice, search, and workflow automation.

Market

Inception competes in the foundation-model and inference market, with a specific focus on low-latency, cost-efficient production LLMs and coding models. Its differentiated category claim is diffusion-based language generation, while its practical alternatives include speed-optimized autoregressive models from OpenAI and Anthropic, such as GPT-4o Mini, GPT-5 mini, Claude 3.5 Haiku, and Claude Haiku 4.5. The company’s pitch is not simply higher model quality; it is frontier-quality output at materially lower latency and inference cost, enabling applications that need responses in real time.

The evidence indicates an early commercial company rather than a purely pre-revenue research project. Inception launched its first commercial-scale Mercury model in February 2025, subsequently raised a $50 million financing round led by Menlo Ventures with participation from investors including M12, Mayfield, Snowflake Ventures, Databricks, Innovation Endeavors, Andrew Ng, and Andrej Karpathy, and says its models are being deployed at Fortune 500 companies. Its models are available through major cloud channels and a public API, but the available evidence does not disclose revenue, customer counts, or contract size, so commercial traction is visible while financial traction remains undisclosed.

Founders & Leadership

Stefano ErmonFounder
Chief Executive Officer
Aditya GroverFounder
Co-founder & Chief Technology Officer

Funding History

2024-07
SeedUndisclosed

Not publicly disclosed

2025-11
Seed$50M

Menlo Ventures

Recent News

2026-08-01product
Introducing Mercury 2

Inception Labs introduced Mercury 2, describing it as a reasoning language model built to make production AI feel instant. The announcement reports throughput of 1,009 tokens per second.

2026-07-29product
More builders. More throughput. Better Mercury 2.

Inception Labs announced that every new Inception API key includes 100 million free tokens, giving developers capacity to benchmark Mercury 2 against existing production stacks.

2026-06-20
Inception Labs' Mercury 2 AI outperforms Google's DiffusionGemma

Crypto Briefing covered Mercury 2's performance and reported that the model, launched in February 2026, was designed to compete with speed-optimized models such as Claude 4.5 Haiku and GPT-5 Mini.

2026-02-24
Mercury 2: The LLM That Doesn't Generate Like an LLM

Developers Digest profiled Mercury 2 as a diffusion-based model capable of reasoning-style iterative refinement and generating at more than 1,000 tokens per second.

2025-11-19partnership
Mercury Diffusion LLM Now Available on Azure AI Foundry

Inception Labs announced that Mercury became available on Azure AI Foundry, bringing its commercial-scale diffusion language model to enterprise developers.

2025-11-07funding
Inception Labs banks $50M to make diffusion LLMs 10x faster

Tech Funding News reported that Palo Alto-based Inception Labs AI raised $50 million in fresh funding to advance diffusion-based large language models.

2025-11-06funding
The Next Step for dLLMs: Scaling up Mercury

Inception Labs announced an upgraded Mercury model and disclosed a $50 million financing round led by Menlo Ventures. The model was made available through Inception's API Platform and partners including OpenRouter and Models.dev.

2025-09-25partnership
Introducing the Inception API

Inception Labs launched its API platform and announced that Mercury Coder was available through OpenRouter and integrated with popular IDEs, including through a partnership with Continue for VS Code.

2025-09-25product
Ultra-Fast Apply-Edit with Mercury Coder

Inception Labs reported that Mercury Coder applied edits accurately 92% of the time while running 46 times faster than comparable frontier-model workflows.

Active Roles

19
San Mateo/Engineering/21d ago
San Mateo/Forward-Deployed Engineer/104d ago
San Mateo/Marketing/148d ago
San Mateo/Engineering/176d ago
San Mateo/Engineering/178d ago
San Mateo/Engineering/178d ago
San Mateo/Engineering/178d ago
San Mateo/Data & Analytics/178d ago
San Mateo/Engineering/178d ago
San Mateo/Engineering/178d ago
San Mateo/Product/196d ago

Business Model

The company monetizes Mercury through its API platform on a usage-based, per-million-token basis, with a 10-million-token free allowance for new accounts. It also offers enterprise revenue paths through AWS Bedrock and Azure Foundry availability, fine-tuning, private deployments, and custom service-level agreements.

Products

Diffusion-based language modelsMercury ChatAPI PlatformModel and developer-platform offerings

Customers

SearchBloxRadientBuildglare

Tech Stack

Diffusion-based language modelsLarge language models (LLMs)Parallel token processingMultimodal AI for text, code, audio, and imagesSchema-adherent, fine-grained output control

Competitors

Goodfire
Bria
Cognition

Key Investors

Menlo Ventures, Mayfield, Innovation Endeavors, M12, Snowflake Ventures, Databricks Ventures, NVentures, Andrew Ng, Andrej Karpathy