
Inception, the company developing diffusion large language models (dLLMs), has raised $50 million in new funding. Menlo Ventures led the round, with participation from Mayfield, Innovation Endeavors, NVentures (NVIDIA's venture capital arm), M12 (Microsoft's venture capital fund), Snowflake Ventures, and Databricks Investment.
The funding arrives as enterprises struggle with the limitations of current AI systems. Traditional large language models rely on autoregression, generating words one at a time in sequence. This creates a structural bottleneck that slows response times, drives up costs, and restricts how companies can deploy AI at scale.
Inception takes a fundamentally different approach. The company's dLLMs use diffusion technology—the same method behind image and video tools like DALL·E, Midjourney, and Sora—to generate text in parallel rather than sequentially. The result is a 10x improvement in both speed and efficiency without sacrificing quality.
Mercury, Inception's first commercially available model, operates 5-10x faster than speed-optimized offerings from OpenAI, Anthropic, and Google while matching their accuracy. This performance makes the technology particularly valuable for applications where latency matters: interactive voice agents, live code generation, and dynamic user interfaces.
"The team at Inception has demonstrated that dLLMs aren't just a research breakthrough; it's a foundation for building scalable, high-performance language models that enterprises can deploy today," said Tim Tully, Partner at Menlo Ventures. "With a track record of pioneering breakthroughs in diffusion models, Inception's best-in-class founding team is turning deep technical insight into real-world speed, efficiency, and enterprise-ready AI."
The speed gains also reduce GPU requirements. Organizations can run larger models at the same latency and cost, or serve more users with existing infrastructure.
CEO and co-founder Stefano Ermon sees the funding as a way to tackle what's becoming AI's primary cost driver. "Training and deploying large-scale AI models is becoming faster than ever, but as adoption scales, inefficient inference is becoming the primary barrier and cost driver to deployment," he explained. "We believe diffusion is the path forward for making frontier model performance practical at scale."
The new capital will accelerate product development, expand the research and engineering teams, and support work on diffusion systems for real-time text, voice, and coding applications.
Speed and efficiency are just the starting point. Inception is working on additional capabilities that diffusion models enable, including built-in error correction to reduce hallucinations and improve reliability. The technology also supports unified multimodal processing, allowing smooth integration of language, image, and code. Precise output structuring is another area of focus, particularly for function calling and structured data generation.
The company was founded by professors from Stanford, UCLA, and Cornell who helped develop some of AI's most important technologies. These include diffusion, flash attention, decision transformers, and direct preference optimization.
Ermon himself is a co-inventor of the diffusion methods that power systems like Midjourney and OpenAI's Sora. The engineering team brings experience from DeepMind, Microsoft, Meta, OpenAI, and HashiCorp.
Inception's models are available through multiple channels: the Inception API, Amazon Bedrock, OpenRouter, and Poe. They function as drop-in replacements for traditional autoregressive models, making integration straightforward for developers.
Early customers are testing the technology for real-time voice interactions, natural language web interfaces, and code generation. The company is based in Palo Alto, California.



