
DeepInfra, a cloud platform built specifically to run AI inference at high volume, has raised $107 million in a Series B round. The company was founded in 2022 by the team behind imo, a messaging app with over 200 million users worldwide, and has designed its infrastructure from scratch to handle the kind of continuous, large-scale AI workloads that most cloud platforms were not built to support.
AI is no longer mostly about building models. It is increasingly about running them. Most cloud platforms were designed with general computing in mind, not the sustained, high-frequency demands of production AI. As businesses move beyond pilots and into full deployment, the cost and reliability of running AI around the clock have become pressing problems. The global AI inference market was valued at around $104 billion in 2025 and is projected to reach $312 billion by 2034, according to Fortune Business Insights. The round was co-led by 500 Global and Georges Harik, one of Google's earliest engineers, with participation from A.Capital Ventures, Crescent Cove, Felicis, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90.
A standard AI interaction involves a single question and a single response. Agentic AI systems work differently. They operate on their own, breaking complex tasks into steps and making 50 to 100 or more separate model calls to complete a single job. That creates a very different kind of demand: continuous, high-volume, and difficult to predict in advance.
Open-source models are also changing the picture for enterprise buyers. Models such as Llama, Mistral, and NVIDIA's Nemotron family have improved substantially, and more companies are now choosing them over proprietary alternatives to keep costs down and avoid lock-in to a single vendor. DeepInfra processes nearly five trillion tokens per week and estimates that around 30% of that volume already comes from agent-based systems. That gives the company a live view of where AI infrastructure demand is heading.
"Demand for AI is causing every layer of the AI stack to innovate, and inference is no exception. In the agentic age, new workflows are arising on a rapid basis, as evidenced recently by OpenClaw and AutoResearch. Enterprises and developers building with open source and agent-driven AI need infrastructure that was designed to be flexible, fast and reliable. We backed DeepInfra because, in our assessment, this team has already proven they can build and operate distributed systems at global scale, and because we believe purpose-built inference infrastructure will be fundamental to the next phase of AI as compute was to the last." — Tony Wang, Managing Partner, 500 Global.
The new capital will go toward expanding compute capacity internationally, improving the tools developers use to build on the platform, and extending support for new models as they become available. DeepInfra currently runs GPU infrastructure across eight U.S. data centers and has more locations in development.
The platform supports over 190 open-source models through OpenAI-compatible APIs. In practice, that means a developer already building with OpenAI's tools can switch to DeepInfra without rewriting their code. The company has also worked closely with NVIDIA to support the Nemotron model family, the NemoClaw agent framework, and NVIDIA Dynamo distributed-inference software. Early access to Blackwell and Vera Rubin GPU architectures has contributed to what the company reports as a 20x improvement in inference cost efficiency.
Revenue has tripled since the start of 2026, which DeepInfra attributes to growing demand from agentic AI workloads. The funding gives the company room to scale hardware procurement and expand geographically at a time when production-grade infrastructure has become a deciding factor for enterprise AI projects.
DeepInfra was founded in September 2022 by the team that previously built and operated the backend infrastructure for imo, a messaging platform with over 200 million users. Running infrastructure at that scale requires the kind of operational discipline that is hard to acquire any other way, and the founders carried that experience directly into building DeepInfra.
CEO and co-founder Nikola Borisov has said the company was built around the belief that inference, not model training, would eventually become the main bottleneck in enterprise AI. The platform is vertically integrated: DeepInfra owns and operates its own GPU hardware rather than leasing capacity from larger cloud providers. That ownership model gives the company more control over pricing, latency, and reliability than providers who depend on rented or spot capacity. The platform includes enterprise security features, zero data retention, and holds both SOC 2 and ISO 27001 certifications.
"When we launched nearly four years ago, we believed inference would become the dominant driver of enterprise AI workloads – and we are now at this inflection point. What's happening now is incredibly exciting – open-source models are rapidly reaching parity with proprietary systems, unlocking a new wave of innovation at a fraction of the cost and enabling widespread adoption. At the same time, agent-based systems are driving continuous, high-volume demand. Inference is no longer a thin layer – it's the system constraint that will define the majority of workloads." — Nikola Borisov, Co-founder and CEO, DeepInfra.
The Series B was co-led by 500 Global and Georges Harik, an early Google engineer. 500 Global is a multi-stage venture capital firm managing $2.2 billion in assets, with a portfolio of more than 5,000 founders operating across 80 countries.
The rest of the round includes A.Capital Ventures, Crescent Cove, Felicis, NVIDIA, Peak6, Samsung Next, Supermicro, and Upper90. NVIDIA's participation carries particular weight given that DeepInfra is already working with its hardware and agent frameworks in production.



