
Deep Cogito builds systems that use large scale reinforcement learning to help AI models keep improving after their initial training is finished. The company was founded by Drishan Arora and Dhruv Malrana, who previously helped build Google's AI Search products, including AI Mode and AI Overviews. Their idea is simple to state, even if the engineering behind it is not. Once a model has been trained on a huge amount of data, a second stage called post-training decides what that model can actually do with what it knows.
That focus fits a pattern showing up across the AI industry. Many labs are now putting more computing power into this second stage, using reinforcement learning to teach models how to work through harder problems, rather than pouring all their resources into the first stage of training. A recent industry analysis found that Cursor has disclosed spending more compute on this kind of post-training than on the initial training of its Composer 1.5 model. The same analysis pointed to similar moves at OpenAI and Anthropic as part of a wider shift, describing post-training as the new competitive edge among AI labs in 2026.
Deep Cogito has now raised $43 million in a Series A round led by TQ Ventures. Benchmark, Nexus Venture Partners, Atreides Management, and South Park Commons also took part, along with Zscaler, a cloud security company that is both a customer of Deep Cogito and a strategic investor in the round. The new funding brings the company's total raised to more than $56 million.
This shift in how models are trained is changing what investors look for. Instead of judging a company mainly by how much data or computing power it can access, some investors now pay closer attention to whether a team knows how to refine a model once it already exists. That is the case Deep Cogito is making to its backers. As models grow more capable, the company argues, the differences between labs will come down less to the size of their initial training and more to what happens afterward.
The company points to its own track record as evidence. Through its open weight Cogito models, which range from 3 billion to more than 600 billion parameters, Deep Cogito says it has already shown that its post-training methods can improve strong models at many different sizes. That history is part of what it is now offering to enterprise clients, and part of what drew its new investors in.
"Very few teams outside the largest AI labs have demonstrated the ability to post-train models at this scale," said Schuster Tanger, Co-Founding Partner at TQ Ventures. "Deep Cogito has done that in public through its model releases, and is now bringing the same capability to companies that want intelligence built around their own products. We believe that combination of frontier research and real-world deployment is extremely powerful."
The company plans to put the money into three main areas. It will expand its research and engineering team, scale up the computing infrastructure it needs to train models at frontier scale, and keep developing new releases in its Cogito model family. Alongside that research, Deep Cogito wants to grow its work with enterprises that need AI models built around their own data and outcomes, not just general purpose tools.
Zscaler's experience helps explain why a company might take that route. According to Zscaler, general purpose AI models were useful, but not precise enough for the level of specialization its security products required. That is why it worked directly with Deep Cogito rather than adopting an off the shelf model.
"Frontier models were useful, but they were not enough for the level of specialization we needed," said Dhawal Sharma, Executive Vice President of AI Security and Strategic Initiatives at Zscaler. "Deep Cogito stood out because they went deeper than lightweight customization. They worked closely with us to understand our products and the metrics we care about and helped train that intelligence into the model itself."
Deep Cogito is based in San Francisco. Arora and Malrana worked together at Google before starting the company, where Arora led post-training work for Gemini within Google's AI Search products and Malrana led product for that same effort from the beginning. In plain terms, the company takes AI models that already hold a lot of general knowledge and teaches them to reason better. It does this by training them on progressively harder problems and, over time, feeding their own improvements back into the model.
One method the company uses is called Iterated Distillation and Amplification. In this process, a model is given extra computing time to work out better answers than it could produce right away. Those improved answers are then folded back into the model, so the gains stick around instead of disappearing once the extra computation ends. Deep Cogito's longer term goal is for models to keep improving this way on their own, eventually going beyond what can be learned from human written text alone. The same system behind its open weight Cogito models is now being adapted for enterprise clients who want a similar process applied to their own products.
"Pre-training gives a model an enormous amount of knowledge and capability. Post-training determines what that model can actually become," said Drishan Arora, co-founder and CEO of Deep Cogito. "We believe the next frontier is in finding ways for models to improve their own intelligence, internalize those improvements, and become increasingly capable over time."
TQ Ventures led the Series A, joined by Benchmark, Nexus Venture Partners, Atreides Management, and South Park Commons. Zscaler also joined the round. It began as a customer of Deep Cogito before becoming a strategic investor. Together, this group has brought the company's total funding past $56 million.
Benchmark, one of the new participants, said its decision to invest came down to both the ambition of what Deep Cogito is trying to build and the technical work it has already shown through its public model releases.
"What stood out to us was not only the technical depth of the team, but the scope of what they're trying to build," said Eric Vishria, General Partner at Benchmark. "Post-training is becoming one of the most important layers in AI. Deep Cogito has demonstrated that it can operate at the frontier of that layer and translate that capability into intelligence that companies can actually own."