The Sarvam AI Origin Story: A New Era for Indian Tech
What Is the Sarvam AI Origin Story, in Brief?
Sarvam AI was founded in August 2023 by two researchers who had each spent over a decade on opposite halves of the same problem: Vivek Raghavan building digital infrastructure for a billion Indians, and Pratyush Kumar measuring exactly how broken Western AI was for Indian languages. That is the Sarvam AI origin story at its core — not a startup trend, but a structural inevitability.
Sarvam AI is a Bengaluru-based, full-stack sovereign AI company founded by Vivek Raghavan and Pratyush Kumar in August 2023 to build large language models, speech systems, and AI infrastructure specifically for India’s 22 official languages and 1.4 billion people. Within five months, it closed $41 million — the largest early-stage raise for an India sovereign LLM startup at the time. By April 2025, the Indian government had selected it from 67 applicants to build India’s first sovereign model under the IndiaAI Mission. By June 2026, Sarvam crossed a $1.5 billion valuation, becoming India’s newest AI unicorn. The numbers are striking. What sits beneath them is more interesting.
Two Parallel Careers That Made One Sarvam AI Origin Story Inevitable
Most AI startup founding stories start with a ChatGPT demo and a pitch deck. The Sarvam AI origin story starts much earlier, on two separate tracks — which is what distinguishes the Vivek Raghavan Pratyush Kumar AI company from every other Indian LLM venture launched in the post-GPT-4 wave.
Vivek Raghavan: Twelve Years Building Population-Scale Infrastructure at Aadhaar
Raghavan holds a B.Tech in Electrical Engineering from IIT Delhi and a PhD in Electrical and Computer Engineering from Carnegie Mellon University. After graduating, he built two Electronic Design Automation companies over nearly two decades — software that enables chip design — with Nvidia among his early customers. His return to India in 2007, following the passing of his mother, redirected everything.
At the UIDAI, he served as Chief Product Manager and Biometric Architect on Aadhaar — the world’s largest biometric identity program. The engineering work was specific and brutal: reliably enrolling and authenticating over a billion people via fingerprint and iris scans, on infrastructure that had to hold up in rural government offices with intermittent connectivity. He did this largely without meaningful financial compensation for years.
After Aadhaar, he moved through a sequence of roles at the intersection of digital public infrastructure and language AI: Chief AI Evangelist at EkStep Foundation, advisor to the National Language Translation Mission (Bhashini), Chief Mentor at the Nilekani Center at AI4Bharat at IIT Madras, and a member of the Supreme Court’s AI committee overseeing the SUVAS judgment translation initiative. That resume is not a list of prestigious positions. It is twelve years of preparation in how to build technology that works for everyone — including people with bad connectivity, non-English scripts, and no tolerance for systems that fail.
Pratyush Kumar and the AI4Bharat Founders Startup Journey
Kumar’s parallel track was technical and rigorous. He earned his B.Tech in Electrical Engineering from IIT Bombay, then his PhD in Computer Engineering from ETH Zurich in 2014. After three years at IBM Research and two at Microsoft Research, he joined IIT Madras as adjunct faculty. There, he co-founded AI4Bharat — a research lab with one specific mandate: build open-source AI tools, datasets, and models for India’s languages.
The AI4Bharat founders startup journey produced work that most breathless startup profiles have never cited once. The lab built systematic open datasets for machine translation, automatic speech recognition, and optical character recognition across more than 20 Indian languages. It released IndicBERT, IndicXTREME, and a series of benchmarks that constituted the first rigorous measurement of how badly Western models performed on Indic language tasks. That measurement work was not advocacy. It was peer-reviewed documentation of a structural problem, completed before that problem was commercially interesting to anyone outside India.
Peak XV Partners’ managing director described their combination as making them “uniquely positioned to build population scale AI applications” — a read-out that was technically accurate, not investor hype. Raghavan understood infrastructure at population scale. Kumar understood exactly where the language gap was quantifiable, measurable, and not closing on its own.
What Is the Tokenization Tax and Why Does It Matter for Indian Language AI?
The tokenization tax in Indian language AI is the most important concept for understanding why a homegrown AI model for Indian languages is structurally necessary — not politically desirable, structurally necessary.
Every large language model processes text by converting it into tokens — small chunks the model actually reads. Tokenizers are trained on data. When that training data is overwhelmingly English, the tokenizer becomes highly efficient at English and highly inefficient at everything else. A common English word maps to roughly one token. A Hindi or Kannada word on a Western tokenizer maps to eight or more.
A 2025 research paper quantifying this problem found that under the cl100k_base tokenizer used by GPT-3.5 and GPT-4, Indian languages experience an average 8.0x tokenization tax relative to English — reaching 13.0x for Malayalam — leaving users with as little as 12% of the usable context window available to English speakers for equivalent semantic content. That is not a performance gap. It is a wall.
The downstream commercial consequences compound immediately. Every API call in Hindi costs more than the equivalent call in English, for identical information. Context windows shrink, so the model processes less document per request. Latency increases. For a business in rural Maharashtra building a voice-based loan repayment reminder in Marathi, those compounding costs make Western frontier models economically indefensible — before considering quality.
Sarvam’s approach was native tokenization from first principles. Sarvam-1’s tokenizer achieved near-English token fertility rates of 1.4 to 2.1 tokens per word for Indic scripts, compared to the 8x-plus penalty imposed by Western alternatives. That single architectural decision alone justifies building a model from scratch. It is not nationalism. It is arithmetic.
Why August 2023? The Four Conditions That Forced the Founding Decision
Sarvam was not an opportunistic play on a ChatGPT trend. Four structural conditions converged precisely in 2023, and any one of them alone would have been insufficient. Together, they made the August 2023 founding close to overdetermined.
Vernacular internet scale had reached critical mass. India had hundreds of millions of smartphone users accessing the internet primarily in regional languages. The English-first internet assumption was already broken before Sarvam’s incorporation documents were filed.
Aadhaar and UPI had proved something more significant than their individual functions. They demonstrated that India could build and scale population-level digital public infrastructure — not just consume it from abroad. A biometric identity system for over a billion people and a payments network processing hundreds of millions of daily transactions gave investors and policymakers a credible precedent for the same approach applied to AI.
Voice had become the natural default interface for non-English-first users. Text prompts are a Western affordance built for people who type fluently in English. For hundreds of millions of Indians, speaking is faster, more natural, and more accessible than switching keyboard scripts. Any model that couldn’t handle voice-first interaction in Indian languages couldn’t reach most of its intended users.
Finally, and most decisively, LLMs had made language itself the primary interface between humans and software. Before the LLM era, language gaps could be patched — build an English backend, add a translated UI. After GPT-3 and its successors, the underlying model is the product. A model that processes Hindi at 12% efficiency compared to English doesn’t just perform worse. It fails structurally. Building a dedicated homegrown AI model for Indian languages stopped being idealistic and became the only technically sound option.
From $41 Million to a $1.5 Billion Unicorn: How Sarvam AI Scaled in Under Three Years
Five months after founding, Sarvam announced $41 million across seed and Series A rounds in December 2023, led by Lightspeed Venture Partners, with Peak XV Partners and Khosla Ventures participating — the largest early-stage raise for an Indian AI startup at the time.
The Series B, announced on June 15, 2026, validated the execution behind that thesis. Sarvam raised $234 million at a $1.5 billion valuation, becoming India’s newest AI unicorn. HCLTech anchored the round with $150 million for a 10.46% stake. Bessemer Venture Partners joined as a new institutional investor. Khosla and Peak XV returned.
The deployment numbers announced alongside the raise tell the real story of what justified the valuation step-up. Sarvam’s conversational platform handles over 2 million interactions daily — a figure that doubled in two months. Its inference platform processes 10 million API calls per day, tripling in three months. Speech models transcribe over 500,000 hours of audio monthly. Document AI systems digitize over 35 million pages of records. A leading fintech deploys Sarvam’s agentic platform to support a 350,000-person sales network. These are not pilot program metrics. They are commercial numbers from a company operating AI at state scale.
HCLTech’s decision to anchor this round matters beyond the dollar figure. A company with $14.7 billion in annual revenue and 227,000 employees taking a double-digit equity stake in a three-year-old domestic AI startup is an enterprise infrastructure positioning decision — not a financial bet. The plan is to combine Sarvam’s models with HCLTech’s enterprise relationships, engineering depth, and global client base.
How Was Sarvam AI Selected by the Government to Build India’s Sovereign LLM?
On April 26, 2025, MeitY announced that Sarvam had been selected under the IndiaAI Mission to build India’s sovereign large language model — the first company chosen from 67 applicants. That IndiaAI Mission sovereign model selection delivered access to 4,096 NVIDIA H100 SXM GPUs hosted at Yotta’s Shakti cluster, approximately ₹99 crore in compute subsidies. In exchange, the government took an equity stake, and the sovereign model must be built, deployed, and governed within India’s borders.
That compute access was transformative. Training a 105B-parameter model without those GPUs would have been financially prohibitive for a startup at Sarvam’s pre-Series B stage. The government mandate converted Sarvam from a well-funded startup into national infrastructure — a status change that reshaped both investor perception and enterprise sales conversations overnight.
Is Sarvam AI’s Model Truly Built from Scratch, or Fine-Tuned from Western AI?
This is the question current top-ranked pages consistently dodge. The honest answer requires telling two separate stories.
In May 2025, Sarvam released Sarvam-M — a 24-billion-parameter model post-trained on Mistral Small, a base architecture from French AI firm Mistral. The reaction was fast and unsparing. Deedy Das of Menlo Ventures called the release “embarrassing,” citing only 23 downloads in two days after it went live on Hugging Face. Critics labeled it a wrapper — Indic post-training layered on a French foundation — and asked the question that cut deepest: how can a model built on foreign weights be called sovereign? Zoho’s Sridhar Vembu publicly defended Sarvam, arguing that post-training is a substantive engineering discipline, not a shortcut. The criticism had already landed.
The 105B model was Sarvam’s technical answer. Unveiled at the India AI Impact Summit in February 2026 — on stage alongside Prime Minister Modi, covered by Bloomberg, TechCrunch, and Business Standard — Sarvam-105B was trained from scratch using compute provisioned through the IndiaAI Mission at Yotta’s Shakti cluster. The company’s own technical documentation confirms that all stages of the training pipeline were developed and executed in-house: model architecture, data curation pipelines, reasoning supervision frameworks, and reinforcement learning infrastructure. All of it.
The architecture is substantive. The model uses sparse mixture-of-experts design with 128 experts per layer, activating approximately 9 to 10 billion parameters per token from a 105-billion total — keeping per-token compute costs comparable to a dense model roughly one-third the size. Its 128,000-token context window makes it practical for large document analysis: balance sheets, legal filings, government records. The companion Sarvam-30B, trained on 16 trillion tokens, handles real-time tasks where latency matters more than depth.
Sarvam’s three-part sovereignty definition — data, compute, utility — now holds for the 105B model. Training data is curated in India, compute is Indian-hosted, and the model targets Indian deployment contexts. The unresolved question, per analysts who’ve examined this closely, is whether enterprise revenue will scale fast enough to sustain training costs once the government compute subsidy ends. The H100 allocation is time-limited. The next frontier model run must be funded by commercial contracts, not public subsidies.
One additional debate: both models were released under Apache 2.0, downloadable by anyone globally. Some critics ask whether a globally available open-weight model is truly sovereign. Sarvam’s position — that sovereignty means controlling the capacity to build and train, not gatekeeping the output — is coherent. Whether that satisfies every policy stakeholder is a separate question worth watching.
How Sarvam AI Compares to ChatGPT and Google Gemini for Indian Languages
Any useful comparison requires defining what’s actually being compared. Sarvam-105B versus GPT-4o on a general English reasoning benchmark is a category error. Sarvam-105B versus GPT-4o on a Hindi document understanding task with code-mixed queries and data residency constraints is an entirely different question.
On Indic-specific benchmarks, Sarvam’s published claims are concrete. The 105B model reportedly outperforms DeepSeek R1 and Gemini 2.5 Flash on the Tau 2 benchmark for reasoning and agentic tasks. On OCR and document layout understanding for Indian scripts, the models show measurable advantages over general-purpose Western alternatives. For code-mixed language — the natural way millions of Indians blend Hindi and English in a single sentence — Sarvam’s training data reflects actual Indian usage patterns in ways American models’ training corpora cannot replicate by design.
Where OpenAI and Google retain genuine advantages: raw reasoning on global English benchmarks, coding across international programming tasks, tool ecosystem maturity, and developer mindshare accumulated over years. A startup building a global SaaS product in English is probably better served by GPT-4o. A government ministry in Tamil Nadu deploying citizen-facing voice services in Tamil, with strict data residency requirements and no option to route sensitive records through American APIs, has no realistic alternative at Sarvam’s price point and language fidelity.
That domain specificity is a structural position, not a consolation prize. Agricultural data collection, insurance policy renewals, KYC workflows in regional languages, court judgment translation — at India’s scale, these are not niche applications. They are the primary use cases for AI across the country.
What the Sarvam AI Origin Story Signals for India’s Next Technology Era
India has a documented pattern for building foundational technology at population scale: identity rails first through Aadhaar, then payment rails through UPI, then document and data infrastructure. The Sarvam AI origin story fits that sequence almost too cleanly — intelligence rails, built by a researcher who designed the first layer and understood its architecture from the inside.
Raghavan’s years at UIDAI were not incidental background. They were specific preparation. The philosophy that produced a biometric identity system for a billion people — build at the foundational layer, treat cost as a design constraint from day one, deploy at population scale before declaring success — runs directly through Sarvam’s approach to model training, API pricing, and voice-first product architecture.
HCLTech anchoring the Series B with $150 million signals that India’s established technology sector now treats domestic sovereign AI as infrastructure worth owning equity in, not a product category to procure from foreign vendors. When a company with 227,000 employees takes a double-digit stake in a three-year-old AI startup, that is not a venture bet. It is an infrastructure positioning decision with a long horizon.
For multilingual markets in the Global South watching this experiment — Nigeria, Indonesia, Brazil — the Sarvam AI origin story offers a replicable template: start with open-source language research, measure the performance gap rigorously before claiming to close it, build on existing digital public infrastructure, and treat sovereignty as a sustained engineering commitment rather than a founding-day press release.
The full Series B close will bring total capital raised to approximately $275 million. What that capital builds over the next 24 months — specifically whether commercial revenue scales fast enough to sustain frontier model training without government subsidy — is the next chapter. The founding chapter now has the technical and commercial evidence to stand behind it. The scale chapter is still being written.
Frequently Asked Questions
Who founded Sarvam AI and what is their academic and research background?
Sarvam AI was co-founded in August 2023 by Dr. Vivek Raghavan, an IIT Delhi and Carnegie Mellon PhD graduate who served as UIDAI’s Chief Product Manager and Biometric Architect on Aadhaar, and Dr. Pratyush Kumar, an IIT Bombay and ETH Zurich PhD graduate who previously worked at IBM Research, Microsoft Research, and co-founded AI4Bharat at IIT Madras.
What is the tokenization tax and why does it matter for Indian language AI?
The tokenization tax is the performance and cost penalty imposed on non-English languages by tokenizers trained predominantly on English data. Under GPT-3.5 and GPT-4 tokenizers, Indian languages face an average 8.0x token penalty relative to English, reaching 13.0x for Malayalam. This means higher API costs, sharply reduced context windows, and degraded output quality for every Indian-language application using Western models.
How did Sarvam AI go from founding to a $1.5 billion unicorn in under three years?
Sarvam closed $41 million in December 2023, five months after founding — India’s largest early-stage AI raise at the time. In April 2025, the Indian government selected it from 67 applicants for the sovereign LLM mandate, granting 4,096 NVIDIA H100 GPUs. In June 2026, HCLTech led a $234 million Series B at a $1.5 billion valuation, validated by deployments reaching 17 million farmers and 45 million insurance policyholders.
How was Sarvam AI selected by the government to build India’s sovereign LLM?
On April 26, 2025, MeitY selected Sarvam from 67 IndiaAI Mission applicants as the first startup to build India’s indigenous foundational model. The selection delivered 4,096 NVIDIA H100 SXM GPUs via Yotta Data Services, approximately ₹99 crore in compute subsidies, and a government equity stake. Sarvam’s research track record through AI4Bharat and the founders’ digital public infrastructure experience were central to that selection decision.
Is Sarvam AI’s model truly built from scratch or fine-tuned from Western AI?
Both, depending on which model. Sarvam-M (May 2025) was post-trained on Mistral Small, a French base model, drawing sharp criticism for lacking true indigenous origins. Sarvam-105B (February 2026) was trained from scratch on Indian compute under the IndiaAI Mission, with architecture, data curation, and reinforcement learning all executed in-house — making it Sarvam’s first genuinely sovereign model by its own three-part definition of data, compute, and utility.
How does Sarvam AI compare to ChatGPT and Google Gemini for Indian languages?
On Indic benchmarks, Sarvam-105B reportedly outperforms DeepSeek R1 and Gemini 2.5 Flash on the Tau 2 reasoning benchmark, and leads on OCR and document understanding for Indian scripts. OpenAI and Google retain clear advantages on global English benchmarks, coding, and developer tooling. Sarvam wins on Indian language fidelity, data residency compliance, and cost efficiency for Indian government and enterprise use cases.
What role did AI4Bharat and Aadhaar play in shaping the Sarvam AI vision?
AI4Bharat, co-founded by Pratyush Kumar at IIT Madras, built the open-source datasets and benchmarks that first rigorously documented the performance gap between English and Indian-language AI systems. Aadhaar, where Vivek Raghavan served as Chief Product Manager and Biometric Architect, established the “build at the foundational layer, deploy at population scale” philosophy that defines Sarvam’s entire technical and commercial approach from day one.
What is the IndiaAI Mission and how does Sarvam fit into India’s AI strategy?
The IndiaAI Mission is a government initiative backed by ₹10,371 crore to build India’s sovereign AI infrastructure, including compute access, model development, and domestic AI capability. Sarvam was selected as the first startup under the mission to train an indigenous foundational model, positioning it as the model layer for India’s broader digital public infrastructure — analogous to how Aadhaar became the identity layer and UPI became the payments layer.
