Sarvam AI, India's first AI unicorn, announced at its Epoch 2026 developer conference in Bengaluru (July 30, 2026) that it is building a trillion-plus parameter foundation model from scratch in India — backed by a named frontier researcher, an actual compute roadmap, and pricing that undercuts global rivals by 5.5x. For builders, developers, and businesses evaluating AI infrastructure in 2026, the announcements that matter most are not the headline parameter count but the shipping products behind it: a $0.80-per-million-token inference platform hosted on Indian soil, speech AI covering all 22 scheduled Indian languages, and an on-premise defense-grade stack already in pilot with state governments.
Last verified: July 31, 2026 — Pricing, model specs, and product availability are volatile; Sarvam has not published a launch date for the trillion-parameter model.
TL;DR — what shipped and what's promised
- Shipping now (confirmed): Sarvam 105B inference at $0.80/M blended tokens, Sarvam Inference platform (India-hosted, serves GLM 5.2 and Gemma 4), Saaras V4 speech-to-text (22 languages + multi-speaker), Bulbul V4 text-to-speech with emotional range, Chanakya on-premise AI stack, Sarvam Code, Sarvam Work agents. (Inc42, Moneycontrol)
- Promised (no ship date): One trillion-plus parameter model — roughly 3 trillion total / 100 billion active parameters — designed for coding, cybersecurity, and scientific research. (The Hindu BusinessLine)
- Talent signal: Devendra Singh Chaplot, founding team member at Mistral AI and former xAI pre-training lead, joined as advisor. (Business Standard)
- Funding context: $234 million Series B first close at $1.5 billion valuation (June 2026), led by HCLTech. (Sarvam AI official, TechCrunch)
- Adoption claimed: 1 million+ registered developers, 500+ enterprises, 325 million conversation minutes. (Inc42)
Who is Sarvam AI and why does this matter?
Sarvam AI is a Bengaluru-based full-stack AI company founded in August 2023 by Vivek Raghavan and Pratyush Kumar, both formerly associated with AI4Bharat at IIT Madras. The company builds across the entire AI stack: training infrastructure, foundation models, inference platforms, speech AI, document intelligence, and enterprise applications. (Inc42)
The reason Sarvam matters beyond India is simple: sovereign AI. Governments and enterprises worldwide are increasingly concerned about dependence on foreign AI models for critical workloads. India's IndiaAI Mission has committed ₹10,372 crore (~$1.24 billion) to subsidize GPU compute and domestic model development, and Sarvam is one of 12 selected organisations, receiving ₹246.72 crore in government support. (NiftyTrader) If a company like Sarvam delivers frontier-scale models on Indian infrastructure, it changes the procurement math for every regulated Indian business that needs data residency.
What did Sarvam announce at Epoch 2026?
Sarvam made nearly 20 announcements at its first developer conference spanning models, infrastructure, products, and defense. Here's the full picture.
| Category | Product | What it does | Status |
|---|---|---|---|
| Foundation model | Trillion-plus parameter model | 3T total / ~100B active params; coding, cybersecurity, scientific simulation | Announced, no ship date |
| Inference | Sarvam Inference | India-hosted platform serving Sarvam 105B, GLM 5.2, Gemma 4 | Live |
| Speech recognition | Saaras V4 | STT across all 22 scheduled Indian languages + multi-speaker diarisation | Live |
| Text-to-speech | Bulbul V4 | TTS with emotional expression, laughter, emphasis | Announced |
| Coding | Sarvam Code | Coding agent for dev, ML, cybersecurity work | Invite-only |
| Enterprise | Sarvam Work | Dataset analysis agent, deployable to Slack or on-premise | Announced |
| Voice agents | Samvaad platform | Voice/WhatsApp agents in 11 languages, sub-500ms latency | Live |
| Defense | Chanakya | Air-gapped, on-premise AI stack for defense/government | Launched March 2026 |
| Developer platform | Epoch Builder Edition | GPU access, training tooling, Indian-language datasets | Private preview August 2026 |
Sources: Inc42, ChatMaxima, Explainx.ai, NiftyTrader
How does the trillion-parameter model compare to what exists today?
The trillion-parameter announcement shifts Sarvam into a small global club. For context, here's where frontier models stand as of July 2026:
| Model | Total parameters | Active parameters | Architecture | Built in |
|---|---|---|---|---|
| Sarvam (planned) | ~3 trillion | ~100 billion | Mixture-of-Experts | India |
| Sarvam 105B (current) | 105 billion | 9 billion | Mixture-of-Experts (128K context) | India |
| GPT-5.6 (OpenAI) | Not disclosed | Not disclosed | Dense/MoE (proprietary) | US |
| Gemini 4 (Google) | In pre-training | Not disclosed | Proprietary | US |
| Grok 4.7 (xAI) | Not disclosed | Not disclosed | Proprietary | US |
| Ling 3.0 Flash | 124 billion | ~5.1 billion | MoE | China |
Sources: Sarvam AI, CloudPrice, AI Flash Report
The key engineering detail: Chaplot said a model with roughly 3 trillion total parameters and 100 billion active could be pre-trained in about 2 months using 10,000 NVIDIA Blackwell GPUs, with the full training cycle (including reinforcement learning) taking 4 to 6 months. (Inc42) That's the Mixture-of-Experts architecture at work — only a fraction of the 3 trillion parameters activate for any given token, making training and inference tractable.
What is the compute and funding behind this plan?
Two numbers anchor the feasibility question.
GPUs: Sarvam currently has access to roughly 2,000 NVIDIA Blackwell GPUs, with a stated path to 10,000. (ChatMaxima) The gap between 2,000 and 10,000 is the single best proxy for whether the timeline is real. Watch for data centre commissioning announcements from its sovereign data centre partnership with HCLTech in Odisha.
Funding: Sarvam raised $234 million in the first close of a $300 million Series B at a $1.5 billion post-money valuation, announced June 15, 2026. HCLTech contributed $150 million for a 10.46% stake, with Bessemer Venture Partners also participating. (Sarvam AI, The Hindu, TechCrunch) Prior to this, Sarvam raised approximately $41 million in seed and Series A from Lightspeed, Peak XV, and Khosla Ventures in December 2023. (TechCrunch)
An honest pause: one source reports roughly $12 million in ARR against a commitment to train a trillion-parameter model in six months. (ChatMaxima) IndiaAI Mission compute subsidies and the sovereign data centre partnership explain part of the gap. But reading every forward-looking claim through the revenue-vs-ambition lens is the right discipline.
Who is Devendra Singh Chaplot and why does his hire matter?
Devendra Singh Chaplot joined Sarvam as a part-time advisor, announced at Epoch 2026. His career arc reads like a map of frontier AI labs:
- Mistral AI — founding team member; helped develop Mistral 7B, Mixtral 8x7B, and Mistral Large; led multimodal research; established the company's US office
- xAI — pre-training lead (joined March 2026)
- Thinking Machines Lab (Mira Murati's venture) — Member of Technical Staff
- Facebook AI Research (FAIR) — AI research scientist
Sources: The Hindu BusinessLine, Inc42, Firstpost
Chaplot said at the event: "The model is not the goal. The ability to build models is the goal." He challenged the perception that frontier-scale AI requires resources beyond India's reach. (Inc42)
The important nuance: he joined as a part-time advisor operating from Sarvam's new San Francisco office, not as a relocated full-time pre-training lead in Bengaluru. That's a meaningful commitment-level distinction for any timeline assessment.
How much does Sarvam 105B cost compared to global models?
Sarvam's headline pricing comparison puts its 105B model at roughly 5.5x cheaper than comparable global models:
| Provider | Model | Price ($/M_BLEND tokens) | Hosted in |
|---|---|---|---|
| Sarvam AI | Sarvam 105B | $0.80 | India |
| OpenAI | GPT-5.4 Mini | $4.50 | US |
| Gemini 3.5 Flash | $9.00 | US |
Source: Moneycontrol, ChatMaxima
Three caveats that keep this honest:
- Blended token pricing is a marketing construct. It assumes a particular input-to-output ratio that vendors pick to flatter themselves. Check the underlying input/output split before comparing. (ChatMaxima)
- Cheaper per token is not cheaper per resolution. A model that costs a fifth as much but needs three extra turns to reach the same answer has erased its own advantage. The only metric worth optimising is cost per resolved conversation. (ChatMaxima) — see our AI token cost optimization guide for the full framework.
- Inference cost is rarely the dominant cost in a support stack. Agent salaries, CRM licenses, telephony minutes, and WhatsApp conversation charges typically dwarf token spend. The 5.5x saving matters at genuine scale or in token-heavy workloads (bulk document extraction, transcript summarisation, agentic reasoning loops).
If you want to understand how AI inference bills actually work and how to cut them 3-10x, our enterprise AI token cost optimization guide breaks down the math with real workload examples.
What are the speech AI upgrades and why do they matter for Indian businesses?
The product announcements most likely to change what you can ship this quarter are not the trillion-parameter model — they're the speech upgrades.
Saaras V4 (speech recognition):
- Covers all 22 scheduled Indian languages, including Odia, Sanskrit, and Manipuri that global models can't transcribe
- Adds multi-speaker transcription for overlapping conversations
- Claims state-of-the-art performance on global English benchmarks
Bulbul V4 (text-to-speech):
- Adds emotional expression: laughter, excitement, emphasis
- Focuses on natural vocal range and delivery style, not just correct pronunciation
- The v3 baseline supports 10 Indian languages + English across 30+ curated voices
Sources: NewsBytes, Explainx.ai, HeadsUpAI
The reason this matters: global speech models handle Indian English acceptably, Hindi passably, and fall apart on everything else — especially code-mixed speech. Real Indian customers speak Hinglish, switch languages mid-sentence, and use English nouns inside Tamil grammar. A speech model trained on Western data transcribes that as noise, and the downstream AI agent responds to garbage. Sarvam's India-first training data is the core differentiator here, not the model size. (ChatMaxima)
What is Chanakya and why is defense-grade AI a separate product line?
Chanakya is Sarvam's dedicated applied-AI vertical for air-gapped, on-premise deployments in defense, government, and critical infrastructure. Launched March 29, 2026, it sits atop Sarvam's existing model stack (Sarvam-30B and Sarvam-105B) and is explicitly designed for environments where public cloud is not an option. (NiftyTrader, Particle News)
Three capabilities define it:
- Air-gapped deployment — zero external network access; sensitive data never leaves the facility
- Multimodal ingestion — processes both text and images
- Production-grade agentic workflows — multi-step AI pipelines designed for environments where failure is not acceptable
As of April 2026, no signed defense contract has been disclosed. The vertical's defense positioning is based on its technical architecture and stated dual-use design intent. For businesses in BFSI, healthcare, and government-adjacent sectors, the existence of an air-gapped Indian AI stack means the data residency objection that has blocked many AI deployments now has a credible answer. For a broader look at how India is defending critical infrastructure from cyberattacks, see our critical infrastructure cybersecurity guide.
What developer tools launched alongside the models?
Sarvam announced several developer-facing products:
Sarvam Inference — an India-hosted inference platform serving Sarvam 105B alongside frontier open models including GLM 5.2 and Gemma 4. The company claimed an agentic optimisation system delivering up to 15x inference speed improvement for certain models. (Inc42) This vendor-quoted figure needs third-party verification.
Sarvam Code — a coding agent aimed at software development, ML, and cybersecurity work. It features long-running work with checkpoints, steering, and a pay-for-completed-work pricing model. CLI and graphical access are in an invite-only phase. Benchmark claims involving GLM-5.2 and Terminal-Bench have been circulated but not independently reproduced. (Explainx.ai)
Epoch Builder Edition — a platform giving developers and enterprises GPU capacity, training tooling, curated Indian-language datasets, and safety testing to build India-centric models. Private preview opens August 2026, with wider rollout planned for Q4 2026. (ChatMaxima)
What this means for you
If you're a developer in India: Sarvam Inference gives you frontier-class open models (Sarvam 105B, GLM 5.2, Gemma 4) serving from Indian infrastructure with sub-500ms latency. The Epoch Builder Edition (private preview August 2026) could let you build and fine-tune India-specific models without exporting your data. Start with the developer platform and test one workload.
If you're a business deploying AI agents: The $0.80/M token pricing is your benchmark for your next vendor renewal conversation — even if you never switch. More importantly, if data residency has blocked your AI projects, Sarvam's on-premise options (Chanakya for defense/government; managed private cloud for enterprises) are now a credible answer. Read our AI agent infrastructure guide for broader context on serving models at scale.
If you're evaluating the trillion-parameter model for production planning: Don't. No ship date has been given, and Sarvam itself did not disclose a timeline at Epoch 2026. (Inc42) Plan around what has shipped — Sarvam 105B inference, the speech stack, the agent platform. Watch what has been promised — the GPU ramp toward 10,000 and Chaplot's involvement level.
If you're a global AI strategist: India's sovereign AI push is not just a national narrative. It's backed by ₹10,372 crore in government compute subsidies, a $1.5 billion private company, and models trained on 22 Indian languages that global frontier models do not cover well. The gap between global and Indian models on code-mixed Indic language tasks is real — and closing fast. For context on how Chinese AI labs are pursuing a similar sovereign strategy, see our Kimi K3 business guide.
FAQ
Q: When will Sarvam AI's trillion-parameter model launch?
A: Sarvam has not disclosed a launch date. At Epoch 2026 (July 30, 2026), co-founder Pratyush Kumar confirmed the company is building the model but did not give a timeline. Advisor Devendra Singh Chaplot estimated that pre-training could take about 2 months with 10,000 Blackwell GPUs, with the full training cycle lasting 4-6 months. (The Hindu BusinessLine, Inc42)
Q: How much does Sarvam 105B cost?
A: Sarvam prices its 105B model at $0.80 per million blended tokens, compared to $4.50 for OpenAI's GPT-5.4 Mini and $9.00 for Google's Gemini 3.5 Flash — roughly 5.5x cheaper than the nearest comparable option. Blended pricing assumes a specific input-output ratio, so check the underlying split before comparing. (Moneycontrol)
Q: What is the Mixture-of-Experts architecture used in Sarvam's models?
A: MoE models have a large total parameter count but only activate a fraction of those parameters for any given token. Sarvam's current 105B model has 105 billion total parameters but only 9 billion active per token, with a 128K context window. The planned trillion-parameter model would have roughly 3 trillion total parameters and about 100 billion active. This makes training and inference more efficient than a dense model of the same total size. (CloudPrice, AI Flash Report)
Q: Can Sarvam's AI be deployed on-premise for sensitive workloads?
A: Yes, through the Chanakya vertical, launched March 29, 2026. Chanakya supports air-gapped deployments in facilities with zero external network access, processing text and image data through Sarvam's model stack entirely on-premise. It is designed for defense, government, and regulated enterprises that cannot use public cloud infrastructure. No signed defense contracts have been disclosed as of April 2026. (NiftyTrader)
Q: Do Sarvam's speech models support Indian languages that global models can't?
A: Saaras V4 covers all 22 scheduled Indian languages, including Odia, Sanskrit, and Manipuri that most global speech recognition models cannot transcribe reliably. Bulbul V4 (text-to-speech) adds emotional expression on top of the v3 baseline's 10 Indian languages plus English. The key advantage is training on code-mixed speech (Hinglish, Tanglish) that real Indian users actually speak. (NewsBytes, Explainx.ai)
Q: Is Sarvam AI profitable or is it burning through funding?
A: Sarvam has not publicly disclosed profitability. One report estimated roughly $12 million in annual recurring revenue, with conversational AI accounting for about 80% of that. The company raised $234 million in Series B funding (June 2026) at a $1.5 billion valuation, led by HCLTech. Government compute subsidies under the IndiaAI Mission also offset infrastructure costs. (ChatMaxima, TechCrunch)

Discussion
0 comments