About
Gradium builds audio language models and infrastructure for real-time voice applications, selling primarily to developers and enterprises. Its differentiation is natural, expressive voice interaction with ultra-low latency, spanning text-to-speech, speech-to-text, voice cloning, and related voice tasks.
Market
Gradium competes in the developer and enterprise voice-AI infrastructure market, supplying foundational speech models and APIs for real-time agents, transcription, synthesis, translation, and conversational applications. Its positioning emphasizes an audio-native model architecture, expressive multilingual output, and production-scale, low-latency WebSocket streaming; it differentiates from broader TTS providers and large model makers through tightly integrated TTS/STT services, stable latency, voice cloning, on-device deployment, and enterprise/on-premises options.
Gradium primarily targets developers and product or engineering teams at startups, scale-ups, and enterprises building real-time voice experiences, especially in customer support, healthcare, gaming, media, robotics, and automotive. Its enterprise offering also targets organizations needing scalable API access, usage-based plans, unlimited contracts, or on-premises deployment.
At a Glance
Problem
Voice applications are difficult to make feel genuinely conversational when model responses are delayed or systems cannot support interactions at scale. Gradium targets this developer pain by focusing on ultra-low-latency voice AI, where responsiveness and scalable delivery are central to the usefulness and economics of voice interfaces. The clearest killer use case is real-time voice applications and agents that need fast, natural back-and-forth interaction.
Product / Service
Gradium provides foundational voice AI models and infrastructure for developers. Its offering includes real-time speech-to-text, text-to-speech, and voice-generation capabilities, delivered as a real-time voice AI platform rather than simply as a standalone research model. The intended benefit is to give developers the building blocks for voice products that respond quickly and can operate at scale.
Market
Gradium competes in the emerging real-time voice AI and generative-audio infrastructure market, alongside platforms that provide low-latency speech models and developer APIs. Its own positioning includes a comparison to Cartesia, making Cartesia the clearest named competitor in the available evidence, although the research does not establish a complete competitor set.
The company was founded in September 2025 by researchers associated with Kyutai and had raised $100 million in seed funding by July 2026, with backing from Nvidia as well as FirstMark Capital, Eurazeo, and DST Global Partners. That financing is meaningful early traction, but the available research does not report revenue, customer counts, or usage, so it is not possible to determine whether Gradium is pre-revenue or already generating commercial sales.
Founders & Leadership
Funding History
FirstMark Capital, Eurazeo
NVIDIA
Recent News
Gradium selected Stripe for billing, payments, invoicing, and tax compliance as it scaled internationally after launching its voice models. The integration reportedly took one week and supports products including Gradium Translate and Phonon.
Agora and Gradium announced an integration that will make Gradium’s low-latency text-to-speech available in Agora’s Conversational AI Engine. The integration is intended to support responsive voice AI, voice cloning, and cloud, private-cloud, and on-premises deployments.
Gradium reopened its seed round and reached $100 million in total funding with new investors including NVIDIA. The company plans to use the capital to expand AI research, product development, international operations, and a Bay Area office.
Gradium announced that its extended seed financing brought total funding to $100 million, with NVIDIA among the new investors. The funding will support research, product development, international expansion, and a new San Francisco Bay Area office.
Gradium announced that it powers RMC BFM Drive, an AI-generated personalized radio experience available in Renault connected vehicles since March 2026. Its text-to-speech models generate real-time transitions using voice clones of RMC BFM journalists.
InteractionLabs and Gradium partnered to power the Ongo living-lamp robot with Gradium’s voices, speech-to-text, and text-to-speech technology. The companies said the initial integration would lead to deeper work on natural, expressive human-robot interaction.
Acolad and Gradium announced a partnership combining Acolad’s language and AI-enabled workflows with Gradium’s real-time streaming text-to-speech and speech-to-text models. The collaboration targets secure, scalable, multilingual AI-powered interpreting, while also building on Acolad’s data-services relationship with Gradium.
Gradium launched out of stealth as a Kyutai spinout with a $70 million seed round led by FirstMark Capital and Eurazeo. The company introduced multilingual real-time voice models aimed at faster, more accurate speech experiences for developers.
Gradium officially launched with real-time multilingual transcription and speech synthesis in English, French, German, Spanish, and Portuguese. Its APIs were positioned for developers and enterprises building natural voice agents and other production voice applications.
Active Roles
8Business Model
Gradium monetizes its voice AI through a free tier, recurring monthly or yearly plans, and pay-as-you-go credits for text-to-speech, speech-to-text, and voice-cloning services. Its developer/API access appears to be the primary commercial product.
Products
Customers
Tech Stack
Similar Companies
Competitors
Key Investors
Eurazeo, FirstMark, DST Global, Xavier Niel, Eric Schmidt