About
Miso Labs develops emotive foundation voice models and API access for builders and teams deploying voice agents. Its Miso One model emphasizes 110ms real-time latency, ten-second one-shot voice cloning, open-source models, and on-premises deployment for expressive applications and sensitive enterprise workloads.
Market
Miso Labs competes in the foundation voice-model and real-time voice-agent market, positioning itself around highly emotive, human-like speech rather than conventional functional text-to-speech. Its stated differentiators are 110 ms response latency, ten-second one-shot voice cloning, open-source weights, and local/on-premises deployment, compared with the higher latency cited for ElevenLabs and Sesame and the privacy constraints of closed-source alternatives.
Miso Labs targets product and engineering teams at companies building real-time voice agents, especially enterprise teams handling sensitive voice or customer data. Its positioning is particularly relevant to organizations that need expressive, low-latency speech plus local or on-premises deployment and support.
At a Glance
Problem
Miso Labs addresses a core weakness of AI voice agents: latency and insufficiently human expression make live conversations feel awkward and untrustworthy. The company says most agents lag by 700 milliseconds or more, creating pauses that disrupt conversational flow. Its main use case is therefore real-time conversational voice agents—automated callers, assistants, and other systems that need to respond naturally enough for users to keep engaging. The economic pain is not quantified publicly, but poor turn-taking and robotic delivery directly reduce the usefulness, trust, and likely completion rates of automated conversations; sensitive voice data also creates a deployment and compliance concern for enterprises.
Product / Service
Miso’s solution is an expressive voice foundation model, branded as MisoTTS 8B and also introduced as Miso One. It generates speech from text plus optional audio context, allowing it to model conversational tone and perform one-shot voice cloning from a short sample. The company advertises approximately 110-millisecond hosted response latency, compared with 700 milliseconds for ElevenLabs and 300 milliseconds for Sesame, alongside consistent voice identity and more natural emotional delivery.
The delivery model is hybrid: the weights are open source under a modified MIT license and can be run locally, while Miso offers on-premises hosting and enterprise support contracts on request. The public implementation is English-only and currently models individual, half-duplex turns rather than full conversational turn-taking; the company’s June 2026 announcement said API access was coming soon. The result is a privacy-oriented voice layer that can be self-hosted for sensitive applications, although the documented 110-millisecond figure refers to a hosted production API on H100-class hardware rather than ordinary local inference.
Market
Miso Labs competes in expressive text-to-speech, real-time voice AI, and foundation-model infrastructure for voice agents. Its own comparison positions ElevenLabs and Sesame as direct alternatives, while its differentiation is the combination of low latency, emotional expressiveness, rapid voice cloning, open weights, and on-premises deployment. This places it at the model layer beneath voice-agent applications rather than as a complete vertical customer-support or contact-center application.
The company appears to be at an early commercialization stage. Y Combinator lists it as an active Spring 2026 company, and its open-source GitHub repository had accumulated 3,133 stars and 320 forks in the retrieved record, providing meaningful developer interest. However, the public materials reviewed disclose no named customers, revenue, or contract wins; they describe API access as coming soon and identify the company as founded in 2025 with a two-person team. The best-supported classification is therefore pre-revenue or pre-scale based on public evidence, rather than an established revenue business.
Founders & Leadership
Funding History
Y Combinator
Recent News
Designatives announced the launch of Miso Labs, a San Francisco AI lab focused on emotive voice foundation models. Designatives said it partnered with Miso from a blank page through branding, website design, and development, and that the new site was live at misolabs.ai.
i-Scoop profiled Miso Labs’ voice AI approach, highlighting its Miso One eight-billion-parameter text-to-speech model, reported 110-millisecond latency, one-shot voice cloning, and open-source deployment.
MarkTechPost covered MisoTTS, an open-weights eight-billion-parameter model that generates expressive speech from text and audio context. The report highlighted local deployment and the modified MIT license, while noting that the initial release was half-duplex and API access was not yet available.
ExplainX described Miso-TTS v1 as Miso Labs’ open-source voice foundation model, emphasizing reported 110-millisecond real-time latency and one-shot voice cloning.
Kingy AI reported on Miso One as an eight-billion-parameter open-source text-to-dialogue RVQ Transformer. The coverage said Miso Labs released the model weights on launch day and described optional audio-context conditioning and local deployment.
Miso Labs introduced MisoTTS, an eight-billion-parameter transformer that generates speech from text and audio context using hierarchical residual vector quantization. The model’s weights were released on Hugging Face under a modified MIT license, with API access announced as forthcoming.
Active Roles
1Business Model
Miso Labs monetizes its voice API through monthly subscription plans that include usage minutes, with additional per-minute charges. It also offers custom-priced annual enterprise contracts featuring volume pricing, on-premises deployment, dedicated voice fine-tuning, and support.