Companies

Miso Labs

misolabs.ai

Miso Labs builds expressive, low-latency foundation voice models and APIs for next-generation voice agents.

HQSan Francisco, California, United States
Employees1-50
1 active role
Jobs checked 15h ago
AI / MLFoundation Model Provider

About

Miso Labs develops emotive foundation voice models and API access for builders and teams deploying voice agents. Its Miso One model emphasizes 110ms real-time latency, ten-second one-shot voice cloning, open-source models, and on-premises deployment for expressive applications and sensitive enterprise workloads.

Market

Miso Labs competes in the foundation voice-model and real-time voice-agent market, positioning itself around highly emotive, human-like speech rather than conventional functional text-to-speech. Its stated differentiators are 110 ms response latency, ten-second one-shot voice cloning, open-source weights, and local/on-premises deployment, compared with the higher latency cited for ElevenLabs and Sesame and the privacy constraints of closed-source alternatives.

Target Customers

Miso Labs targets product and engineering teams at companies building real-time voice agents, especially enterprise teams handling sensitive voice or customer data. Its positioning is particularly relevant to organizations that need expressive, low-latency speech plus local or on-premises deployment and support.

At a Glance

Problem

Miso Labs addresses a core weakness of AI voice agents: latency and insufficiently human expression make live conversations feel awkward and untrustworthy. The company says most agents lag by 700 milliseconds or more, creating pauses that disrupt conversational flow. Its main use case is therefore real-time conversational voice agents—automated callers, assistants, and other systems that need to respond naturally enough for users to keep engaging. The economic pain is not quantified publicly, but poor turn-taking and robotic delivery directly reduce the usefulness, trust, and likely completion rates of automated conversations; sensitive voice data also creates a deployment and compliance concern for enterprises.

Product / Service

Miso’s solution is an expressive voice foundation model, branded as MisoTTS 8B and also introduced as Miso One. It generates speech from text plus optional audio context, allowing it to model conversational tone and perform one-shot voice cloning from a short sample. The company advertises approximately 110-millisecond hosted response latency, compared with 700 milliseconds for ElevenLabs and 300 milliseconds for Sesame, alongside consistent voice identity and more natural emotional delivery.

The delivery model is hybrid: the weights are open source under a modified MIT license and can be run locally, while Miso offers on-premises hosting and enterprise support contracts on request. The public implementation is English-only and currently models individual, half-duplex turns rather than full conversational turn-taking; the company’s June 2026 announcement said API access was coming soon. The result is a privacy-oriented voice layer that can be self-hosted for sensitive applications, although the documented 110-millisecond figure refers to a hosted production API on H100-class hardware rather than ordinary local inference.

Market

Miso Labs competes in expressive text-to-speech, real-time voice AI, and foundation-model infrastructure for voice agents. Its own comparison positions ElevenLabs and Sesame as direct alternatives, while its differentiation is the combination of low latency, emotional expressiveness, rapid voice cloning, open weights, and on-premises deployment. This places it at the model layer beneath voice-agent applications rather than as a complete vertical customer-support or contact-center application.

The company appears to be at an early commercialization stage. Y Combinator lists it as an active Spring 2026 company, and its open-source GitHub repository had accumulated 3,133 stars and 320 forks in the retrieved record, providing meaningful developer interest. However, the public materials reviewed disclose no named customers, revenue, or contract wins; they describe API access as coming soon and identify the company as founded in 2025 with a two-person team. The best-supported classification is therefore pre-revenue or pre-scale based on public evidence, rather than an established revenue business.

Founders & Leadership

Aoden TeoFounder
Co-Founder and CEO
Cassidy DalvaFounder
Co-Founder and President

Funding History

2026-06
Pre-SeedUndisclosed

Y Combinator

Recent News

2026-06-11partnership
San Francisco-Based Miso Labs Goes Live

Designatives announced the launch of Miso Labs, a San Francisco AI lab focused on emotive voice foundation models. Designatives said it partnered with Miso from a blank page through branding, website design, and development, and that the new site was live at misolabs.ai.

2026-06-06
Miso Labs, the Rise of Emotive Voice AI

i-Scoop profiled Miso Labs’ voice AI approach, highlighting its Miso One eight-billion-parameter text-to-speech model, reported 110-millisecond latency, one-shot voice cloning, and open-source deployment.

2026-06-04product
Miso Labs Releases MisoTTS: An 8B Emotive Text-to-Speech Model with Open Weights

MarkTechPost covered MisoTTS, an open-weights eight-billion-parameter model that generates expressive speech from text and audio context. The report highlighted local deployment and the modified MIT license, while noting that the initial release was half-duplex and API access was not yet available.

2026-06-04product
Miso One: 110ms Real-Time TTS Voice Model Guide 2026

ExplainX described Miso-TTS v1 as Miso Labs’ open-source voice foundation model, emphasizing reported 110-millisecond real-time latency and one-shot voice cloning.

2026-06-03
AI Launch Tracker - Miso One: The 8B Open-Source Voice Model That Wants to Out-Emote Humans

Kingy AI reported on Miso One as an eight-billion-parameter open-source text-to-dialogue RVQ Transformer. The coverage said Miso Labs released the model weights on launch day and described optional audio-context conditioning and local deployment.

2026-06-03product
Releasing the MisoTTS

Miso Labs introduced MisoTTS, an eight-billion-parameter transformer that generates speech from text and audio context using hierarchical residual vector quantization. The model’s weights were released on Hugging Face under a modified MIT license, with API access announced as forthcoming.

Active Roles

1
San Francisco, CA, US/Data & Analytics/1d ago

Business Model

Miso Labs monetizes its voice API through monthly subscription plans that include usage minutes, with additional per-minute charges. It also offers custom-priced annual enterprise contracts featuring volume pricing, on-premises deployment, dedicated voice fine-tuning, and support.

Products

Miso One — an 8-billion-parameter, highly expressive text-to-speech model designed for low-latency voice agents and one-shot voice cloning.MisoTTS — an open-source 8-billion-parameter speech-and-dialogue model that conditions generation on both text and audio context, with API access announced as forthcoming.

Tech Stack

8B transformer-based speech modelsResidual Vector Quantization (RVQ) audio tokenizationInterleaved text-and-audio conditioning7.7B temporal backbone plus 300M depth decoderOpen-source model weights with Hugging Face distributionOn-premises/local deployment

Competitors

ElevenLabs
Sesame
Inworld AI
Cartesia
Hume
Deepgram