Companies

Baseten

baseten.co

Baseten provides infrastructure for training, deploying, and serving production AI models efficiently at scale.

HQSan Francisco, California, United States
Employees51-200
Funding$585M
Valuation$13B
Revenue$50m ARR
86 active roles
Profile 6mo agoJobs checked 16h ago
AI / MLAI InfrastructureB2B SaaSSeries E+$200M-$1B

About

Baseten builds a training and inference platform that turns open-source, fine-tuned, or custom AI models into production API endpoints with autoscaling, observability, multi-cloud GPU scheduling, and optimized serving. It sells primarily to engineering and machine-learning teams at startups and AI companies, differentiating through applied performance research, distributed infrastructure, and developer tooling focused on high performance and cost efficiency.

Market

Baseten competes in the AI infrastructure and MLOps market with an inference-first cloud platform for deploying, serving, scaling, and optimizing open, custom, and proprietary models in production. It positions around high-performance, reliable, and cost-efficient LLM inference, differentiating through GPU-accelerated infrastructure and software optimization—including NVIDIA and TensorRT-LLM integrations—alongside Kubernetes-based operations and multi-cloud model serving.

Target Customers

Ideal customers are AI-native application companies and enterprises that run open-source, custom, or proprietary AI models—especially LLMs—in production at scale. Core users and buying centers include data science and machine-learning teams, engineering, research, infrastructure, product, security, and enterprise operations.

At a Glance

Problem

Bringing AI models from experimentation into reliable production is difficult because teams must solve deployment, scaling, hardware, latency, and cost problems at once. Baseten addresses this infrastructure gap for companies building AI products, especially those serving open-source, fine-tuned, or custom models where inference performance and unit economics directly affect product quality and gross margins. The central use case is turning a model into a dependable, high-throughput production API without forcing an application team to build and operate its own GPU-serving stack.

The pain is especially acute for variable or rapidly growing workloads: idle capacity wastes money, while traffic spikes can create latency and reliability problems. Baseten’s positioning emphasizes high performance and cost efficiency, and its infrastructure is designed for workloads that need to scale across clouds and regions rather than remain a single-model experiment.

Product / Service

Baseten is a managed AI training and inference platform. A customer brings an open-source model from Hugging Face, a fine-tuned checkpoint, or a custom model; Baseten packages it as a production API endpoint with autoscaling, observability, and optimized serving infrastructure. The platform handles containerization, GPU scheduling across multiple clouds, and model-engine optimizations such as TensorRT-LLM compilation, allowing developers to focus on the model and application rather than deployment operations.

The delivery model combines software, infrastructure, and applied performance expertise. Deployments automatically adjust replicas to traffic, can scale to zero when idle, and can scale back up when demand arrives; Baseten also schedules workloads across cloud providers and regions. The intended benefit is faster time to market with lower operational burden, better latency and throughput, and less spending on unused GPU capacity.

Market

Baseten competes in AI infrastructure, more specifically managed model deployment, inference serving, and GPU-backed AI application infrastructure. Its adjacent and direct competitors include Fireworks AI, Together AI, and Runpod, which also help developers run models at production scale; the competitive axis is inference latency, throughput, performance, availability, and cost. Baseten’s positioning also overlaps with broader cloud AI platforms and infrastructure providers, but it differentiates around specialized serving optimization and multi-cloud execution.

The company is not presented as pre-revenue in the available research. Greylock reports that the platform processes more than one billion inference calls per day across 18 cloud providers, indicating substantial production usage. Baseten’s own announcement says it raised a $1.5 billion Series F at a $13 billion valuation, providing evidence of significant investor traction, although the available sources do not disclose revenue, profitability, or customer-level financial metrics.

Founders & Leadership

Tuhin SrivastavaFounder
CEO and Co-Founder
Amir HaghighatFounder
CTO and Co-Founder
Philip HowesFounder
Chief Scientist and Co-Founder
Pankaj GuptaFounder
Co-Founder
Dannie HerzbergPresident

Funding History

2022-04
Seed$8M

Greylock, South Park Commons Fund

2022-04
Series A$12M

Greylock Partners

2024-03
Series B$40M

IVP, Spark Capital

2025-02
Series C$75M

IVP, Spark Capital

2025-09
Series D$150M

BOND

2026-01
Series E$300M

IVP, CapitalG

2026-06
Series F$1.5B

Altimeter Capital, Conviction Partners, Spark Capital, Sands Capital, Wellington Management

Recent News

2026-07-23
How to choose an AI model: lessons from Notion and Gamma

Baseten shared model-selection lessons from Notion and Gamma, emphasizing workflow-specific model choice, model switching for cost and reliability, and the growing viability of open-weight models.

2026-06-22funding
Baseten Raises $1.5 Billion to Power the Next Era of AI Inference

Baseten announced a $1.5 billion Series F at a $13 billion valuation, led by Altimeter Capital, Conviction Partners, and Spark Capital. The company said revenue had grown 20x and inference volume 40x over the prior year.

2026-06-11partnership
Mercury 2, the first reasoning diffusion LLM, is now on Baseten

Baseten and Inception announced that Mercury 2 is live on Baseten, making the platform the first inference provider to offer production-grade diffusion large language models to developers.

2026-06-02partnership
MAI-Thinking-1 is coming to Baseten

Baseten and Microsoft AI announced that Microsoft AI’s MAI-Thinking-1 reasoning model would be available through Baseten, offering a commercial-grade model with post-training customization options.

2025-12-10
Baseten Acquires Parsed to Enable Companies to Own Their Intelligence

Baseten announced its acquisition of Parsed, a reinforcement-learning startup focused on post-training and continual learning for large models. Parsed’s team and technology were to be integrated into a single platform spanning inference, data, evaluation, and post-training.

Active Roles

86
San Francisco/Operations/Today
San Francisco/Engineering/1d ago
San Francisco/Data & Analytics/1d ago
San Francisco/Marketing/1d ago
San Francisco/Operations/2d ago
San Francisco/Finance/2d ago
San Francisco/Operations/7d ago
San Francisco/Design/8d ago
San Francisco/Design/8d ago
San Francisco/Design/8d ago
San Francisco/Sales/10d ago
New York/Sales Engineer/10d ago
San Francisco/Operations/14d ago
San Francisco/Marketing/15d ago
San Francisco/Engineering/15d ago
San Francisco/Forward-Deployed Engineer/16d ago
San Francisco/Finance/17d ago
San Francisco/Operations/17d ago

Business Model

Baseten primarily monetizes production inference and model API usage through usage-based pricing: customers pay as they go, including per-token charges for Model APIs, rather than a standard platform fee. Enterprise customers are handled under a separate Enterprise plan.

Products

TrussModel APIsDedicated deploymentsInference optimizationMulti-cloud model serving

Customers

AbridgeClayCursorDecagonHarveyHubSpotLovableMercorNotionOpenEvidence

Tech Stack

KubernetesAmazon EKSAmazon EC2NVIDIA GPUsTensorRT-LLMKarpenterLarge language models (LLMs)MLOps

Competitors

OpenRouter
Modal
BentoML
Replicate
fal
Fireworks

Key Investors

IVP, CapitalG, Nvidia, Altimeter, Battery Ventures, Bond Capital, Conviction, 01A, Greylock, Spark Capital, BoxGroup, Premji Invest