Companies

FriendliAI

friendli.ai

FriendliAI provides optimized cloud infrastructure for fast, reliable, cost-efficient large-scale AI model inference.

HQSan Francisco, California, United States
Employees11-50
Funding$26m
20 active roles
Profile 6mo agoJobs checked 15m ago
AI / MLAI InfrastructureB2B SaaSSeed$10M-$50M

About

FriendliAI builds an AI inference cloud and infrastructure platform for deploying, scaling, and monitoring large language and multimodal models. It sells to enterprise teams and AI organizations, differentiating through optimized inference, high throughput, low latency, GPU-cost efficiency, and enterprise-grade reliability including 99.99% uptime SLAs.

Market

FriendliAI competes in the generative-AI infrastructure and managed model-inference market, offering OpenAI-compatible model APIs plus serverless shared-GPU and dedicated reserved-GPU endpoints for production deployment. It positions itself around vertically optimized, high-throughput and low-latency inference, claiming 2×+ faster performance, 50–90% GPU cost savings, and 99.99% uptime. Its differentiation is the combination of specialized inference-engine optimizations, multimodal support, enterprise reliability, and a path from elastic serverless access to isolated dedicated capacity.

Target Customers

FriendliAI primarily targets enterprise teams, AI organizations, and high-volume production AI teams that need to deploy and scale foundation models reliably. Its buyer personas are likely ML/AI engineering, platform, and infrastructure leaders in industries such as telecommunications, healthcare, finance, and e-commerce, with both elastic serverless and isolated dedicated-GPU deployment needs.

At a Glance

Problem

Running generative AI in production creates an inference “last-mile” problem: teams need fast, reliable responses at high and unpredictable volumes while controlling scarce-GPU costs and the operational burden of deploying, scaling, and recovering large models. The economics become painful for high-traffic applications: FriendliAI cites chatbot GPU-cost reductions of 50–90%, while ScatterLab reported GPU costs reaching up to 70% of operating costs and NextDay AI processed more than three trillion tokens monthly, creating substantial H100 expense.

The clearest killer use case is high-volume conversational AI, including character chatbots, real-time assistants, and increasingly agentic applications that require low latency, streaming, tool use, and dependable uptime. These workloads can overwhelm an in-house serving stack because every additional token and peak-traffic request translates into GPU capacity and infrastructure expense.

Product / Service

FriendliAI provides a managed “Frontier AI Inference Cloud” for deploying and serving open-weight, multimodal, and custom models. Its stack combines custom GPU kernels, smart caching, continuous batching, speculative decoding, parallel inference, and multi-cloud scaling; the company says this produces 2×-plus faster inference and higher GPU utilization. Customers can one-click deploy hundreds of thousands of Hugging Face models without handling manual optimization, deployment, scaling, or performance tuning.

The delivery model covers both elastic and dedicated workloads. Serverless Endpoints use shared GPUs for experimentation and low, bursty, or unpredictable traffic, while Dedicated Endpoints reserve GPU capacity for consistent performance and high-volume production workloads; FriendliAI also offers Friendli Container for customer-controlled infrastructure and supports custom or fine-tuned models. The benefit is a managed path to lower latency, higher throughput, automated scaling and fault recovery, and lower infrastructure cost, with the company claiming 2–5× faster output-token speed and a 99.99% uptime SLA for production traffic.

Market

FriendliAI competes in managed AI inference, generative-AI infrastructure, and GPU model-serving. The category includes API-first and serverless providers such as Together AI, whose service emphasizes running open-source models without infrastructure management, and Fireworks AI, which offers pay-per-token serverless inference. FriendliAI’s differentiation is its emphasis on production-grade latency-throughput trade-offs, deep serving optimizations, dedicated GPU deployments, and cost efficiency for demanding enterprise and agent workloads rather than simply providing model access.

The company shows operating traction rather than evidence of an unlaunched, pre-revenue project. In August 2025 it announced a $20 million seed-extension round led by Capstone Partners, with Sierra Ventures, Alumni Ventures, KDB, and KB Securities participating. Its cited customers include LG AI Research, SK Telecom, NextDay AI, ScatterLab, TUNiB, and Upstage; reported outcomes include SK Telecom’s fivefold throughput increase and threefold cost savings, LG’s tripling of traffic after deploying large EXAONE models, and NextDay AI’s roughly 50% GPU-cost reduction. The available sources do not disclose revenue or profitability, so the strongest conclusion is customer and funding traction with financial scale not publicly quantified in the reviewed material.

Founders & Leadership

Byung-Gon ChunFounder
Founder and CEO
Gyeong-In YuCTO
Brian YooChief Business Officer

Funding History

2021-XX (month not disclosed)
Seed$6M

Capstone Partners

2025-08
Seed extension$20M

Capstone Partners

Recent News

2026-07-01partnership
How Kilo Code and FriendliAI Bring Open Source AI Coding Agents to Production

FriendliAI and Kilo Code announced a partnership aimed at making production-grade AI coding agents faster, more accurate, and more cost-efficient.

2026-06-17product
GLM-5.2, Day-0 on FriendliAI Model APIs

FriendliAI made Z.ai’s GLM-5.2 flagship open-weight model available on its Model APIs on day zero, targeting agentic coding workloads.

2026-04-15product
FriendliAI Now Supports Anthropic Messages API

FriendliAI announced support for the Anthropic Messages API, allowing developers to use open models such as DeepSeek without rewriting their existing code.

2026-04-14partnership
FriendliAI and Samsung Cloud Platform Forge Strategic Alliance to Power Frontier Model AI Inference on NVIDIA B300 GPUs

FriendliAI and Samsung Cloud Platform formed an alliance combining FriendliAI’s inference stack with Samsung Cloud Platform’s scalable NVIDIA B300 GPU infrastructure.

2026-03-15partnership
Integrating FriendliAI with OpenClaw

FriendliAI published an integration guide for using its models with OpenClaw, including a multi-agent setup supporting multiple models and fallback behavior.

2026-02-11product
GLM-5 Now Available on FriendliAI with Day 0 Support

FriendliAI announced day-zero support for GLM-5 for reasoning, agentic workflows, and coding, available through Serverless and Dedicated Endpoints.

2025-12-15partnership
Enabling the Next Level of Efficient Agentic AI: FriendliAI Supports NVIDIA Nemotron 3 Nano Launch

FriendliAI served as an official launch partner for NVIDIA’s Nemotron 3 Nano, providing high-performance inference for efficient agentic AI applications.

2025-11-19partnership
FriendliAI Partners with Nebius to Deliver High-Performance, Cost-Efficient AI Inference

FriendliAI announced a strategic partnership with Nebius, integrating its inference-optimization technology into Nebius’ large-scale AI cloud infrastructure.

2025-08-29funding
FriendliAI raises $20M in funding to accelerate AI inference workloads

SiliconANGLE reported that FriendliAI raised $20 million in a seed extension led by Capstone Partners, with participation from Sierra Ventures, Alumni Ventures, KDB, and KB Securities.

2025-08-28funding
FriendliAI Secures $20M to Redefine AI Inference

FriendliAI announced a $20 million seed extension led by Capstone Partners to scale its AI inference platform, reduce infrastructure costs, and accelerate generative AI deployment.

Active Roles

20
San Francisco/Product/Today
Seoul/Engineering/17d ago
San Francisco/Engineering/21d ago
San Francisco/Engineering/33d ago
San Francisco/Engineering/33d ago
Seoul/Engineering/33d ago
San Francisco/Engineering/33d ago
San Francisco/Engineering/33d ago
Seoul/Customer Success Manager/33d ago
San Francisco/Sales/33d ago
San Francisco/Solutions Engineer/172d ago
San Francisco/Solutions Engineer/172d ago
Seoul/Engineering/172d ago
San Francisco/Engineering/172d ago
San Francisco/Engineering/172d ago
Seoul/Engineering/172d ago
Seoul/Product/172d ago

Business Model

FriendliAI monetizes inference through usage-based model APIs priced by input and output tokens, as well as dedicated GPU endpoints priced hourly. It also offers customizable enterprise plans and sales-assisted pricing for larger deployments.

Products

Friendli Inference serving engineFriendli Model APIsFriendli Serverless EndpointsFriendli Dedicated EndpointsFriendli Container

Customers

LG AI ResearchSK TelecomNextDay AIUpstage

Tech Stack

Large language model (LLM) servingMultimodal AI APIs for text, images, audio, and videoCustom GPU kernelsSmart cachingContinuous batchingGPU-accelerated inference, including NVIDIA H100 supportQuantized and mixture-of-experts (MoE) model supportOpenAI-compatible APIs

Competitors

Together AI
Fireworks AI
Baseten
Modal

Key Investors

Alumni Ventures, Sierra Ventures, Capstone Partners, KB Investment, Capstone Partners Korea