Companies

Patronus AI

patronus.ai

Patronus AI builds Digital World Models and evaluation infrastructure to train, test, and improve AI agents.

HQSan Francisco, California, United States
Employees11-50
Funding$40.1M
Revenue$1m ARR
11 active roles
Profile 6mo agoJobs checked 21h ago
AI / MLAI ApplicationB2B SaaSPre-Seed / Seed$1M-$10M

About

Patronus AI builds AI evaluation, monitoring, and simulation infrastructure for enterprise development teams, AI companies, and financial-services organizations. Its differentiation is the use of Digital World Models and realistic simulated workflows to train, test, and improve long-horizon AI agents before deployment.

Market

Patronus AI competes in the enterprise LLM evaluation, AI quality and safety, observability, and agent-testing market, while expanding into simulation infrastructure for training and stress-testing AI agents. It differentiates through research-led proprietary evaluators such as Lynx and GLIDER, the Percival agent debugger, and Digital World Models that extend beyond the narrower evaluation or observability focus of many point solutions.

Target Customers

Patronus AI primarily targets enterprise AI engineering, ML platform, quality, and safety teams deploying LLM applications and multi-step agents at scale, including customer-service, financial-services, software, and product-application use cases. Its newer simulation products also target foundation-model developers, frontier AI labs, and research teams seeking realistic environments for agent training and evaluation.

At a Glance

Problem

Patronus AI addresses the reliability gap between impressive benchmark scores and AI systems that work correctly in production. Enterprises deploying large language models and agents face hallucinations, safety vulnerabilities, unpredictable behavior, and the engineering burden of testing these systems across realistic workflows. The economics are tied to the cost of production failures and to maintaining testing infrastructure; Patronus positions its usage-based API as a more accessible alternative to managing open-source evaluation models and infrastructure.

The killer use case is stress-testing an AI agent before deployment on complex, real-world tasks. Patronus creates simulated replicas of websites and internal systems so agents can be tested for shortcuts, hacks, and incomplete task execution rather than merely being judged on static benchmark questions.

Product / Service

Patronus provides an automated AI evaluation, security, and guardrails platform delivered through a self-serve API, web dashboard, and enterprise services. Development teams can score model performance, generate adversarial test cases, benchmark different LLMs, monitor evaluation results, compare experiments over time, and run targeted tests using curated datasets such as FinanceBench, EnterprisePII, and SimpleSafetyTests.

Its evaluation models include Lynx for hallucination detection, configurable LLM judges for custom capability, safety, and alignment criteria, and real-time or offline evaluators. The newer agent-testing offering uses “digital world models” to reproduce websites and internal systems, allowing companies and model makers to evaluate agent behavior in realistic environments and improve reliability without relying entirely on human testing.

Market

Patronus competes in the automated AI evaluation, LLM security, guardrails, and AI-agent testing market, adjacent to the broader AI observability category. Ragas is cited as a comparable evaluation solution, while Galileo and other evaluation-first or observability platforms represent adjacent alternatives. Patronus’s differentiation is its combination of automated evaluators, production guardrails, benchmarking and monitoring workflows, and simulated environments for testing agents on real-world tasks.

The company is not pre-revenue. Patronus reported customers including AngelList, Pearson, and HP, and said its partner ecosystem included NVIDIA, MongoDB, and IBM. By June 2026, TechCrunch reported that virtually every frontier AI lab and many emerging startups were customers, revenue had grown fifteenfold year over year, and the company had raised a $50 million Series B, bringing total funding to approximately $70 million.

Founders & Leadership

Anand KannappanFounder
Co-Founder & CEO
Rebecca QianFounder
Co-Founder & CTO
Duncan CurtisSVP Product & Operations

Funding History

2023-09
Seed$3M

Lightspeed Venture Partners

2024-05
Series A$17M

Notable Capital

2026-06
Series B$50M

Greenfield Partners

Recent News

2026-07-18funding
Announcing our $50M Series B to Simulate the Entire World’s Intelligence and Unveiling our First Digital World Model for AI Agent Training and Simulation

Patronus AI announced a $50 million Series B and unveiled its first Digital World Model for AI-agent training and simulation. The announcement marks the company’s expansion from LLM evaluation into simulation infrastructure for agentic systems.

2026-06-25
Patronus AI lands $50M to build 'digital worlds' that stress-test AI agents

TechCrunch reported that Patronus AI is building Digital World Models that replicate websites and internal systems, allowing AI agents to be stress-tested in realistic environments.

2026-06-25funding
Patronus AI Raises $50 Million Series B and Unveils First Digital World Models for AI Agent Training and Simulation

Patronus AI announced a $50 million Series B led by Greenfield Partners and introduced Digital World Models designed to support AI-agent training and simulation.

2026-05-22product
Patronus AI Launches Self-Serve API for AI Evaluation and Guardrails

Patronus AI launched a self-serve API giving developers access to evaluation models trained by its research team, including its Lynx hallucination-detection model.

2025-10-19product
Introducing MEMTRACK: A Benchmark for Agent Memory

Patronus AI introduced MEMTRACK, a benchmark for studying long-term and cross-platform memory in agentic systems. The benchmark simulates a software-development environment to evaluate agent memory capabilities.

2025-08-20product
Patronus Evaluators

Patronus AI highlighted its Evaluators as a core platform capability for automatically assessing models across specific quality and performance dimensions.

Active Roles

11
San Francisco, CA/Engineering/39d ago
San Francisco, CA/Data & Analytics/89d ago
San Francisco, CA/Marketing/89d ago
San Francisco, CA/Operations/183d ago
San Francisco, CA/Engineering/190d ago
San Francisco, CA/Engineering/190d ago
San Francisco, CA/Engineering/196d ago
San Francisco, CA/Forward-Deployed Engineer/197d ago
San Francisco, CA/Engineering/197d ago
San Francisco, CA/Engineering/197d ago
San Francisco, CA/HR & Recruiting/197d ago

Business Model

Patronus AI monetizes its evaluation and guardrails platform through a self-serve, usage-based API, including pay-as-you-go charges per evaluator call and evaluation explanation, with free introductory credits. It also targets enterprise customers through its broader platform, although enterprise contract terms are not publicly disclosed.

Products

Core Evaluation and Monitoring Platform for experiments, comparisons, logging, traces, benchmarking, and LLM-as-a-judge scoringLynx hallucination detector and GLIDER general-purpose 3B evaluatorPercival evaluation copilot and agent debugger for detecting 20+ agentic failure modesDigital World Models for realistic digital workflowsGenerative Simulators and reinforcement-learning environments for adaptive agent trainingEvaluation benchmarks and datasets, including FinanceBench and MemTrack

Customers

Nova AIEmergence AIWeaviateEtsyGammaExaHospitable.comAlgomoDatabricks

Tech Stack

Large language models (LLMs) and LLM-as-a-judge evaluationProprietary evaluator models: Lynx and GLIDER, a 3B-parameter modelMultimodal AI evaluation for image-to-text systemsAgentic trace logging, monitoring, and integrations with LangChain, HuggingFace, CrewAI, and OpenAIRAG evaluation, adversarial test generation, and benchmark datasetsDigital World Models, reinforcement-learning environments, and generative simulators

Competitors

Braintrust
Galileo
Arize AI
Langfuse
Portkey
LangChain

Key Investors

Notable Capital, Lightspeed Venture Partners, Datadog, Factorial Capital