Companies

Arthur

arthur.ai

Arthur provides model-agnostic software to evaluate, monitor, govern, and improve enterprise AI systems.

HQNew York City, New York, United States
Employees51-200
Funding$63M
2 active roles
Profile 6mo agoJobs checked 22h ago
AI / MLAI ApplicationB2B SaaSSeries A$50M-$200M

About

Arthur builds a continuous AI evaluation, monitoring, governance, and optimization platform for traditional machine learning, generative AI, and agentic systems. It sells to enterprise security, audit, governance, engineering, and product teams, differentiating through a model-agnostic control plane for discovering, monitoring, governing, and improving AI across its lifecycle.

Market

Arthur competes in the enterprise AI observability, evaluation, security, and governance market, covering traditional ML, GenAI, and agentic systems through a model-, framework-, and cloud-agnostic control plane. Its differentiation is breadth: it combines model and agent observability, continuous evaluation, LLM firewall and guardrail capabilities, policy enforcement, and agent discovery in one platform rather than focusing only on prompt-level evaluation or runtime defense. Its competitive set therefore overlaps with Arize AI, Fiddler AI, Galileo AI, LangSmith, and Weights & Biases across AI monitoring, evaluation, and governance.

Target Customers

Arthur targets large enterprises deploying AI in production, particularly organizations in financial services, insurance, healthcare, and other regulated or high-stakes industries. Its primary users and buyers include ML engineers, security teams, compliance and risk officers, audit and governance teams, and product or engineering leaders responsible for reliable enterprise AI.

At a Glance

Problem

Arthur addresses the enterprise trust and control gap around AI. Organizations are deploying models and agents, but the teams building them and the teams responsible for security, compliance, and governance often lack a shared operating system. The economics are consequential: Arthur states that only 25% of AI projects return their investment, while unreliable or poorly governed systems can suffer from regressions, bias, weak explainability, and customer distrust.

The clearest use case is monitoring and improving customer-facing AI agents before and after deployment. In Arthur’s Upsolve case study, the company needed to ensure that its Analysis AI Agent answered questions reliably, diagnose failures immediately, and provide measurable evidence of correctness so customers would adopt it with confidence.

Product / Service

Arthur provides a model-, framework-, and cloud-agnostic AI governance and performance platform. Its observability product covers LLM, tabular, computer-vision, and NLP models, helping teams track accuracy and data drift, investigate explainability, detect fairness and bias issues, and improve model performance. Its broader platform adds continuous evaluations across development and runtime, tracing, production experimentation, versioned prompts with rollback, alerts, and organization-wide policy controls.

The product is delivered through multi-tenant SaaS, with Premium and Enterprise options, and can also support single-tenant SaaS, customer VPCs, BYO cloud, or on-premises deployments. Arthur’s deployable runner can execute evaluations and guardrail checks inside the customer environment while sending only aggregated metrics and metadata to the control plane. The benefit is earlier detection of quality or safety regressions, better accountability, and a practical way for technical, product, security, and governance teams to operate AI at scale.

Market

Arthur competes in the overlapping markets for AI governance, AI observability, model monitoring, continuous evaluation, and MLOps. Its positioning has expanded from machine-learning monitoring into a broader enterprise control plane for traditional ML, generative AI, and agentic systems. Market profiles identify Fiddler Labs, Portkey, and Arize AI as competitors; adjacent MLOps platforms and monitoring tools also compete for parts of the same budget.

The evidence indicates meaningful commercial traction rather than an idea-stage or pre-revenue profile, although no current revenue figure is disclosed. Arthur says it has served customers including Humana, Zesty.ai, and Truebill since its 2018 inception, and its 2025 Upsolve case study demonstrates continued use in production-oriented agent development. The company announced a $42 million Series B in 2022, more than $60 million in total funding at that point, and reported 58% quarterly growth and 445% growth over the preceding year.

Founders & Leadership

Adam WenchelFounder
CEO
John DickersonFounder
Startup advisor and angel investor
Liz O'SullivanFounder
Current title not identified in the available evidence
Priscilla AlexanderFounder
AI product and technology executive
Zach FryEngineering leader
Eve StaszczyszynOperations leader

Funding History

2019-08
Seed$3.3M

Work-Bench, Index Ventures

2020-12
Series A$15M

Index Ventures

2022-09
Series B$42M

Acrew Capital, Greycroft Ventures

Recent News

2026-04-10
What "Building an Agent" Means & Why Most Get It Wrong

Arthur published an explainer on what practical agent development entails, highlighting Louisa, an open-source workflow built by Arthur that automatically generates release notes.

2026-04-06
AI Agent Guardrails: Pre-LLM & Post-LLM Best Practices

Arthur outlined best practices for implementing pre-LLM and post-LLM guardrails for AI agents, including PII redaction, hallucination detection, and self-correction loops.

2026-01-07partnership
Arthur Launches Agent Discovery and Governance Platform on Google Cloud Marketplace

Arthur made its agent discovery and governance platform available through Google Cloud Marketplace. The offering enables enterprises to discover and govern AI agents within their existing cloud environment.

2025-12-22
Arthur 2025 Recap: Building Trust & Governance for Agentic AI

Arthur recapped its 2025 expansion from model monitoring into agentic governance, including the ADG platform, open-source Evals Engine, and Agent Development Lifecycle methodology. The company also highlighted deeper partnerships across AWS, Google Cloud, developer tooling, and open standards.

2025-12-17product
Arthur Releases the First Enterprise Platform Built for Agentic Discovery and Governance

Arthur launched its Agent Discovery & Governance platform to discover agents across compute environments, build an agent inventory, and support governance. The platform provides a unified foundation for traditional, generative, and agentic AI systems.

2025-11-25partnership
Stronger Engines and a Look at What Is Coming in 2026

Arthur’s November platform update improved setup for its open-source Evals Engine and previewed agentic governance capabilities planned for early 2026. The update also highlighted momentum for the Arthur Start Up Partner Program, which offers platform credits, guidance, and support.

2025-11-03product
Introducing The Agent Development Lifecycle (ADLC)

Arthur introduced the Agent Development Lifecycle, a methodology for building reliable agentic AI systems and moving agent projects from functional pilots toward production readiness.

Active Roles

2
Main (Hybrid)/Forward-Deployed Engineer/34d ago
New York City, USA/Product/34d ago

Business Model

Arthur uses a freemium SaaS model with a free tier, a $60-per-month Premium tier, and Enterprise plans for larger deployments. Enterprise offerings support multi-tenant or single-tenant SaaS, self-managed VPC/BYOCloud/on-premises deployments, and professional services.

Products

Arthur Observability for monitoring, measuring, and improving traditional ML, GenAI, and agentic AI systemsArthur Shield, an LLM firewall for securing deployed LLM applicationsArthur Evals Engine, a free open-source toolkit for evaluating AI modelsArthur Bench, an open-source tool for evaluating LLMs for production use casesArthur Agent Discovery & Governance, for discovering, cataloging, monitoring, evaluating, and governing production AI agents

Customers

HumanaExpelTruebillUpsolveAxios HQ

Tech Stack

Large language models and generative AI, including Anthropic, OpenAI, Meta Llama, Google Gemini, and Together AI modelsTraditional machine-learning models across tabular, computer-vision, and NLP use casesAgentic AI systems, including agent tools, sub-agents, data sources, and orchestrationOpenTelemetry and OTLP/HTTP tracingPython and JavaScript SDKsAWS integrations, including Bedrock, SageMaker, EKS, IAM, and VPCGoogle Cloud, Vertex AI, Kubernetes, and custom model pipelinesModel-, framework-, and cloud-agnostic deployment architecture

Competitors

Arize AI
Fiddler AI
Galileo AI
LangSmith
Weights & Biases

Key Investors

Index Ventures, Work-Bench, Plexo Capital, Acrew Capital