About
Olam Labs builds multi-agent simulated environments and behavioral evaluations for frontier AI labs and researchers. Its platform uses games involving humans and AI agents to measure social behavior, safety, negotiation, collaboration, and other capabilities, while preserving detailed behavioral traces that conventional benchmarks often miss.
Market
Olam Labs competes in the AI model evaluation, AI safety, alignment, and agent-testing market, with a specialized focus on measuring social intelligence and qualitative behavior before deployment. Its differentiation is the use of multi-agent games and long-horizon simulated societies—often involving humans, multiple agents, and resource constraints—to expose emergent social behaviors and produce legible behavioral data, rather than relying only on conventional task or prompt evaluations.
Olam Labs primarily serves frontier AI labs and AI researchers that need pre-deployment evaluations of model safety, alignment, social behavior, and performance. Likely buyers are model-evaluation, AI-safety, alignment, and research teams rather than general-purpose enterprise software departments.
At a Glance
Problem
Olam Labs addresses a blind spot in conventional AI evaluation: most testing emphasizes single-agent task performance, scientific benchmarks, or real-world data, while offering little visibility into how models behave in complex social systems. As AI agents become more autonomous and interact with people and other agents, labs need to know whether they can negotiate, collaborate, socialize, make sound judgments, or deceive before deployment. The core pain is therefore safety and deployment risk: a model may score well on isolated tasks yet behave unpredictably when incentives, communication, competition, and other agents interact.
The economic value proposition is scalable risk reduction and better test coverage rather than a stated line-item cost saving. Olam Labs argues that real-world social evaluations are difficult to quantify at scale, whereas simulated environments can produce large samples with legible game states, actions, reasoning traces, and causal context. Its killer use case is a pre-deployment behavioral evaluation that exposes socially consequential model tendencies—especially deception, negotiation, and cooperation—before those tendencies appear in autonomous products or institutions.
Product / Service
The flagship product is Social Arena, a public and research-oriented multi-agent environment in which humans play games such as Poker, Risk, and Catan against talking AI agents. The visible game is an engaging human interface; behind it, Olam runs model- and provider-agnostic multi-agent simulations that collect behavior over many matches and long sequences of interaction. Olam also runs internal environments such as simulated towns and ultra-long-horizon games, producing agent and human traces, datasets, evaluations, and behavior reports for research labs.
Its evaluation pipeline turns those interactions into performance and behavioral measures. For example, the Deception Index uses transcripts, reasoning traces, table talk, and known cards to identify deliberate misrepresentation, with human-curated rubrics, judge models, quality checks, and repeated benchmark runs. The result is a leaderboard and evaluation layer that lets researchers compare models on both competitive performance and social traits, while the public arena supplies human-informed interaction data.
Market
Olam Labs competes in the emerging AI model-evaluation, agent-evaluation, AI safety, and alignment market, with a distinctive focus on social intelligence in multi-agent environments. It is not a like-for-like replacement for every evaluation platform: adjacent alternatives include METR’s autonomous-capability evaluations, the UK AI Security Institute and Meridian Labs’ Inspect framework for frontier evaluations, Scale’s SEAL model leaderboards, and Braintrust’s production tracing and evaluation platform. Olam’s differentiation is the combination of human-versus-agent social games, continuous multi-agent interaction, and behavior metrics such as deception rather than only task accuracy or production monitoring.
The company is early but has meaningful product and research traction. It was founded in 2026, is an active YC Summer 2026 company with a two-person team, and publicly released Social Arena and its first benchmark, the Deception Index, on August 1, 2026. The benchmark had been built from roughly 55,000 human-involved poker hands, while subsequent evaluation pages report 69,140 hands; Olam says it works with frontier researchers and labs on pre-deployment evaluations and behavior reports. The results are explicitly labeled preliminary, and no public revenue figure or named customer list is disclosed, so the best-supported characterization is an early commercial/research-stage company that is revenue-undisclosed and may still be pre-revenue rather than a scaled vendor.
Founders & Leadership
Funding History
Y Combinator
Recent News
ExploreYC profiled Olam Labs as a Y Combinator Summer 2026 company evaluating model behavior and performance through multi-agent simulated games. The profile highlights Social Arena, where humans play games such as Catan, Risk, or Poker against AI opponents.
Olam Labs announced Social Arena, a platform where humans play social games including Risk, Catan, and Poker against AI agents. It also released the Deception Index, an AI-model benchmark based on approximately 55,000 Poker hands involving at least one human participant.
Active Roles
0No active roles right now.
Get notified when they postBusiness Model
Olam Labs appears to operate a B2B SaaS and AI-data model, selling behavioral evaluations, multi-agent environments, and datasets to frontier AI labs and researchers. Public evidence indicates free-tier access alongside paid plans, although detailed pricing and enterprise terms are not disclosed.