About
SafetyKit builds AI agents that automate risk, fraud, compliance, content moderation, and trust-and-safety reviews for large technology, finance, marketplace, and payments companies. Its differentiation is replacing difficult, expensive manual review work with vertical AI trained on the founders’ experience building risk-review systems at Stripe and Airbnb.
Market
SafetyKit competes in enterprise trust-and-safety, fraud-prevention, and governance, risk, and compliance SaaS, positioning itself as vertical, agentic AI for marketplaces, payments companies, and other digital platforms. Its differentiation is replacing large manual review operations with AI agents that investigate content, merchants, transactions, and compliance issues while supporting workflows such as case management, network takedowns, and continuous monitoring. This creates broader workflow coverage than point fraud or moderation tools, although Sift, Fingerprint, and SEON overlap on fraud while ActiveFence and Spectrum Labs overlap on content moderation and trust-and-safety.
SafetyKit primarily serves large digital marketplaces, payments companies, and other platform businesses with substantial user or merchant volumes and complex fraud, trust-and-safety, and compliance workloads. Likely buyers are trust-and-safety, fraud/risk, compliance, and operations leaders seeking to scale investigations and reduce manual review; named customers include Patreon, Eventbrite, Upwork, Character.ai, Substack, and Faire.
At a Glance
Problem
SafetyKit addresses the operational problem of fraud, abuse, and policy violations on large online platforms. Bad actors continually change their tactics, forcing trust-and-safety teams into a reactive cycle: discover an attack, identify a fix, patch defenses, and then respond as attackers adapt. Manual review is expensive and inconsistent; SafetyKit cites typical BPO human accuracy of about 80%, while a single scammer or illegal product can materially damage a platform. A central use case is reviewing marketplace content and activity before harm occurs—for example, reviewing every Upwork job post before publication—or investigating scams, account takeovers, fake listings, and coordinated fraud networks at scale.
Product / Service
SafetyKit combines data infrastructure with purpose-built AI agents delivered through a platform integration. Its lightweight SDK and API ingest user actions, models map relationships among users, actions, interactions, and content, and behavior models evaluate activity in real time for fraud and abuse signals. Agents then investigate cases end-to-end, resolve routine cases automatically, and escalate only matters requiring human judgment, with explanations that reviewers can inspect and challenge.
The benefit is broader coverage, faster response, lower manual-review load, and more consistent decisions across text, images, listings, transactions, and other modalities. SafetyKit says its agents can evaluate all platform activity, adapt to emerging patterns, and automate complex workflows such as scam detection and region-specific policy compliance; OpenAI reports SafetyKit evaluations showing more than 95% accuracy while reviewing 100% of customer content.
Market
SafetyKit competes in the AI-enabled trust-and-safety, fraud detection, risk, and compliance-operations market for marketplaces, payment platforms, fintechs, and other user-generated-content businesses. Its positioning is broader than a point fraud tool: it combines behavioral data infrastructure, multimodal content moderation, investigation workflows, and policy/compliance agents. Named alternatives in the fraud-detection category include Sift, Fingerprint, and SEON, while SafetyKit also says its AI agents can replace or reduce reliance on outsourcing firms such as Genpact and Accenture for risk and safety work.
The company is commercially deployed rather than merely pre-launch. SafetyKit identifies Patreon, Eventbrite, Upwork, Character.ai, Substack, Faire, and other large marketplaces and payments companies as production users; it says Upwork measured 95% accuracy for SafetyKit versus 80% for humans. OpenAI reports that SafetyKit processes more than 16 billion tokens daily, up from 200 million six months earlier, and protects hundreds of millions of end users. SafetyKit has publicly announced $27 million in funding, but the available evidence does not disclose pricing, revenue, or profitability.
Founders & Leadership
Funding History
Y Combinator
Not publicly disclosed
Ribbit Capital
Recent News
SafetyKit introduced continuous merchant monitoring for portfolio risk management. Its AI agents investigate what merchants sell, who they are connected to, and what they may be concealing.
SafetyKit highlighted its real-time AI fraud-prevention capabilities for detecting scams, fake shops, and coordinated fraud networks. The offering includes merchant intelligence and agentic investigations that can scale human review.
SafetyKit launched or published an MCC-classification capability for merchant risk assessment, automatically identifying merchants in elevated-risk categories such as adult content, dating services, gambling, gaming, and cryptocurrency.
OpenAI profiled SafetyKit’s integration of GPT-5, GPT-4.1, deep research, and Computer Using Agent technology into multimodal risk and compliance workflows. SafetyKit reported 95%+ accuracy while reviewing all customer content, processing 16 billion tokens daily, and achieving more than 10-point gains on its hardest vision tasks.
Active Roles
9Business Model
SafetyKit monetizes through enterprise B2B SaaS contracts for deploying AI agents across risk, fraud, compliance, and trust-and-safety operations. Its customers include large technology, finance, marketplace, and payments companies; public per-seat or usage-based pricing was not identified.