Companies

Datoric

datoric.com

Datoric delivers licensed, ethically sourced multimodal data for frontier AI development.

HQSan Francisco, California, United States
Employees1-50
Data InfrastructureData Labeling / TrainingInfrastructure

About

Datoric builds licensed, consented multimodal training data—including voice, video, computer-use, robotics, and physical-AI datasets—for frontier AI teams developing models and intelligent systems. It differentiates itself through origin-level contributor consent, fair compensation, provenance records, clean licensing, and isolated collection workflows rather than open-web scraping.

Market

Datoric competes in the enterprise AI training-data, data-annotation, and custom multimodal dataset market, serving teams building frontier models and AI systems for computer use, voice, video, robotics, physical AI, and gaming. Its differentiation is a compliance- and provenance-first model: data is sourced from consented and fairly compensated contributors at the origin, delivered under clear licenses, and supported by build-to-spec collection, structured labeling, and QA rather than relying primarily on scraped or generic corpora.

Target Customers

Datoric primarily serves frontier-AI model developers and specialized AI teams working on robotics, world models, voice AI, computer-use agents, physical AI, and gaming AI. Its strongest buyer profile is a technically sophisticated, compliance- or security-conscious organization whose ML, data, procurement, and legal teams need custom or licensed training data that can withstand scrutiny.

At a Glance

Problem

Datoric addresses a central bottleneck in frontier multimodal AI: labs need authentic human-generated data, but open-web and conventional third-party sources often provide unclear rights, missing consent, weak provenance, and exposure to legal, reputational, fraud, and security risk. Quality and provenance also become harder to control as collection scales, while compromised training data can contribute to unsafe model behavior and expensive rework.

The killer use case is difficult, project-specific data that off-the-shelf corpora do not provide, especially for voice models, robots, world models, physical-AI systems, and computer-use agents. Examples include multilingual speech, egocentric video, action-conditioned interactions, and browser or GUI trajectories collected under tightly defined requirements.

Product / Service

Datoric is a licensed multimodal data company that combines a dataset catalog with build-to-spec collection programs. It sources contributors directly, establishes consent and licensing at the point of contribution, offers fair compensation, and links samples to verifiable provenance. Its private apps and isolated project environments are designed to keep customer programs separate and limit unauthorized or manipulated submissions.

The delivery model runs from specification through recruitment, collection, labeling, modality-specific quality review, piloting, and delivery with rights attached. Customers define the modalities, tasks, labels, environments, and acceptance criteria; contributors then record the required voice, video, computer-use, robotics, or physical-AI data with structured metadata. The result is data that is more tailored to the model objective and easier for legal, security, and model-quality teams to evaluate and defend.

Market

Datoric competes in the AI training-data, data-collection, annotation, and multimodal data-infrastructure market, with a specialization in licensed, provenance-first data for frontier models. Its broader competitors include established providers such as Appen, which offers collection and annotation across text, image, audio, video, and geospatial data, and Scale AI, whose Data Engine supplies labeled data and domain-expert workflows. Datoric's differentiation is its emphasis on origin-level consent, engagement-specific licensing, project isolation, and custom multimodal data that may not exist yet.

Datoric is not pre-revenue. In its July 29, 2026 Y Combinator launch, the company reported nearly seven figures of revenue in the preceding 30 days, a major quality improvement versus other providers, and a network of more than 300,000 active contributors. These figures are company-reported rather than independently audited, and the public evidence does not identify customer names or contract economics. The broader data-collection and labeling market was estimated at $3.8 billion in 2024 and projected to reach $17.1 billion by 2030, providing a large market backdrop for Datoric's specialized offering.

Founders & Leadership

Nikhil ReddyFounder
CEO
Jeffrey LinFounder
CTO

Funding History

2026-01
Accelerator/Incubator$500K

Y Combinator

Recent News

2026-07-29
Datoric | Trustworthy data for the next generation of models

Y Combinator featured Datoric's launch, describing its custom, security-first training data for voice models, robotics, and world models. The company said its private contributor systems had improved quality for customers and generated nearly seven figures in revenue during the preceding 30 days.

2026-06-10product
VidWork-Bench: A five-axis benchmark for procedural video understanding

Datoric Research published VidWork-Bench, a benchmark covering step recognition, temporal ordering, causal reasoning, cross-modal grounding, and error detection. It evaluates 171 procedural video clips through 2,092 QA items and 10,686 scored model responses across six vision-language models; the page also identifies a 2026-07-21 report edition.

Active Roles

0

No active roles right now.

Get notified when they post

Business Model

Datoric sells catalog datasets and custom data-collection programs under direct, engagement-specific commercial licenses. Customers pay for licensed multimodal datasets or build-to-spec collection engagements, with delivery accompanied by provenance, consent, and usage-rights documentation.

Products

Licensed off-the-shelf multimodal datasets: Computer-Use Traces, TTS Voice, Egocentric Residential Video, and Audio-Video ConversationalCustom data-collection programs for voice, video, computer-use, robotics, and physical AIComputer-use agent training-data programsPhysical AI and VLA training dataVoice AI training data for TTS, ASR, and voice agentsGaming AI training dataResearch benchmarks including VideoTruth-Bench, VidWork-Bench, GlobalVoice-Bench, and VoicePro-Bench

Tech Stack

Multimodal AI training-data and data-annotation platformComputer-use agent trajectory capture, including screen recordings, UI action logs, structured context, and outcome labelsSpeech and synchronized audio-video data for TTS, ASR, and voice agentsEgocentric video and multimodal human-manipulation data for physical AI and VLA modelsStructured metadata, labeling, provenance, consent, licensing, and modality-specific QA workflowsAction-conditioned gameplay data for gaming AI, world models, and game-playing agents

Competitors

Nexdata
Appen
Defined.ai
Shaip
Scale AI