About
Markov builds datasets and environments for training computer-use AI, capturing real-world computer workflows and converting them into data models can learn from. Its apparent customers are teams developing computer-use agents and models; its differentiation is large-scale behavioral data spanning professional software and other interactive environments.
Market
Markov competes in the AI training-data and agent-environment market, with a specific focus on computer-use models that operate through graphical user interfaces. Its differentiation is a data-first, open-source approach centered on large volumes of screen-recording workflows across professional software and games, supplemented by action/event metadata; this contrasts with broader managed annotation providers such as Appen and with competitors such as Scale AI and Surge AI that emphasize simulated or expert-built reinforcement-learning environments and evaluations.
Markov’s ideal customers are AI model developers—especially frontier AI labs and AI-agent or robotics startups—whose research and engineering teams need data to train and evaluate computer-use models. The primary buyers are likely leaders responsible for model training, reinforcement learning, data acquisition, or agent evaluation.
At a Glance
Problem
Computer-use AI models need to operate graphical interfaces rather than merely generate text, but training them requires large volumes of authentic, task-specific demonstrations: screen states, mouse and keyboard actions, and annotations tied to real workflows. Markov addresses this data bottleneck by sourcing high-quality tasks and capturing how people use computers. The economic pain is the difficulty and cost of collecting enough reliable, domain-specific interaction data, especially for complex software and long-horizon workflows.
The killer use case is training agents that can perform white-collar software work through a GUI, including workflows in Salesforce, Blender, Photoshop, AutoCAD, Excel, and VS Code. Markov frames the opportunity as automating “the next level of white-collar work,” turning demonstrations of expert behavior into training data that helps models reliably click, type, scroll, and complete real tasks.
Product / Service
Markov offers data and environments for training computer-use AI. It captures how people use computers and converts that activity into model-ready assets, including screen recordings, synchronized mouse and keyboard inputs, annotations, and open datasets. Its public delivery model combines open-source releases through Hugging Face with a contact-led offering for specialized data; for example, the company solicits interest in expert CAD computer-use data.
The flagship Computer Use Large dataset contains 48,478 screen-recording videos totaling approximately 12,300 hours across professional software categories such as AutoCAD, Blender, Excel, Photoshop, Salesforce, and VS Code. The dataset is designed to train and evaluate agents that interact with desktop software through GUI actions. Markov also publishes a 494.7-hour gaming dataset spanning 776 workflows and 168 games, demonstrating the breadth of computer-interaction data it can assemble.
Market
Markov competes in the emerging AI training-data, data-labeling, and reinforcement-learning-environment market, with a narrower focus on computer-use agents. The company’s materials do not name direct competitors. Adjacent alternatives include Scale AI’s realistic RL environments and open research environments or benchmarks such as OSWorld, while broader human-data suppliers represent another substitute for sourcing and labeling demonstrations. These are best viewed as neighboring offerings rather than confirmed head-to-head competitors.
Public traction is primarily open-source adoption and dataset scale rather than disclosed commercial revenue. Markov’s site reports more than 150,000 Hugging Face downloads, while its public datasets show substantial usage and scale; Y Combinator lists the company as active, founded in 2026, in its Summer 2026 batch, with a two-person team and zero listed jobs. No public source reviewed disclosed revenue, named customers, pricing, or funding, so Markov is best characterized as an early-stage company that appears pre-revenue or commercially undisclosed.
Founders & Leadership
Funding History
Not publicly disclosed
Recent News
Y Combinator’s company profile describes Markov as an active San Francisco startup sourcing high-quality tasks and data to train computer-use AI models. The profile identifies Markov as a Summer 2026 batch company.
Markov published a browsable collection of computer-use workflow samples with synchronized video, narration annotations, and raw input events. The samples cover coding, design, browser use, gaming, video editing, and productivity workflows.
A post linking to Markov’s computer-use-large dataset announced an open-source computer-use recordings dataset with more than 10,000 hours across applications including Salesforce, Blender, and Photoshop.
Markov co-founder Dev Mandal announced a dataset of computer-use recordings intended to help build the next generation of computer-use agents, covering more than 300 tasks.
Markov announced what it described as an advanced computer-use dataset, built by capturing how people use computers and converting those interactions into data for model training.
Active Roles
0No active roles right now.
Get notified when they postBusiness Model
Markov appears to use a B2B data and licensing model, selling computer-use datasets, task data, and training environments to AI developers through negotiated engagements. Its public materials show data samples and a “Talk to us” sales CTA, but do not publish specific pricing.