Companies

Markov

markovstudios.com

Markov creates real-world datasets and environments that help AI agents learn to operate computer software.

HQSan Francisco, California, United States
Employees1-50
Jobs checked 19h ago
Data InfrastructureData Labeling / Training

About

Markov builds datasets and environments for training computer-use AI, capturing real-world computer workflows and converting them into data models can learn from. Its apparent customers are teams developing computer-use agents and models; its differentiation is large-scale behavioral data spanning professional software and other interactive environments.

Market

Markov competes in the AI training-data and agent-environment market, with a specific focus on computer-use models that operate through graphical user interfaces. Its differentiation is a data-first, open-source approach centered on large volumes of screen-recording workflows across professional software and games, supplemented by action/event metadata; this contrasts with broader managed annotation providers such as Appen and with competitors such as Scale AI and Surge AI that emphasize simulated or expert-built reinforcement-learning environments and evaluations.

Target Customers

Markov’s ideal customers are AI model developers—especially frontier AI labs and AI-agent or robotics startups—whose research and engineering teams need data to train and evaluate computer-use models. The primary buyers are likely leaders responsible for model training, reinforcement learning, data acquisition, or agent evaluation.

At a Glance

Problem

Computer-use AI models need to operate graphical interfaces rather than merely generate text, but training them requires large volumes of authentic, task-specific demonstrations: screen states, mouse and keyboard actions, and annotations tied to real workflows. Markov addresses this data bottleneck by sourcing high-quality tasks and capturing how people use computers. The economic pain is the difficulty and cost of collecting enough reliable, domain-specific interaction data, especially for complex software and long-horizon workflows.

The killer use case is training agents that can perform white-collar software work through a GUI, including workflows in Salesforce, Blender, Photoshop, AutoCAD, Excel, and VS Code. Markov frames the opportunity as automating “the next level of white-collar work,” turning demonstrations of expert behavior into training data that helps models reliably click, type, scroll, and complete real tasks.

Product / Service

Markov offers data and environments for training computer-use AI. It captures how people use computers and converts that activity into model-ready assets, including screen recordings, synchronized mouse and keyboard inputs, annotations, and open datasets. Its public delivery model combines open-source releases through Hugging Face with a contact-led offering for specialized data; for example, the company solicits interest in expert CAD computer-use data.

The flagship Computer Use Large dataset contains 48,478 screen-recording videos totaling approximately 12,300 hours across professional software categories such as AutoCAD, Blender, Excel, Photoshop, Salesforce, and VS Code. The dataset is designed to train and evaluate agents that interact with desktop software through GUI actions. Markov also publishes a 494.7-hour gaming dataset spanning 776 workflows and 168 games, demonstrating the breadth of computer-interaction data it can assemble.

Market

Markov competes in the emerging AI training-data, data-labeling, and reinforcement-learning-environment market, with a narrower focus on computer-use agents. The company’s materials do not name direct competitors. Adjacent alternatives include Scale AI’s realistic RL environments and open research environments or benchmarks such as OSWorld, while broader human-data suppliers represent another substitute for sourcing and labeling demonstrations. These are best viewed as neighboring offerings rather than confirmed head-to-head competitors.

Public traction is primarily open-source adoption and dataset scale rather than disclosed commercial revenue. Markov’s site reports more than 150,000 Hugging Face downloads, while its public datasets show substantial usage and scale; Y Combinator lists the company as active, founded in 2026, in its Summer 2026 batch, with a two-person team and zero listed jobs. No public source reviewed disclosed revenue, named customers, pricing, or funding, so Markov is best characterized as an early-stage company that appears pre-revenue or commercially undisclosed.

Founders & Leadership

Harish AshokFounder
Co-Founder
Dev MandalFounder
Co-Founder & CEO

Funding History

2026-08
No publicly disclosed funding round identified as of 2026-08Not publicly disclosed

Not publicly disclosed

Recent News

2026-06-23
Markov joins Y Combinator Summer 2026

Y Combinator’s company profile describes Markov as an active San Francisco startup sourcing high-quality tasks and data to train computer-use AI models. The profile identifies Markov as a Summer 2026 batch company.

2026-04-21product
Workflow Samples | Markov

Markov published a browsable collection of computer-use workflow samples with synchronized video, narration annotations, and raw input events. The samples cover coding, design, browser use, gaming, video editing, and productivity workflows.

2026-03-17product
Markov AI releases the computer-use-large dataset

A post linking to Markov’s computer-use-large dataset announced an open-source computer-use recordings dataset with more than 10,000 hours across applications including Salesforce, Blender, and Photoshop.

2026-02-13product
Markov launches a dataset of computer-use recordings

Markov co-founder Dev Mandal announced a dataset of computer-use recordings intended to help build the next generation of computer-use agents, covering more than 300 tasks.

2025-08-05product
Markov announces a computer-use dataset

Markov announced what it described as an advanced computer-use dataset, built by capturing how people use computers and converting those interactions into data for model training.

Active Roles

0

No active roles right now.

Get notified when they post

Business Model

Markov appears to use a B2B data and licensing model, selling computer-use datasets, task data, and training environments to AI developers through negotiated engagements. Its public materials show data samples and a “Talk to us” sales CTA, but do not publish specific pricing.

Products

Computer Use Large: a large-scale dataset of professional-software screen recordings for training and evaluating GUI agentsGaming-500-hours: gameplay screen recordings with synchronized interaction events across hundreds of gamesCustom computer-use data and environments for training AI modelsPublic developer tooling, including the Markov Python SDK and frontend components

Customers

No publicly disclosed enterprise customers or customer logos identified

Tech Stack

Computer-use AI and GUI-interaction data pipelinesScreen-recording video, including H.264Mouse and keyboard event capture with JSON/NDJSON metadataPython public SDKOpenCV/computer visionFrontend UI toolingHugging Face dataset distribution

Competitors

Surge AI
Scale AI
Appen
OpenTrain AI