Companies

hud

hud.so

HUD provides reinforcement-learning environments and evaluation infrastructure for training and testing AI agents.

HQSan Francisco, California, United States
Employees1-50
13 active roles
Jobs checked 15h ago
AI / MLAI InfrastructureB2B SaaSPre-Seed / Seed

About

HUD builds reinforcement-learning environments and evaluation infrastructure for AI agents, helping businesses and AI labs train specialized agents, assess models, and produce post-training datasets. Its differentiation is a combination of open-source development tools, scalable cloud evaluations, production telemetry, and enterprise RL workflows.

Market

HUD competes in AI-agent evaluation and reinforcement-learning environment infrastructure, with an emphasis on frontier-model post-training data. Its differentiation is an integrated workflow that lets builders create environments, publish them for labs to run without custom integration or separate hosting, generate trajectory data and reward signals, and feed those results into training pipelines; competitors such as Harbor and Prime Intellect are more focused on customizable containerized RL workflows or open-source model training.

Target Customers

HUD primarily targets frontier AI and model-development labs, especially teams responsible for agent evaluation, post-training, reinforcement learning, and model reliability. It also serves businesses and technical teams that build RL environments and want to distribute them to AI labs.

At a Glance

Problem

AI-agent developers and research teams need realistic, repeatable environments in which to evaluate and train agents, but building those environments, task sets, graders, and scalable execution infrastructure is technically demanding and slow. HUD targets the resulting data bottleneck: labs need high-quality, domain-specific post-training data and reliable reward signals, while teams otherwise face long evaluation cycles and the cost of operating large numbers of environments. HUD highlights that its infrastructure can run thousands of environments concurrently, with full benchmarks typically costing roughly $1–$10.

The central use case is turning a real software workflow—such as a browser, desktop application, coding environment, spreadsheet, API, or chat interface—into an agent-training and evaluation environment quickly. That lets a team benchmark an agent on realistic tasks, inspect failures, and reuse the resulting graded rollouts as training data rather than treating evaluation and training as separate workflows.

Product / Service

HUD combines an open-source SDK with a hosted cloud platform and enterprise services. Developers define environments, tasks, capabilities, and verifiers locally, then deploy them to HUD, run evaluations in parallel, compare models, and inspect telemetry and traces. Its protocol-first design exposes an environment manifest, task prompts, and rewards, so different models and agent harnesses can operate against the same environment; supported capabilities include shells, MCP tools, browsers, computer interfaces, and robotics integrations.

The delivery model is usage-based cloud infrastructure at $0.50 per environment hour, with a free local SDK, production benchmarks, live debugging, and parallel runs. Enterprise customers can purchase private benchmarks and RL workflows, custom environments, on-premise deployment, and engineering support. Because every rollout returns a reward and trace, HUD positions the same environment as a reusable asset for both model evaluation and reinforcement-learning training.

Market

HUD competes in the emerging AI-agent evaluation, benchmarking, reinforcement-learning environment, and post-training-data market. Its differentiation is the combination of environment creation, scalable execution, evaluation telemetry, and reward-generating training workflows, along with a vendor platform intended to connect post-training suppliers with AI labs. Adjacent or overlapping alternatives include Braintrust for agent-evaluation workflows, Browserbase for browser-agent training and evaluation, and Scale’s RL-environment offering; broader AI-evaluation tools such as Arize Phoenix, Promptfoo, Galileo, and related platforms also occupy parts of the surrounding category.

HUD appears to have moved beyond a purely pre-revenue or speculative stage, although the available evidence does not provide audited revenue. Y Combinator reports that more than 50 businesses use HUD to build environments, sell them to AI labs, or train their own models, while the company was founded in 2025 and listed a 15-person team. HUD also announced a $16 million Series A led by Standard Capital on June 30, 2026. A third-party estimate reported approximately $1.1 million in 2025 ARR, but that figure should be treated as directional rather than verified company disclosure.

Founders & Leadership

Jay RamFounder
Founder & CEO
Lorenss MartinsonsFounder
Founder & CPO
Parth PatelFounder
Co-Founder & CTO

Funding History

2025-01
Seed$500K

Y Combinator

2026-06
Series A$16M

Standard Capital

Recent News

2026-07-07product
How to Sell Startup Assets to Model Labs

HUD introduced its vendor-platform process for helping shutting-down startups sell codebases and internal engineering data to AI labs. The process covers intake, NDAs, scoping, delivery, licensing, and payout.

2026-06-30funding
Announcing HUD's $16M Series A

HUD announced a $16 million Series A led by Standard Capital. The company said the funding will support software for turning data into better AI and a platform where users can build and sell reinforcement-learning environments.

2026-04-29
HUD Rating, Review, Features

A third-party AI-evals profile described HUD as a Y Combinator W25, approximately 15-person open-source platform for building reinforcement-learning environments and evaluations for computer-use agents.

2026-04-06product
Best Platforms for Publishing RL Environments to Model Labs

HUD published a guide ranking platforms for publishing reinforcement-learning environments and positioned hud.ai as the leading option for environment builders seeking use by frontier labs. The guide highlights evaluation, training-data generation, deployment, and existing model-lab usage, including the Autonomy-10 benchmark used to evaluate OpenAI Operator.

2026-01-20
Building an RL Environment to Train Agents for Production Debugging

HUD’s official blog listed a research note on building a reinforcement-learning environment for training agents to handle production-debugging workflows.

2025-10-01partnership
Evaluating Agents on Financial Analyst Workflows (SheetBench)

HUD and Sepal AI created SheetBench-50, a financial-analyst-grade benchmark that evaluates AI agents on real spreadsheet and financial workflows. The case study represents a partnership or collaborative benchmark-development announcement.

Active Roles

13
San Francisco/Legal/2d ago
San Francisco/Engineering/33d ago
San Francisco/Engineering/33d ago
San Francisco/Engineering/33d ago
San Francisco/Engineering/33d ago
San Francisco/HR & Recruiting/33d ago
San Francisco/Engineering/33d ago
San Francisco/Marketing/33d ago
San Francisco/Engineering/33d ago
San Francisco/Sales Engineer/33d ago
San Francisco/Marketing/33d ago
San Francisco/Forward-Deployed Engineer/33d ago

Business Model

HUD offers a free open-source SDK, charges $0.50 per environment hour for its managed cloud evaluation platform, and provides custom-priced enterprise solutions. Enterprise offerings include private benchmarks, RL workflows, on-premise deployment, and dedicated engineering support.

Products

Environment SDK for defining evaluations, environments, and verifiersTraining + Eval Platform for running and evaluating models on environmentsCUA Evals framework for evaluating computer-use agentsHUD Vendor Platform for connecting post-training environment suppliers with AI labs

Customers

No named enterprise customers publicly disclosed in the reviewed sources

Tech Stack

PythonRustPostgresDockerAWSNext.jsReactMCPReinforcement learningLLM/AI-agent evaluationGRPO

Competitors

Harbor
Prime Intellect
Mechanize
AfterQuery
Bespoke Labs
Huzzle Labs