Companies

Refresh

refresh.dev

Refresh builds reinforcement-learning environments that help AI labs train coding and computer-use models.

HQSan Francisco, California, United States
Employees1-50
Jobs checked 15h ago
AI / MLData Labeling / TrainingB2B SaaS

About

Refresh builds reinforcement-learning environments for coding and computer use, targeting AI labs that develop and train advanced models. Its differentiation is combining AI and human expertise to create environments with verifiable reward signals and rich feedback for model learning.

Market

Refresh competes in the reinforcement-learning environment, agent evaluation, and post-training infrastructure market for coding and computer-use AI systems. It differentiates through realistic simulated software worlds, automatically verifiable rewards, and failure-driven tasks sourced from expert engineers and validated against frontier models, rather than relying only on generic human-generated trajectories or isolated coding sandboxes.

Target Customers

Refresh primarily serves frontier AI labs and enterprise AI/software teams that are building, evaluating, and post-training coding or computer-use agents. The likely buyers are model-training, agent-platform, and applied-AI leaders; the evidence identifies labs and enterprises but does not specify a narrower company-size segment.

At a Glance

Problem

Frontier AI labs need models to perform extended, real software-engineering and computer-use tasks, not merely demonstrate general language competence. Refresh addresses the gap between models that learn fundamentals during pretraining and agents that can reliably use specialized tools over long horizons, such as debugging failing CI pipelines, patching security vulnerabilities, building through hundreds of tool calls, or operating connected enterprise software. The economic pain is that labs need realistic tasks with objective success criteria, while enterprises need to convert high-value internal workflows into measurable, repeatable agent tasks rather than rely on subjective judgment.

Product / Service

Refresh provides simulated software worlds and reinforcement-learning environments that reproduce terminals, applications, and human workflows. It packages them as evaluations, which benchmark how well frontier models perform real computer work, and as training gyms, where verifiable rewards provide a learning signal. Its environments include Terminal-Bench-style long-horizon engineering tasks, MCP tool gyms, and high-fidelity clones of connected software suites such as EHR and enterprise applications.

Its delivery model combines expert task sourcing with bespoke environment building. Engineers in Refresh’s network identify real tasks that break leading models; Refresh reproduces the failure, codifies it into a task with a verifiable reward, and validates that frontier models fail at high rates before inclusion. Labs use the environments to measure and improve model, tool, and prompt performance, while enterprises can turn high-value workflows into measurable, trainable tasks and request bespoke datasets.

Market

Refresh operates in AI infrastructure and enterprise AI tooling, specifically simulation environments, reinforcement-learning training gyms, and evaluations for coding and computer-use agents. Its primary buyers are frontier AI labs training models and enterprises evaluating or improving agents for economically valuable workflows. The evidence does not provide a definitive Refresh-specific competitor list; a separate third-party snippet names CoLab Software, GrabCAD, and Encube, but that record is labeled “Rev1,” so those companies are best treated as adjacent or indicative rather than confirmed direct competitors.

Traction is early but not merely conceptual. Refresh entered YC’s Spring 2025 batch and is listed as an active San Francisco company with eight employees; its own materials say it is trusted by frontier AI labs and enterprises and is actively working with frontier labs. Public funding databases indicate seed financing but disagree on the amount: PitchBook reports $500,000 and Y Combinator investment, while Tracxn reports $1.3 million from a January 2026 seed round backed by Black Nova and Archangel. No revenue figure is disclosed in the available evidence, so the company cannot be confirmed as either revenue-generating or definitively pre-revenue.

Founders & Leadership

Christopher SettlesFounder
Co-founder & CEO
Erik QuintanillaFounder
Co-founder & CTO

Funding History

2025-03
Y Combinator accelerator financing$500K

Y Combinator

Recent News

2026-07-16
Refresh: RL environment vendor profile | RL List

RL List profiled Refresh (YC X25) as a provider of simulation engines and reinforcement-learning environments with verifiable rewards for coding and computer use. The profile says Refresh partners with frontier labs and enterprises to train software-engineering and computer-use capabilities.

2026-04-08
Jobs at Refresh (P25) | Y Combinator's Work at a Startup

Y Combinator’s company profile describes Refresh as building training gyms for computer use and software-engineering work, with reinforcement-learning environments for coding and computer-use capabilities in LLMs. It also identifies Refresh as an eight-person San Francisco company founded by Erik Quintanilla and Christopher Settles.

2026-02-17product
Gauntlet 4K RLVR

Refresh launched Gauntlet 4K RLVR, an RLVR dataset curated from Gauntlet and delivered in Harbor format for reinforcement learning. This is the clearest product or dataset launch identified in the period.

2025-12-20
Open Roles | Refresh

Refresh announced hiring for ex-founders, go-to-market engineers, and research engineers to build the simulation engines used by frontier labs. The hiring update signals continued expansion of its coding and computer-use environment business.

Active Roles

0

No active roles right now.

Get notified when they post

Business Model

Refresh uses a B2B paid developer-tool model: a third-party pricing listing reports plans starting at $29 per month, while company records describe partnerships with AI labs working on coding or computer use. Enterprise contract terms and other revenue streams are not disclosed in the available evidence.

Products

Coding and computer-use simulation environments and evaluationsRL training gyms with verifiable rewardsTerminal-Bench-style, long-horizon software-engineering tasksLong-horizon MCP tool gymsHigh-fidelity computer-use software worlds for enterprise workflowsHarbor framework for agent evaluations and RL environmentsWeb-eval-agent MCP server for autonomous web-application evaluation

Customers

Unnamed frontier AI labs and enterprises; no customer names or logos were publicly disclosed in the reviewed sources.

Tech Stack

Reinforcement learning (RL) environmentsLLM and AI-agent evaluation and post-trainingSimulated software worlds for coding and computer useVerifiable reward instrumentationMCP (Model Context Protocol) tool environmentsTerminal-based software-engineering environmentsWeb-application evaluation agents

Competitors

OpenTrain
Scale AI
Rubrik