About
Robocurve is a San Francisco-based Public Benefit Corporation building open-source tools and independent benchmarks for evaluating AI-controlled robots in real-world tasks. It serves robotics and AI researchers, developers, and organizations seeking trustworthy capability measurements, differentiating itself through reproducible, real-world evaluations rather than unverified demonstration videos.
Market
Robocurve competes in the emerging physical-AI and robotics evaluation/benchmarking market, helping organizations measure whether AI models and robot embodiments can perform economically relevant real-world tasks. It positions itself as an independent, open-source, public-benefit reference layer rather than as a robot manufacturer or foundation-model vendor. Its differentiation is a real-world-first, embodiment-agnostic evaluation approach with reproducible logs, cross-model and cross-robot execution, and integrations spanning ROS, Isaac Lab, simulation, VLAs, and frontier LLMs.
Robocurve primarily serves frontier AI labs, robotics companies, and academic or industrial research teams developing LLM/VLA models and general-purpose robots; likely buyers are AI/robotics research, evaluation, and engineering leads seeking independent, reproducible capability scores. Its broader stakeholder audience includes governments, policymakers, and market participants that need forecasts of economically relevant robot capabilities.
At a Glance
Problem
Robocurve addresses a measurement problem in physical AI: robotics laboratories often evaluate their systems privately, while public understanding still relies heavily on impressive but unverified demonstration videos. That creates an opaque market in which customers, investors, researchers, and competing labs cannot reliably determine what a robot or the model controlling it can actually do. The economic pain is uncertainty, duplicated in-house testing, and the risk of deploying expensive robotic systems before their capabilities are proven.
The main use case is independent, reproducible scoring of frontier AI models on real robots performing useful physical jobs. Robocurve’s benchmark portfolio spans tasks such as kitchen manipulation, laundry, coffee making, and data-center construction, including server racking and cable routing. DataCenterBench is a particularly consequential example because it targets the physical work required to expand compute infrastructure.
Product / Service
Robocurve combines open-source evaluation infrastructure with independent real-world benchmarks. Its Inspect Robots framework is designed to run any model on any robotic embodiment and benchmark, producing full trace logs and live visualization. It is real-world first but supports simulation, integrates with robotics and model-development ecosystems such as ROS and Isaac Lab, and is released under an MIT license. World Evals provides a catalog of reproducible benchmarks across arms, dexterous hands, humanoids, mobile robots, and quadrupeds.
The delivery model appears deliberately hybrid: researchers and robotics companies can use the open-source tools themselves, while Robocurve invites organizations to have their models or robots evaluated and has already run a pilot scoring a frontier model on a physical robot. The benefit is a trusted, comparable performance record that can replace subjective demos and support deployment, model selection, and forecasting of when particular physical tasks are ready for automation.
Market
Robocurve competes in the emerging physical-AI and robotics-evaluation market, at the intersection of hard-tech robotics, AI evaluations, benchmarking, and open-source developer infrastructure. It is not primarily a robot manufacturer; its proposed position is as an independent third-party reference for robot capability. The company identifies METR, Epoch, and Artificial Analysis as analogues that became trusted evaluators for language models. No equally established direct physical-robot evaluator is identified in the available company materials, suggesting that the category is still being formed.
The company is at an early stage. Y Combinator lists Robocurve as an active Summer 2026 company with a two-person team; the company says it has shipped version one of Inspect Robots and completed its first physical-robot pilot. Dealroom lists a $125,000 Y Combinator seed entry in June 2026. No public revenue, customer count, or repeat-commercial-contract figure is disclosed in the reviewed sources, so for diligence purposes Robocurve is best treated as pre-scale and likely pre-revenue rather than as an established commercial business.
Founders & Leadership
Funding History
Y Combinator
Recent News
Menlo Times highlighted Robocurve as a YC S26 company founded by Jay Chooi and Aris Zhu. It described the company’s open-source tools and independent benchmarks for measuring how well robots perform real-world jobs.
Y Combinator’s Launch YC listing announced Robocurve’s focus on measuring how capable AI and robots are in the physical world.
This article describes Inspect Robots as Robocurve’s open-source, MIT-licensed framework for evaluating vision-language actions models and large language models on real robots.
Y Combinator’s company launch profile reported that Robocurve had released v1 of Inspect Robots, an open-source framework for evaluating VLAs and LLMs on robots. The launch materials say the framework supports transcript logging, tracebacks, Rerun, digital-twin simulation, and integrations covering dozens of robotic models.
Active Roles
0No active roles right now.
Get notified when they postBusiness Model
Robocurve’s public pricing and revenue model are not disclosed. Its visible offering is MIT-licensed, open-source evaluation software and benchmarks, while its website invites organizations to discuss evaluations and funding support, suggesting a services- or funding-supported model rather than a published subscription product.