
Robocurve
运营中Real-world evals for robots
公司简介
Robocurve builds open-source tools and independent benchmarks to measure how well robots can do real-world jobs. Instead of relying on unverified demo videos from frontier labs, we score their models on reproducible benchmarks that anyone can trust. Today there are no well-run, standardized robotics benchmarks. Labs evaluate in-house, and no independent group has stepped in to run continuous benchmarking as a service. The result is that no one actually knows how good anyone else is, or where the real frontier sits. Good benchmarks require operating and maintaining physical hardware and real-world setups. Simulation only goes so far, since models that look strong in sim can show large performance gaps once deployed in the real world. Building real-world benchmarks means coordinating job-domain experts, evals engineering, and hands-on robotics all at once. Frontier labs are targeting general-purpose robotics by 2028, yet the field of robotics evals barely exists. Whoever builds the trusted measure of robot capability becomes the reference everyone relies on. We combine backgrounds in AI evals and robotics to build and grow this field as fast as possible. We already shipped v1 of our open-source framework, Inspect Robots, and ran our first pilot scoring a frontier model on a real robot.
创始团队
Jay Chooi· FounderFounder and CEO of Robocurve. MA Statistics and BA CS/Math from Harvard. Jay builds evals of frontier AI and forecasts its progress. Previously Research Fellow at MATS, researcher at the UK AI Security Institute, and top contributor to Inspect Evals, the UK government's AI eval framework. Published at ACM EC, ICML, ACL, and EMNLP. Called the 2024 election correctly in all 50 states. Won a Rhodes Scholarship. Gold medal at the International Olympiad on Astronomy and Astrophysics.
Aris Zhu· FounderFounder and CTO of Robocurve. Studied CS & Physics at Harvard. Aris builds robotics systems and AI agents across research and production. At Amazon AGI Labs, her test-time scaling research improved the NovaAct browser agent on WebVoyager. At Yondu Robotics (YC W24), she architected the navigation stack and fleet management system for humanoid robots. At Amazon Robotics she deployed a vision model to edge devices in production. She co-authored a paper in IEEE Robotics and Automation Letters.
产品发布 · 1 次发布
We measure how good AI and robots are in the physical world.
变化历史 · 4 条记录
- 2026年8月30日
- 团队从 2 人扩张到 3 人19:00
- 更新了标签19:00
- 2026年8月28日
- 更新了一句话简介19:00
- 2026年7月17日
- 更新了一句话简介19:00