Refine capture automatically
Egocentric and teleop capture, refined automatically: quality control, hand-mesh extraction, action segmentation and LeRobot-v2 packaging. Run it from the workbench or the centeros refinery CLI.
SVRC Eval closes the sim-to-real loop: refine your data, run your policy against real acceptance criteria, and feed every failure back into targeted collection — one stack across hardware, data and evaluation.
pi0-fast · v0.3WHY TEAMS STALL
Most robotics teams don't stall on model capability — they stall on the loop between a policy that works in the lab and one that survives a real workcell. SVRC Eval is the stack that runs that loop: refine the data, evaluate against the criteria you'll actually be judged on, and feed every failure back into targeted collection.
WHERE PROJECTS BREAK
After two years supplying hardware and data to robotics teams, the place projects stall is remarkably consistent — and it's not where people expect.
A policy trained in sim requests torque profiles real actuators can't sustain — thermal throttling, joint backlash and unit-to-unit variance mean two of the same arm off the line don't share dynamics. "Embodiment-agnostic" turns out embodiment-specific.
Bootstrap demonstrations come from a human operator, not your model. That covariate shift doesn't shrink by collecting more of the same — it shrinks through on-policy rollouts on real hardware, with failures fed back into targeted collection.
Teams ship on lab success rates, then meet the customer's acceptance criteria in the long tail — the wrinkled shirt, the 4pm glare, the pallet 3cm off spec. Without an eval harness that encodes those criteria, readiness is a guess.
THE LOOP WE RUN
Real hardware, real failure distributions, run as a cycle instead of post-hoc debugging. Each stage maps to something we operate today.
Run the policy on the actual arm, hand and camera in the target environment — not a sim twin. Real actuator dynamics, real latency, and the real failure surface.
WHAT'S IN THE PLATFORM
SVRC Eval is built on the data infrastructure we already run in production. Every capability below is live today.
Egocentric and teleop capture, refined automatically: quality control, hand-mesh extraction, action segmentation and LeRobot-v2 packaging. Run it from the workbench or the centeros refinery CLI.
A real operator network and a WebSocket teleop layer with MCAP ingestion and time-synced multi-sensor alignment. This is the source of your on-policy and targeted data.
Live todayRegister and version your policies, then run inference sessions in shadow or execute mode and record every action and outcome for evaluation.
Live todayDefine pass criteria — success rate, latency, recovery — per task and per dataset. Runs are tracked over versions, and a drop below threshold can auto-trigger a warm-start retrain.
Live todayEverything is scriptable through the centeros CLI and exposed as MCP tools, so the loop runs from your own pipeline or from an agent — not just a dashboard.
Index, inspect and share datasets with signed, time-limited access, and pull targeted slices back into the loop as failures dictate.
Live todayHonest boundary: we're not a foundation-model lab or an inference vendor. We own the loop that gets a policy from demo to production — the evaluation and the failure-driven data that make it better.
HOW READINESS SCORING WORKS
Lab success rate is not readiness. We encode customer acceptance criteria — the specific tasks, environments and edge cases you'll be judged on in production — into an eval you can re-run on every model version.
Metric thresholds tracked across versions, so a regression is caught before it ships — the robotics equivalent of a CI gate on your policy.
A shirt folded "correctly" isn't a unit test. Our operator network runs real rollouts and judges outcomes against your criteria, so the score reflects the workcell, not a proxy.
Every failed rollout points at the exact scene, object or condition to collect next. The eval doesn't just grade — it writes your next data order.
THE OUTPUT
SVRC Eval answers the question that matters: under which real conditions can this policy complete this task, reliably and economically?
Top failure mode: operator_rejected · 4 rollouts.
Secondary: timeout · 2 rollouts past action budget.
Next action: targeted capture on flagged scenes → re-evaluate.
Sample output from the eval-runs pipeline scored on real rollout telemetry — not a published benchmark or a claim about a deployed policy. Fully-automated policy scoring across arbitrary tasks is where we're investing next; today, readiness combines automated gates with our operator network so the number means something in the real world.
CONTINUE THE LOOP
WORK WITH US
We're comparing notes with teams building real deployments — and we host builders at 90 Welsh St in San Francisco. If your policy works in the lab but not the workcell, book a walkthrough of the refinery, managed teleop and evaluation runs.