Physical-AI Evaluation · Live at Robotics Center

Lab-ready isn't
deployment-ready.

SVRC Eval closes the sim-to-real loop: refine your data, run your policy against real acceptance criteria, and feed every failure back into targeted collection — one stack across hardware, data and evaluation.

Operator-in-the-loopFailure-driven targetingAcceptance criteria, not vanity metrics
Illustrative example · Sample readiness output
eval-runs / readiness● sample report
Readiness Reportpi0-fast · v0.3
REAL-LIFERL²LOOP
DEPLOYOn real HWTarget environment
CAPTUREFailuresFull rollout log
EVALUATEAcceptanceOperator + gates
TARGETNext dataFailure-driven
success 88.5% 52 actions 2 targets queued

WHY TEAMS STALL

The sim-to-real gap is a loop problem.
Not a model problem.

Most robotics teams don't stall on model capability — they stall on the loop between a policy that works in the lab and one that survives a real workcell. SVRC Eval is the stack that runs that loop: refine the data, evaluate against the criteria you'll actually be judged on, and feed every failure back into targeted collection.

WHERE PROJECTS BREAK

Three failure modes inside the loop.

After two years supplying hardware and data to robotics teams, the place projects stall is remarkably consistent — and it's not where people expect.

01 · PHYSICS

Actuator physics doesn't transfer

A policy trained in sim requests torque profiles real actuators can't sustain — thermal throttling, joint backlash and unit-to-unit variance mean two of the same arm off the line don't share dynamics. "Embodiment-agnostic" turns out embodiment-specific.

02 · DATA

Teleop data is structurally off-policy

Bootstrap demonstrations come from a human operator, not your model. That covariate shift doesn't shrink by collecting more of the same — it shrinks through on-policy rollouts on real hardware, with failures fed back into targeted collection.

03 · CRITERIA

Nobody defines "working" before deploy

Teams ship on lab success rates, then meet the customer's acceptance criteria in the long tail — the wrinkled shirt, the 4pm glare, the pallet 3cm off spec. Without an eval harness that encodes those criteria, readiness is a guess.

THE LOOP WE RUN

RL² — reinforcement learning in real life.

Real hardware, real failure distributions, run as a cycle instead of post-hoc debugging. Each stage maps to something we operate today.

01 / 05

Deploy on real hardware

Run the policy on the actual arm, hand and camera in the target environment — not a sim twin. Real actuator dynamics, real latency, and the real failure surface.

  • On-policy rollouts on real hardware
  • Shadow or execute mode
  • Every action and outcome recorded

WHAT'S IN THE PLATFORM

One stack, from raw capture to a readiness gate.

SVRC Eval is built on the data infrastructure we already run in production. Every capability below is live today.

DATA REFINERY

Refine capture automatically

Egocentric and teleop capture, refined automatically: quality control, hand-mesh extraction, action segmentation and LeRobot-v2 packaging. Run it from the workbench or the centeros refinery CLI.

Live today
MANAGED TELEOP

On-policy & targeted data

A real operator network and a WebSocket teleop layer with MCAP ingestion and time-synced multi-sensor alignment. This is the source of your on-policy and targeted data.

Live today
MODEL REGISTRY

Versioned policies & rollouts

Register and version your policies, then run inference sessions in shadow or execute mode and record every action and outcome for evaluation.

Live today
EVALUATION RUNS

Runs & regression gates

Define pass criteria — success rate, latency, recovery — per task and per dataset. Runs are tracked over versions, and a drop below threshold can auto-trigger a warm-start retrain.

Live today
CLI · API · MCP

Scriptable everywhere

Everything is scriptable through the centeros CLI and exposed as MCP tools, so the loop runs from your own pipeline or from an agent — not just a dashboard.

Live today
DATASET ACCESS

Index, inspect & share

Index, inspect and share datasets with signed, time-limited access, and pull targeted slices back into the loop as failures dictate.

Live today

Honest boundary: we're not a foundation-model lab or an inference vendor. We own the loop that gets a policy from demo to production — the evaluation and the failure-driven data that make it better.

HOW READINESS SCORING WORKS

Acceptance criteria, not vanity metrics.

Lab success rate is not readiness. We encode customer acceptance criteria — the specific tasks, environments and edge cases you'll be judged on in production — into an eval you can re-run on every model version.

01

Automated regression gates

Metric thresholds tracked across versions, so a regression is caught before it ships — the robotics equivalent of a CI gate on your policy.

02

Operator-in-the-loop evaluation

A shirt folded "correctly" isn't a unit test. Our operator network runs real rollouts and judges outcomes against your criteria, so the score reflects the workcell, not a proxy.

03

Failure-driven data targeting

Every failed rollout points at the exact scene, object or condition to collect next. The eval doesn't just grade — it writes your next data order.

THE OUTPUT

Every run leaves a readiness report.

SVRC Eval answers the question that matters: under which real conditions can this policy complete this task, reliably and economically?

Illustrative example · Sample output
SAMPLE READINESS REPORT

pi0-fast · v0.3 · garment fold

PASS vs acceptance criteria
MODELv0.3pi0-fast
SESSIONS3sample runs
ACTIONS52decisive
SUCCESS88.5%46 ✓ / 6 ✗ · 161 ms avg
01

Top failure mode: operator_rejected · 4 rollouts.

02

Secondary: timeout · 2 rollouts past action budget.

03

Next action: targeted capture on flagged scenes → re-evaluate.

Sample output from the eval-runs pipeline scored on real rollout telemetry — not a published benchmark or a claim about a deployed policy. Fully-automated policy scoring across arbitrary tasks is where we're investing next; today, readiness combines automated gates with our operator network so the number means something in the real world.

CONTINUE THE LOOP

Need the hardware to run evals on?

WORK WITH US

Hitting one of these three walls?
Let's run the loop.

We're comparing notes with teams building real deployments — and we host builders at 90 Welsh St in San Francisco. If your policy works in the lab but not the workcell, book a walkthrough of the refinery, managed teleop and evaluation runs.