✳Physical AI / Evaluation protocol v0.1

Capability is a claim. Evidence is the work.

A protocol for testing robot policies where they have to perform. Connect data quality, open benchmarks and controlled site trials to evidence a model team can inspect and an operator can act on.

Public protocol design · October 2, 2026Download protocol template
A montage of bimanual robot demonstrations manipulating everyday objects across varied tabletop scenes.
Different objects. Different scenes. The same demand for evidence.Supplied reference imagery; task outcomes are not inferred from these frames.
01 / Evidence layer

Data integrity

Is the demonstration observable, synchronized and consistently labeled?

02 / Evidence layer

Policy capability

Can a frozen model complete a reproducible benchmark task?

03 / Evidence layer

Field performance

Does it work on the tested robot, at the tested site, with assistance accounted for?

01 / Evaluation notebook

Before the first trial, write the contract.

A benchmark becomes useful when the task, conditions and scoring rule stay fixed. This public design is the starting point for scoping an engagement, with thresholds and operating limits agreed for each robot and site.

  1. 01

    Freeze the question

    Name the workflow, robot embodiment, model checkpoint and task population. Register success predicates, time limit, permitted assistance, stop rules and a trial budget before seeing outcomes.

  2. 02

    Separate the splits

    Keep calibration and tuning outside the held-out test set. Group by site, recording session and object instance; do not scatter adjacent frames or near-duplicate trajectories across splits.

  3. 03

    Match the conditions

    Record start states, resets, camera calibration, payload and control frequency. Randomize or counterbalance model order within matched site, task and shift blocks. Blind reviewers to model identity where practical.

  4. 04

    Run and record

    Log every initiated attempt, including timeouts, interventions and aborts. Timestamp observations, actions, model latency and control ownership against synchronized video. Link each result to its episode manifest.

  5. 05

    Score and adjudicate

    Apply the registered success predicate at the required final state. Report subgoal progress separately. Have human reviewers adjudicate ambiguous cases and audit model-judge outputs against a held-out human-labeled set.

  6. 06

    Publish the evidence

    Report per-task and per-condition denominators, uncertainty, failure categories and exclusions. Release the protocol version, checkpoint identifiers and permitted artifacts so another team can reproduce the comparison.

Current deployment vertical / proposed trial scope

Move a sealed parts tote to a designated service staging bay.

Work near live electrical equipment requires a separately authorized task and site-specific operating limits.

Observable success predicate

Correct tote reaches the marked bay within the time budget; payload remains intact and the route stays inside the authorized zone.

Conditions to vary and report

  • Aisle layout and route
  • Tote instance and payload
  • Lighting and reflective surfaces
  • Pedestrian traffic windows
02 / Evaluation notebook

One episode. Every primitive.

A useful training demonstration needs more than a video. It needs timing, action boundaries, arm identity, object attributes and an auditable relationship between the labels and what happened.

42 sepisode duration
20labeled segments
98.5%reported timeline coverage
0reported overlapping segments
Supplied 42-second bimanual block-sorting annotation timeline: left and right arm tracks show approach, grasp, lift, move over box, place, return home and idle.
Supplied annotation summary for one teleoperated demonstration: 20 segments, 98.5% coverage, zero overlaps, one-frame gaps at 30 fps (approximately 33 ms), and attributes for 15 of 15 objects. Exact timestamps and the underlying annotation records were not supplied, so the timeline is shown as provided.
Supplied example / n = 1 episode

A demonstration, under the microscope.

Teleoperated · model-scored

Two kinds of evidence appear in the supplied figure. Rubric percentages describe a score relative to its maximum; probabilities describe the judge’s belief.

Values transcribed from the supplied TypeSafe JEV figure. The 0.89 task-success judgment is not an observed autonomous success rate. Reported judge confidence is not a statistical confidence interval. The 50% relevance score concerns this toy sorting task.

Make annotation quality measurable.

Compute coverage from the union of labeled intervals over a declared eligible duration, per arm or per episode. Measure overlaps on the same track; simultaneous actions across arms can be valid. Check boundaries with a stated frame tolerance and validate required object attributes.

Audit a held-out subset with two human reviewers and an adjudication rule. Retain disagreements, missing data and rubric versions. Pin model-judge inputs and outputs, then test agreement and calibration against the human reference.

Keep the signals separate.

The example’s completeness score of 91% and its 98.5% temporal coverage measure different things. Coverage asks how much time is labeled; completeness is a rubric judgment about the annotation.

A model’s probability or reported confidence is not a measured robot success frequency. Using a judge for review does not replace controlled autonomous trials. JEV rubric type reference

Supplied montage of a humanoid demonstrating manipulation of fruit, bowls, bags and a small crate on a table.
From demonstrations to an inspectable record of work.Reference montage; evaluation outcomes require episode logs and a scoring rule.
Evidence provenance / image estimates vs telemetry

A high score still needs a traceable source.

The supplied evaluation view pairs task scores with an explicit evidence-quality label. Its motion estimate is sourced from a 2D vision-language model, with unknown scale and a sparse image path measured in frame axes. Those coordinates do not establish metric travel, joint motion or contact force. Keep estimated signals separate from calibrated measurements, and show missing evidence alongside the score.

Supplied evaluation interface with robot episode cards and a 9.4 out of 10 rubric score. A motion provenance panel marks VLM 2D, unknown scale, and thin evidence, with image coordinates distinguished from metric distance and joint telemetry.
Supplied evaluation UI example. The visible 9.4/10 score and 0.03 frame-axis path are displayed by that interface; the scoring run, raw frames and measurement pipeline were not supplied for independent recomputation. Image-space motion is not robot telemetry.
03 / Evaluation notebook

Measure the work. And the help it needed.

A robot can finish a task while an operator does the hard part. Report autonomous completion, assistance, progress and operational burden as separate measures.

Field metrics and their reporting units
MeasureDefinitionReport
Unassisted task successAttempts meeting the final-state predicate within the time budget, with no human takeover, divided by valid initiated attempts after pre-registered infrastructure exclusions. Publish all initiated and excluded counts.k / N, with interval
Assisted completionAttempts that finish after operator input or physical help. Record assistance type, count and duration; keep these separate from autonomous success.k assisted / N
Subgoal completionCompleted required predicates divided by the registered predicate count. A partial score does not count as binary task success.Per-episode fraction
Cycle time & latencyStart-to-terminal-state time, with timeouts retained. Show median and p95 for successful runs, the timeout fraction, and policy latency separately.Seconds / milliseconds
Interventions & recoveryShare of attempts requiring a takeover; time under human control. Recovery is autonomous completion after a defined recoverable disturbance, on a separately declared subset.Episodes, seconds, subset N
Operational incidentsUnexpected contacts, dropped items, boundary violations and stops. Define events and severity before trials; report exposure hours and affected attempts.Count / exposure
Denominator policy

Every initiated attempt has a place in the report.

A policy-caused stop, timeout or human takeover remains a failed unassisted attempt. Assisted completion can be recorded separately. Infrastructure-invalid attempts may be excluded only under rules fixed before testing; preserve their logs and disclose counts, reasons and replacement trials. Include reset time and downtime when reporting tasks per wall-clock hour.

Generalization is a set of tests.

Separate zero-shot performance from performance after site adaptation. Disclose the adaptation budget and known training exposure; label unknown overlap as unknown. Do not average away a weak site or platform.

Holdout design for evaluating generalization
ShiftHold outReport separately
ObjectHeld-out instances, geometry, texture and payloadKnown category / new instance; new category separately
EnvironmentNew site or layout; cameras, lighting and clutterWithin-site changes / held-out sites
InstructionParaphrases, new referents and task compositionsSame intent / new composition
TemporalA later session, shift or deployment daySession-disjoint and time-disjoint results
EmbodimentA different platform, gripper or action interfacePer-platform results; adaptation budget disclosed
04 / Evaluation notebook

More trials. A clearer estimate.

An 80% result from ten attempts and an 80% result from a hundred attempts carry different uncertainty. Publish the numerator, denominator and uncertainty alongside every success percentage.

Mathematical planning illustration

How precision changes with trial count.

At a hypothetical observed success rate of 50 percent, 95 percent Wilson interval half-width falls from 26.3 percentage points for 10 independent trials to 20.1 for 20, 13.4 for 50, 9.6 for 100, 6.9 for 200 and 4.9 for 400.
Hypothetical 50% observed success; two-sided 95% Wilson intervals. Independent binary trials only. This figure is statistical planning, not measured CosmicBrain performance. Method: NIST proportion intervals.
Try the calculation / illustrative inputs

A percentage needs a denominator.

80.0% observed success

95% Wilson interval: 71.1% to 86.7%

This is a local mathematical calculator, not a CosmicBrain result. Assumes independent binary trials. For repeated trials within shared sites or sessions, use an analysis that accounts for clustering.

For comparisons, use matched task blocks and an analysis that respects task, session and site structure. Report macro averages across tasks as well as pooled episode counts, declare weights, and retain slice-level results. Choose sample size and stopping rules before observing results.

05 / Evaluation notebook

Build on open research. Keep its boundaries visible.

Open benchmarks make experiments repeatable and expose distinct failure modes. These references inform the protocol; benchmark adapters and test scope are selected per engagement.

01 / Real hardware · distributed evaluation

RoboArena

Distributed, double-blind comparisons across institutions and real environments on the DROID platform.

Protocol lesson
Borrow matched-condition comparisons and blind review for field trials.
Scope boundary
Its current platform does not establish performance across every robot brand.
02 / Simulation · lifelong manipulation

LIBERO

Task suites vary spatial relationships, objects, goals and longer-horizon behavior.

Protocol lesson
Use separate suites to examine transfer and retention, rather than one aggregate score.
Scope boundary
Simulation success is evidence about that suite, not deployment reliability.
03 / Simulation · kitchen manipulation

RoboCasa

Kitchen tasks with controlled scene and object diversity, plus training and evaluation splits.

Protocol lesson
Pin the task split, scenario seeds and horizon when comparing policies.
Scope boundary
Kitchen simulation does not validate food handling or an operational kitchen.
04 / Simulation · long-horizon activities

BEHAVIOR

Goal-predicate completion and efficiency metrics distinguish progress from complete task execution.

Protocol lesson
Track required subgoals alongside binary completion and time.
Scope boundary
Partial progress and human-normalized efficiency are different metrics.
05 / Simulation · manipulation tooling

ManiSkill

Distinguishes success at any point in a rollout from success at its end.

Protocol lesson
Specify terminal-state criteria, simulator backend and control configuration.
Scope boundary
Backend differences and reset conditions can change results.
06 / Open tooling · policy evaluation

LeRobot

Open policy evaluation tooling records outcomes and rollout videos for inspection.

Protocol lesson
Preserve model, environment and rollout configuration with evaluation artifacts.
Scope boundary
Using the tooling alone does not define a standardized field benchmark.
07 / Real-world data · policy learning

DROID

Real-world demonstration data, policy training code and a reproducible robot setup.

Protocol lesson
Document camera calibration, data provenance and target-domain adaptation.
Scope boundary
A training dataset is not itself a deployment leaderboard.

Primary project and documentation sources reviewed October 2, 2026. Pin a release or commit, task IDs, initial-state distribution, controller, observation mode and episode horizon. Scores from different tasks, embodiments, simulators or time budgets are not directly comparable. Simulation and live-site results belong in separately labeled reports.

06 / Evaluation notebook

A score is the beginning. Evidence is the deliverable.

Give the model team enough detail to debug, and the site operator enough context to judge the tested operating envelope. A result should trace back to a task, a run and an outcome.

A report another team can inspect.

  • Protocol manifestTask predicates, holdouts, horizons, assistance and exclusion rules.
  • System manifestCheckpoint hash, robot profile, cameras, calibration, control and inference runtime.
  • Episode ledgerOutcomes, failures, stops, interventions and every denominator.
  • Replayable evidenceSynchronized video, observation/action logs and annotation provenance, subject to agreed site permissions.
  • Analysis recordMetric code, uncertainty method, aggregation, slice results and judge audits.
Protocol starter files / v0.1

Make the evaluation explicit.

Protocol & episode template Reporting checklist Supplied example data

Templates are proposed specifications. Empty fields are left unmeasured. The example file contains values from the supplied figures, not raw episode records.

For model teams, site operators & hardware partners

Bring a model.
Define the test.

Scope a live-site evaluation around the workflow, platform and conditions that matter to you.

Read Cosmic 0.5