Preview  Clio is being built in the open. Score aggregation is in curation. How we source it

Benchmarks

How the field measures itself — simulation suites, real-robot protocols, distributed arenas, and the pooled datasets underneath them. Start here to know what a reported number means.

Loading benchmarks…
Descriptors are drawn from each benchmark's own paper or site, linked on every row.

What comes next

Aggregated scores. The next release adds reported results per benchmark, each row carrying a source link and one of three marks: official (from the benchmark's own leaderboard), self-reported (from the model's paper or release), or reproduced (by a third party). Numbers without a traceable source do not get published.

Mnesis Arena. Beyond aggregation, Mnesis Labs is building real-robot blind A/B evaluation — paired rollouts under matched conditions, ranked with confidence intervals rather than single-run point scores. Those results will appear here as a separate, clearly-labelled track once the protocol has enough samples to mean anything.