Benchmarks
How the field measures itself — simulation suites, real-robot protocols, distributed arenas, and the pooled datasets underneath them. Start here to know what a reported number means.
What comes next
Aggregated scores. The next release adds reported results per benchmark, each row carrying a source link and one of three marks: official (from the benchmark's own leaderboard), self-reported (from the model's paper or release), or reproduced (by a third party). Numbers without a traceable source do not get published.
Mnesis Arena. Beyond aggregation, Mnesis Labs is building real-robot blind A/B evaluation — paired rollouts under matched conditions, ranked with confidence intervals rather than single-run point scores. Those results will appear here as a separate, clearly-labelled track once the protocol has enough samples to mean anything.