Status

Status & illustrative leaderboard

The headline signal is skill in bits over the own-routine baseline R2, with calibration and permutation specificity as gates. The "overall" column is an illustrative roll-up for display only, not a canonical composite.

No empirical leaderboard results exist for v1.0. Every row below is a synthetic illustrative mock baseline on benchmark protocol v1.0 — not a real submission, not an official result, and not evidence about any system or product. Higher is better for every metric except calibration error and long-horizon degradation (lower is better). A large prediction count from few targets is not a large independent sample.

Loading leaderboard…

Metric definitions
  • Overall score — illustrative pilot roll-up for display only.
  • Target adaptation gain — Skill in bits over the own-routine baseline R2 (headline).
  • Calibration error — top-label ECE; lower is better.
  • Evidence attribution — whether cited evidence supports the prediction.
  • Counterfactual consistency — coherence of action-conditioned predictions.
  • Temporal-order sensitivity / wrong-target penalty — skill lost under shuffled history / the permutation gate (higher = more target-specific).
  • Long-horizon degradation — skill decay at long horizon; lower is better.
  • Modality contribution — marginal skill from added evidence streams.