What TargetSpace-Bench measures
TargetSpace-Bench evaluates target-specific prediction under partial observation: whether a system can maintain and use a target-specific predictive representation for a specific target from passive, multimodal, non-IID observation — measured through calibrated, sealed prediction.
Definition & scope
A target
A persistent, evolving entity tracked over time: a person, agent, system, organization, environment, or process. The flagship track targets a consenting individual; other tracks target patients, energy assets, embodied agents, and projects.
A target state
The latent configuration a target acts to reach or maintain — a commitment, priority, constraint, or regime — evaluated only through its externally observable consequences. The object to prediction is the transition between target states.
Passive longitudinal observation
Evidence accrued over time without acting on the target: metadata, text, audio, passive multimodal streams, location, physiology. The benchmark measures whether richer, correctly-ordered observation improves target-specific predictions.
The capability under test
Can a model maintain and use a target-specific predictive representation, rather than merely a model of generic scenes or generic events? A high score means calibrated, prospective, target-specific skill — nothing about a target's inner life.
A worked example, end to end synthetic · illustrative
One pass through the TargetSpace loop on a made-up target with made-up numbers — no participant data, no empirical result. The loop:
1 · Target & passive evidence stream
The target is a synthetic person. Two instrumented streams record timestamped evidence — calendar metadata and app-focus samples — without the target reporting, recalling, or performing anything for the measurement itself:
- days 1–7 — a recurring “draft spec review” block sits each weekday at 09:00; on days 3, 5, and 7 it is rescheduled to later the same day.
- day 7, 16:10 — app focus shifts to a competing deliverable and stays there for the rest of the session.
- day 7, 18:02 — a meeting request from another team lands on the day-8 09:00 slot.
- day 8, 07:35 — the morning’s first app-focus samples are on the competing deliverable again.
2 · Cutoff, query & possibility space
Evidence seals at day 8, 08:00. The sealed query: does the “draft spec review” commitment complete, defer, cancel, or get replaced? Before the resolution window closes, several future target states remain operationally possible:
The benchmark asks one question: does the system assign better-calibrated probabilities over these possibilities than the baselines do?
3 · Sealed distributions vs the baselines
| Possible state | Model (sealed) | R1 population prior | R2 own-routine |
|---|---|---|---|
| complete | 0.14 | 0.55 | 0.70 |
| defer resolved outcome | 0.78 | 0.25 | 0.18 |
| cancel | 0.04 | 0.10 | 0.06 |
| replace | 0.04 | 0.10 | 0.06 |
4 · Resolution & per-instance score
Illustrative single-instance numbers only. A reportable run aggregates many sealed instances per target, and skill counts only after the calibration and wrong-target permutation gates — genuine target-specific skill collapses under wrong-target permutation.
Evidence cost. Of the two streams, the app-focus samples are the more invasive — and that stream must earn its lift. The objective is validated lift from minimum sufficient observation, not maximal capture: at equal gated skill, the less invasive configuration wins.
A full end-to-end prototype (registry, sealing, scoring, evidence efficiency) runs at targetspace.ai/platform.
Different from generic dynamics
TargetSpace is not a video-realism, robotics-manipulation, or intuitive-physics benchmark. Those evaluate generic dynamics, realism, and control. TargetSpace evaluates target-specific dynamics: whether prediction improves when a model is given the correct target's history in the correct temporal order. A system can render plausible scenes, generate fluent continuations, or execute dexterous control and still be unable to track this target and prediction where it turns next.
How it complements other benchmark families
The families are complementary, not competing. Each is the right instrument for a different question; TargetSpace is compared with them only on shared axes.
| Benchmark family | Primary object of evaluation | Typical input | Typical output | What it misses | How TargetSpace complements it |
|---|---|---|---|---|---|
| Physical reasoning / intuitive physics | physical plausibility, object permanence, causality, spatial continuity | short scenes / clips | plausibility or violation judgment | a persistent target; longitudinal adaptation; calibration over time | adds target-specific dynamics over generic physical law |
| Video generation / world simulation | visual realism, temporal consistency, plausible scene evolution | context frames | generated continuation | target identity; sealed prospective scoring; proper calibration | scores latent target-state transitions, not surface reconstruction |
| Embodied robotics | action utility, policy evaluation, manipulation / control success | proprioception, sensors, actions | actions; achieved configuration | passive longitudinal inference; calibration; an own-routine baseline | scores passive consequence predictions, not control |
| Symbolic / event predictioning | probabilistic prediction of public events | question + context | calibrated probability | a tracked individual target; own-routine R2; permutation specificity | adds the target as the unit, with R2 + permutation controls |
| Agent memory / personalization | recall, preference modeling, retrieval QA | history / profile + query | held-out response / preference | prospective sealing; calibration; transition predictioning | scores why an episode matters and predictions the next transition, sealed |
| TargetSpace-Bench (this work) | target-specific prediction under partial observation | passive multimodal observation up to sealed T | calibrated prediction over target-state transitions | by design: physical realism, control, generation fidelity | is the complementary layer the other families omit |
Built like the benchmarks researchers trust
We borrow the structure and seriousness of established efforts — not their branding.
Challenge & leaderboard ARC-style
A clear mission, a public leaderboard, and explicit benchmark versions — with contamination-resistant, prospective evaluation rather than a static answer key.
Submissions & splits SWE-bench-style
Defined task splits, a submission pipeline, and a verification path so leaderboard credibility rests on reproducibility, not self-report.
Transparent evaluation HELM-style
An explicit evaluation philosophy: scenarios (tracks/splits), multiple metrics reported side by side, and calibration treated as first-class.
Governance & verification MLCommons-style
Versioned rules, official vs unofficial (public/verified/private-eval) submissions, and organizer-run private evaluation for high-stakes claims.