The methodology paper
First time here? Start with the editorial introduction: Why TargetSpace → · How this relates to adjacent benchmarks: Related work →
TargetSpace: Benchmarking Predictive Understanding in Personal AI — a benchmark protocol and validity stack for measuring target-specific predictive lift from longitudinal observation.
Abstract
Personal AI systems increasingly claim to remember users, personalize, act as agents, and model the people they serve, and the same target-modeling claim now appears in enterprise software, healthcare, robotics, vehicles, and other settings where an AI system observes a target over time. Existing benchmarks measure general knowledge, static tasks, conversational quality, tool use, code, or retrospective retrieval; none is designed to test whether a system has learned a useful model of a specific target over time. TargetSpace is a benchmark protocol for that question, with personal AI as its flagship track and a formulation designed to generalize across longitudinally observed targets. Its anchor metric is prospective state-transition prediction: a system observes a consenting target — a person, team, organization, patient, project, device, or environment — up to a sealed time t and commits calibrated probabilistic predictions of the target's future state transitions before outcomes exist; outcomes resolve by deterministic rules and are scored with strictly proper rules. A validity stack determines whether the prediction is meaningful: skill must exceed a population prior (R1), the target's own routine (R2), and retrieval-and-summarization over the same evidence; remain calibrated; collapse under wrong-target substitution; concentrate on regime shifts rather than routine continuation; and attribute its lift to disclosed evidence streams. An observation-science layer prices that evidence, reporting predictive lift against volume, sensitivity, collection burden, energy, latency, consent burden, and privacy risk, with minimum sufficient observation, not maximal capture, as the design objective. Understanding is not reducible to prediction; TargetSpace operationalizes one measurable component of it, predictive understanding — whether a system forms a calibrated, target-specific model with predictive force over time — and surrounds that observable with the controls that make it resistant to shallow correlation. Core, Controlled, and Full deployment levels fix what a reported evaluation must include, from the smallest valid configuration to the complete apparatus. This is a pre-pilot protocol with a synthetic reference harness; no human-subject results are reported, and a high score supports only the bounded protocol claim: calibrated, target-specific predictive force over observable future states.
Positioning and using TargetSpace
A summary of the paper's positioning section.
A complementary, under-measured axis
TargetSpace addresses a distinct, under-measured axis — target-specific prediction under partial observation — that physical-plausibility, generation-fidelity, and control-success benchmarks do not, by construction, score.
Compared on shared axes only
Physical reasoning, video generation, embodied robotics, symbolic forecasting, and agent memory are complementary families. TargetSpace is compared with each only where they overlap — see the comparison table.
A battery of controls
Apply it as a protocol: zero/short/longitudinal history, shuffled-history and wrong-target controls, modality ablation, oracle and human anchors, calibration, and evidence attribution. See Baselines.
Positioning statement
TargetSpace evaluates a missing layer — whether a model can transform passive longitudinal observation into a target-specific predictive model — central for agents that reason about particular people, teams, systems, or environments over time.
Cite
@misc{sylvester2026targetspace,
author = {Sylvester, Yuri Andrade},
title = {TargetSpace: Benchmarking Predictive Understanding in Personal AI},
year = {2026},
note = {Public preprint --- manuscript and benchmark release V1.0},
url = {https://targetspace.org/paper}
}
Version history
Versioning model: Release V1.0 names the joint manuscript-and-benchmark release (paper, protocol, schemas, harness). Earlier manuscript revisions (v1.0–v1.2) predate that release and are retained below as history.