The benchmark for predictive understanding in personal AI

Measure whether AI has learned a useful model of a target.

TargetSpace tests whether a system can turn passive, instrumented longitudinal observation of a specific target — a person, team, or organization — into calibrated predictions of that target's future state transitions, beyond population priors, routine replay, and retrieval. Architecture-neutral: no model class, memory system, or sensor stack is prescribed.

Understanding is the capability. Calibrated prediction is the measurement. The objective is validated lift, not maximal capture. The record is not the person.

Public preprint — manuscript and benchmark release V1.0. Pre-pilot protocol. Synthetic harness only. No human-subject results.

V1.0  pre-pilot protocol release Results  none yet Harness  synthetic only personal track instantiated first
The evaluation gap

Current evaluations miss the target-specific question

Existing evaluations measure retrieval, preference replay, style imitation, or generic prediction. None of them can distinguish those from target-specific predictive skill.

Generic prediction

Predicts what usually happens. Looks calibrated and often correct while knowing nothing about any individual target.

Routine replay

Predicts what this target usually does. Looks personalized while having captured only habit.

Retrieval

Finds facts already stored in memory. Surfaces what happened, not what happens next.

Target-specific prediction

Predicts how this target will transition next, before the outcome exists — the capability TargetSpace isolates and measures.

TargetSpace enters no empty field: routine modeling, behavior forecasting, and personal sensing are decades old. What it standardizes is the conjunction — sealed prospective predictions under own-routine and population baselines, calibration, and wrong-target specificity (the landscape). New to the project? The full argument is a 12-minute read: Why TargetSpace →

How TargetSpace works

Sealed predictions, resolved outcomes, audited skill

1
Observe the target up to time tLongitudinal evidence only; strict walk-forward, no future information.
2
Seal the prediction before the outcome existsTimestamped and hashed (SHA-256); contamination-resistant when seals are committed to an external witness before outcomes exist.
3
Resolve the outcome at time rDeterministic, pre-registered resolution rules over observable evidence.
4
Score with proper scoring rulesLog score (bits) and Brier over the full probability distribution.
5
Compare against R1 and R2Population prior and the target's own routine — skill is measured over both.
6
Check calibration and permutation specificityConfidence must track uncertainty; skill must collapse on the wrong target.
7
Report evidence-tier liftWhat each additional evidence stream buys is measured, never assumed.

The decision rule: a model earns target-specific credit only if it beats the population prior, beats the target's own routine, stays calibrated, and loses skill when its predictions are matched to the wrong target.

A concrete instance. A note-taking app claims memory helps. TargetSpace asks it to predict, before the outcome exists, whether a user will complete, defer, cancel, or replace a recurring commitment. It earns credit only if it beats population rates, beats that user's routine, stays calibrated, and fails when scored against another user — the paper's minimal runnable task, and the shape of every personal-track task.

The validity stack

Each layer removes a way to score well without target-specific skill

1
Prospective sealingRemoves hindsight rationalization and contamination.
adopted
2
Proper scoringScores the full probability distribution, not cherry-picked accuracy.
adopted
3
R1 — population-prior baselineFilters out generic base-rate prediction.
adopted
4
R2 — own-routine baselineFilters out routine replay: skill over R2 is skill the routine does not contain.
anchor
5
Calibration gateFilters out overconfident lucky prediction.
adopted
6
Permutation specificity gateSkill must collapse when predictions are matched to the wrong target.
anchor
7
Evidence-tier ablationMeasures which evidence streams add target-specific information.
adopted

Prospective sealing, proper scoring, calibration, and evidence ablation are adopted from prior work and cited. The contribution is the conjunction, anchored on the own-routine baseline and the permutation specificity gate.

One destination

Every branch of personal AI is climbing the same ladder

Meeting recorders, AI wearables, smart glasses, memory assistants, companions, enterprise copilots, and lifelogging systems ship as separate categories with separate metrics. Each accumulates longitudinal context and advertises the same destination: a persistent model of a particular person. TargetSpace is the shared measurement layer — recording → structured memory → longitudinal evidence → target-specific predictive capability → validated assistance. The destination is a capability, not a representation: a personal world model is one architectural hypothesis; target-specific prospective skill is what gets measured.

Layer 1 · Scientific instrumentation

Can a given evidence stream produce calibrated lift beyond the R1 population-prior and R2 own-routine baselines? Today's hardware suffices — a poor all-day product can still be a good instrument.

Layer 2 · Observation architecture

If Layer 1 validates: how do you acquire faithful, low-burden evidence over months to years? Coverage, energy, missingness, reactivity, and consent dominate — an open research problem, not a product roadmap.

Layer 3 · Interactive personal AI

Assistants, displays, chat, agents — interaction surfaces, downstream of validated capability. The sequence is observe → infer → predict → validate → interact, not observe → interact.

Why this matters to AI

Personal world modeling is a proving ground for advanced machine intelligence

Building a persistent, calibrated model of one real person from partial longitudinal evidence is not sufficient for general intelligence, and we make no such claim. But it demands a difficult conjunction of capabilities that more capable machine intelligence will likely need more broadly — and it is comparatively neglected. TargetSpace supplies an instrument for measuring progress on that conjunction.

Persistent identity

Evidence must bind to the same target over time; the wrong-target gate fails a model whose skill is not about this person.

Long-horizon memory

Banal observations become predictive only when integrated across weeks — the longitudinal arm rewards integration over recency.

Latent-state inference

The scored object is a latent state read only through observable resolutions; a system must infer state it cannot see directly.

Adaptation under drift

The target is non-stationary; skill concentrated in transitions rewards updating over routine replay.

Uncertainty calibration

Strictly proper scoring and the calibration gate require honest probabilities, not confident point guesses.

Specificity & transfer

Skill must beat the person's own routine and collapse on the wrong person; transfer across tasks separates a reusable model from a per-task rule.

A proving ground, not a proof. Passing does not certify mastery of any capability in isolation; the task places simultaneous pressure on all of them, and the failure-mode read-out helps localize which one a system lacks. The connection to world-model research is methodological: JEPA-style systems preserve predictable, task-relevant structure without reconstructing the sensory surface — TargetSpace applies the same intuition at the level of a person, and evaluates such systems rather than competing with them.

What builders can measure

Turn product and architecture decisions into controlled experiments

TargetSpace shows builders what actually improves target-specific prediction: model architecture, memory design, observation streams, calibration, and evidence efficiency. Every decision reduces to one comparison: run the same sealed task with one factor changed, and read the change in gated, target-specific skill (Skill over R2, under the gates) at a disclosed cost — not interaction quality, not offline accuracy. In federated mode you can answer these on your own users' data without exporting raw evidence.

Memory

Does persistent memory improve prospective skill, or only retrieval? Toggle it; read Skill over R2. Memory that lifts R1 but not R2 is retrieval.

Sensors

Does audio, video, location, or physiology add value beyond cheaper evidence? Evidence-tier ablation with cost per tier.

Models

Does a stronger model extract more target-specific signal from the same observations? Architecture comparison at a fixed evidence tier.

Edge AI

Can on-device inference match cloud performance? EE-energy with processing locality in the cost vector, at equal gated skill.

Privacy & retention

Which raw data or representations can be deleted without losing validated skill? The minimal sufficient tier answers by rule.

Product

Does the representation transfer across features, and does better prediction improve downstream assistance? Cross-task transfer and the utility track.

Minimum sufficient observation

The goal is not to record everything

TargetSpace does not reward maximal capture. It rewards the least burdensome governed evidence configuration that preserves validated target-specific skill. Gated skill is expected to saturate while observation cost keeps rising; ties within a pre-registered equivalence margin rank the cheaper, less invasive configuration first. More invasive evidence must earn its place in measured predictive lift — and when it cannot, it loses the comparison by rule, not by exhortation.

Weak skill has two separable causes. An evidence limitation (the signal was never captured) and a model limitation (the signal is present but unused) look identical in a single number and need opposite investments. The model–evidence frontier separates them experimentally — vary evidence with the model fixed, or the model with evidence fixed.

What success and failure would mean. Success would show there is target-specific, dynamic signal beyond routine, that modalities can be ranked, and that personal-model claims can be audited. A clean null would be progress too — it would replace a marketing assertion with a measured bound. The uninformative outcome is the unaudited claim, which is where personal AI is now.

Conceptual framing only; no empirical results. See the hardware & observation-science agenda and the paper.

The record is not the target

Speech, calendar entries, messages, location, tasks, sensors, and self-reports are traces — partial observations of an underlying target-state through particular channels, not the target itself. A calendar entry is not a commitment; a missed reply is not avoidance; a late-night search is not a goal. A system can summarize these traces without modeling the state that generated them. TargetSpace tests whether it can infer enough of the latent target-state to predict future observable transitions — scored only against sealed, externally observable outcomes.

traces speech · calendar location · tasks sensors · self-report latent target-state belief Bₜ sealed prediction P over answers resolution rule deterministic scored outcome log / Brier evidence, not the target the only scored object

Traces are evidence of a latent target-state; only the sealed prediction’s resolved outcome is scored.

The flagship track

The personal track, instantiated first

The first instantiated track evaluates a consenting individual's future state transitions: commitment follow-through, priority maintained vs. displaced, task continuation vs. switch, response behaviour, meeting and event realization, attention shifts. The personal track comes first because personal AI is commercially important and scientifically under-evaluated — the paper claims no cross-domain validation. Here, “personal intelligence” is a bounded, product-facing label for calibrated, target-specific predictive skill about future observable states — not general intelligence, consciousness, or access to inner life.

All tracks & splits

Loading tracks…

Beyond wearables

Wherever longitudinal traces meet resolved outcomes

TargetSpace applies wherever longitudinal observations and resolved outcomes make R1, R2, and permutation tests meaningful.

Assistants & agents

Email, calendar, chat, documents, browsing, and task history.

Enterprise copilots

Organizational memory over trackers, commits, and communications.

Health & care systems

Longitudinal behavioural and physiological traces.

Tutors & learning systems

Student interaction history over weeks and terms.

Robotics & embodied agents

Environments, actions, and configurations observed over time.

Recommenders & attention systems

Attention and behaviour histories; project, market, and process prediction.

These are proposed extension domains, not validated empirical results. Each track is admitted only when it supports a strong own-routine baseline and a meaningful permutation test. Product-category pages are on the industry map.

Claims and boundaries

What the paper claims — and does not

Claimed

  • The TargetSpace benchmark framework, specified in full
  • The personal track, specified as the flagship instantiation
  • A synthetic demonstration harness (sanity check only)
  • The evaluation stack, protocol, and reporting contract

Not claimed

  • Human pilot results (no pilot has run)
  • Cross-domain empirical validation
  • Proof that passive observation improves prediction
  • That attention causes target formation
  • Permission to act on predictions
  • A public raw first-person dataset in the initial release
Pilot plan

The first pilot is harness validation

A small feasibility study of whether the protocol can run: prediction generation, sealing, deterministic resolution, variance and dependence estimation, evidence-tier feasibility, and preliminary signal over R1/R2.

It does not test whether personal intelligence exists, does not validate cross-domain generality, and does not establish population-level effects. Abandonment criteria are pre-registered: no system beats R1 → noise; none beats R2 → target-specific prediction not demonstrated; skill survives permutation → not target-specific.

The staged roadmap: Phase 0 synthetic harness (complete) → Phase 1 feasibility pilot (next) → Phase 2 powered target-specificity study → Phase 3 evidence-tier & model–evidence frontier → Phase 4 cross-task transfer → Phase 5 downstream utility. No phase borrows the license of a later one — full ladder on the research page.

Harness & submissions · Version 1.0

A runnable reference harness, an honest label

The V1.0 release ships the protocol, a minimal synthetic harness, and the submission specification through the public repository and the paper. The synthetic harness verifies schemas, metric computation, R1/R2 baseline execution, calibration checks, permutation controls, and evidence ablation — a smoke test plus reference path, not a validated leaderboard result. Pilot validation is pending.

For a product team, TargetSpace converts the claim “our AI knows the user better over time” into an auditable one: run the same sealed prediction tasks with a memory or context feature enabled and disabled, report Skill over R1 and R2, test calibration, and verify the lift collapses under matched-target permutation. If it does not improve prospective target-specific skill, the feature may still be useful retrieval or personalization — but it has not shown that it models the user’s evolving target state. See the product quickstart.

Run it now

python examples/targetspace_synthetic_demo.py — deterministic, standard-library-only, no human data. See Run the harness for the full command and expected output, the Protocol for the scoring spine, and Schemas for the submission contract.

Real-data submissions

A team with compliant longitudinal passive-observation data can format a submission, run the reference checks, and compare against R1 (population prior) and R2 (own routine) — a claimed improvement must beat R2, and its skill must collapse under target permutation. Consent, privacy, and governance requirements are binding: derived, consented artifacts only; never raw bystander media. Teams building ambient recorders, wearable memory systems, smart glasses, agents, or personal-AI memory layers are especially invited to evaluate against the personal track.

Ethics & governance

The inference object is the privacy boundary

Deleting or never storing raw audio and video is not enough if transcripts, OCR, embeddings, speaker maps, entity graphs, commitments, routines, summaries, and inferred goals persist. TargetSpace treats the derived experiential model — not only the raw signal — as the object requiring governance.

Risks the framework names

  • Bystander consent — people near the system become data subjects
  • Purpose drift — meeting notes become behavioural prediction
  • Deletion ambiguity — erasing media may not erase derived layers
  • Searchability — forgotten remarks become queryable
  • Power asymmetry — memory interrogated by whoever holds power
  • Inference leakage — goals and vulnerabilities exposed by accumulation

Balanced by design

  • Beneficial uses are real: accessibility, memory support, productivity, care coordination, accountability, safety
  • Local processing and open-source implementations are valuable mitigations — but incomplete against bystander consent, purpose drift, and derived-inference risk
  • Consent-first, federated, aggregate-only evaluation; observe-not-intervene during scoring windows
  • Benchmark validity is kept strictly separate from deployment legitimacy
Leaderboard

Skill over the routine, not the crowd

Full leaderboard

The headline signal is skill in bits over the own-routine baseline R2, with calibration and permutation specificity as gates. All rows below are illustrative mock baselines — no empirical results exist yet.

Loading leaderboard…

Paper

The protocol, in full

TargetSpace: Benchmarking Predictive Understanding in Personal AI — a benchmark protocol and validity stack for measuring target-specific predictive lift from longitudinal observation. Public preprint — manuscript and benchmark release V1.0.

@misc{sylvester2026targetspace,
  author={Sylvester, Yuri Andrade},
  title={TargetSpace: Benchmarking Predictive
         Understanding in Personal AI},
  year={2026},
  note={Public preprint --- manuscript and
        benchmark release V1.0},
  url={https://targetspace.org/paper}}
FAQ

Common questions

Loading…

All questions