TargetSpace tests whether a system can turn passive, instrumented longitudinal observation of a specific target — a person, team, or organization — into calibrated predictions of that target's future state transitions, beyond population priors, routine replay, and retrieval. Architecture-neutral: no model class, memory system, or sensor stack is prescribed.
Public preprint — manuscript and benchmark release V1.0. Pre-pilot protocol. Synthetic harness only. No human-subject results.
Existing evaluations measure retrieval, preference replay, style imitation, or generic prediction. None of them can distinguish those from target-specific predictive skill.
Predicts what usually happens. Looks calibrated and often correct while knowing nothing about any individual target.
Predicts what this target usually does. Looks personalized while having captured only habit.
Finds facts already stored in memory. Surfaces what happened, not what happens next.
Predicts how this target will transition next, before the outcome exists — the capability TargetSpace isolates and measures.
TargetSpace enters no empty field: routine modeling, behavior forecasting, and personal sensing are decades old. What it standardizes is the conjunction — sealed prospective predictions under own-routine and population baselines, calibration, and wrong-target specificity (the landscape). New to the project? The full argument is a 12-minute read: Why TargetSpace →
The decision rule: a model earns target-specific credit only if it beats the population prior, beats the target's own routine, stays calibrated, and loses skill when its predictions are matched to the wrong target.
A concrete instance. A note-taking app claims memory helps. TargetSpace asks it to predict, before the outcome exists, whether a user will complete, defer, cancel, or replace a recurring commitment. It earns credit only if it beats population rates, beats that user's routine, stays calibrated, and fails when scored against another user — the paper's minimal runnable task, and the shape of every personal-track task.
Prospective sealing, proper scoring, calibration, and evidence ablation are adopted from prior work and cited. The contribution is the conjunction, anchored on the own-routine baseline and the permutation specificity gate.
Meeting recorders, AI wearables, smart glasses, memory assistants, companions, enterprise copilots, and lifelogging systems ship as separate categories with separate metrics. Each accumulates longitudinal context and advertises the same destination: a persistent model of a particular person. TargetSpace is the shared measurement layer — recording → structured memory → longitudinal evidence → target-specific predictive capability → validated assistance. The destination is a capability, not a representation: a personal world model is one architectural hypothesis; target-specific prospective skill is what gets measured.
Can a given evidence stream produce calibrated lift beyond the R1 population-prior and R2 own-routine baselines? Today's hardware suffices — a poor all-day product can still be a good instrument.
If Layer 1 validates: how do you acquire faithful, low-burden evidence over months to years? Coverage, energy, missingness, reactivity, and consent dominate — an open research problem, not a product roadmap.
Assistants, displays, chat, agents — interaction surfaces, downstream of validated capability. The sequence is observe → infer → predict → validate → interact, not observe → interact.
Building a persistent, calibrated model of one real person from partial longitudinal evidence is not sufficient for general intelligence, and we make no such claim. But it demands a difficult conjunction of capabilities that more capable machine intelligence will likely need more broadly — and it is comparatively neglected. TargetSpace supplies an instrument for measuring progress on that conjunction.
Evidence must bind to the same target over time; the wrong-target gate fails a model whose skill is not about this person.
Banal observations become predictive only when integrated across weeks — the longitudinal arm rewards integration over recency.
The scored object is a latent state read only through observable resolutions; a system must infer state it cannot see directly.
The target is non-stationary; skill concentrated in transitions rewards updating over routine replay.
Strictly proper scoring and the calibration gate require honest probabilities, not confident point guesses.
Skill must beat the person's own routine and collapse on the wrong person; transfer across tasks separates a reusable model from a per-task rule.
A proving ground, not a proof. Passing does not certify mastery of any capability in isolation; the task places simultaneous pressure on all of them, and the failure-mode read-out helps localize which one a system lacks. The connection to world-model research is methodological: JEPA-style systems preserve predictable, task-relevant structure without reconstructing the sensory surface — TargetSpace applies the same intuition at the level of a person, and evaluates such systems rather than competing with them.
TargetSpace shows builders what actually improves target-specific prediction: model architecture, memory design, observation streams, calibration, and evidence efficiency. Every decision reduces to one comparison: run the same sealed task with one factor changed, and read the change in gated, target-specific skill (Skill over R2, under the gates) at a disclosed cost — not interaction quality, not offline accuracy. In federated mode you can answer these on your own users' data without exporting raw evidence.
Does persistent memory improve prospective skill, or only retrieval? Toggle it; read Skill over R2. Memory that lifts R1 but not R2 is retrieval.
Does audio, video, location, or physiology add value beyond cheaper evidence? Evidence-tier ablation with cost per tier.
Does a stronger model extract more target-specific signal from the same observations? Architecture comparison at a fixed evidence tier.
Can on-device inference match cloud performance? EE-energy with processing locality in the cost vector, at equal gated skill.
Which raw data or representations can be deleted without losing validated skill? The minimal sufficient tier answers by rule.
Does the representation transfer across features, and does better prediction improve downstream assistance? Cross-task transfer and the utility track.
TargetSpace does not reward maximal capture. It rewards the least burdensome governed evidence configuration that preserves validated target-specific skill. Gated skill is expected to saturate while observation cost keeps rising; ties within a pre-registered equivalence margin rank the cheaper, less invasive configuration first. More invasive evidence must earn its place in measured predictive lift — and when it cannot, it loses the comparison by rule, not by exhortation.
Weak skill has two separable causes. An evidence limitation (the signal was never captured) and a model limitation (the signal is present but unused) look identical in a single number and need opposite investments. The model–evidence frontier separates them experimentally — vary evidence with the model fixed, or the model with evidence fixed.
What success and failure would mean. Success would show there is target-specific, dynamic signal beyond routine, that modalities can be ranked, and that personal-model claims can be audited. A clean null would be progress too — it would replace a marketing assertion with a measured bound. The uninformative outcome is the unaudited claim, which is where personal AI is now.
Conceptual framing only; no empirical results. See the hardware & observation-science agenda and the paper.
Speech, calendar entries, messages, location, tasks, sensors, and self-reports are traces — partial observations of an underlying target-state through particular channels, not the target itself. A calendar entry is not a commitment; a missed reply is not avoidance; a late-night search is not a goal. A system can summarize these traces without modeling the state that generated them. TargetSpace tests whether it can infer enough of the latent target-state to predict future observable transitions — scored only against sealed, externally observable outcomes.
Traces are evidence of a latent target-state; only the sealed prediction’s resolved outcome is scored.
The first instantiated track evaluates a consenting individual's future state transitions: commitment follow-through, priority maintained vs. displaced, task continuation vs. switch, response behaviour, meeting and event realization, attention shifts. The personal track comes first because personal AI is commercially important and scientifically under-evaluated — the paper claims no cross-domain validation. Here, “personal intelligence” is a bounded, product-facing label for calibrated, target-specific predictive skill about future observable states — not general intelligence, consciousness, or access to inner life.
Loading tracks…
TargetSpace applies wherever longitudinal observations and resolved outcomes make R1, R2, and permutation tests meaningful.
Email, calendar, chat, documents, browsing, and task history.
Organizational memory over trackers, commits, and communications.
Longitudinal behavioural and physiological traces.
Student interaction history over weeks and terms.
Environments, actions, and configurations observed over time.
Attention and behaviour histories; project, market, and process prediction.
These are proposed extension domains, not validated empirical results. Each track is admitted only when it supports a strong own-routine baseline and a meaningful permutation test. Product-category pages are on the industry map.
A small feasibility study of whether the protocol can run: prediction generation, sealing, deterministic resolution, variance and dependence estimation, evidence-tier feasibility, and preliminary signal over R1/R2.
It does not test whether personal intelligence exists, does not validate cross-domain generality, and does not establish population-level effects. Abandonment criteria are pre-registered: no system beats R1 → noise; none beats R2 → target-specific prediction not demonstrated; skill survives permutation → not target-specific.
The staged roadmap: Phase 0 synthetic harness (complete) → Phase 1 feasibility pilot (next) → Phase 2 powered target-specificity study → Phase 3 evidence-tier & model–evidence frontier → Phase 4 cross-task transfer → Phase 5 downstream utility. No phase borrows the license of a later one — full ladder on the research page.
The V1.0 release ships the protocol, a minimal synthetic harness, and the submission specification through the public repository and the paper. The synthetic harness verifies schemas, metric computation, R1/R2 baseline execution, calibration checks, permutation controls, and evidence ablation — a smoke test plus reference path, not a validated leaderboard result. Pilot validation is pending.
For a product team, TargetSpace converts the claim “our AI knows the user better over time” into an auditable one: run the same sealed prediction tasks with a memory or context feature enabled and disabled, report Skill over R1 and R2, test calibration, and verify the lift collapses under matched-target permutation. If it does not improve prospective target-specific skill, the feature may still be useful retrieval or personalization — but it has not shown that it models the user’s evolving target state. See the product quickstart.
python examples/targetspace_synthetic_demo.py — deterministic, standard-library-only, no human data. See Run the harness for the full command and expected output, the Protocol for the scoring spine, and Schemas for the submission contract.
A team with compliant longitudinal passive-observation data can format a submission, run the reference checks, and compare against R1 (population prior) and R2 (own routine) — a claimed improvement must beat R2, and its skill must collapse under target permutation. Consent, privacy, and governance requirements are binding: derived, consented artifacts only; never raw bystander media. Teams building ambient recorders, wearable memory systems, smart glasses, agents, or personal-AI memory layers are especially invited to evaluate against the personal track.
Deleting or never storing raw audio and video is not enough if transcripts, OCR, embeddings, speaker maps, entity graphs, commitments, routines, summaries, and inferred goals persist. TargetSpace treats the derived experiential model — not only the raw signal — as the object requiring governance.
The headline signal is skill in bits over the own-routine baseline R2, with calibration and permutation specificity as gates. All rows below are illustrative mock baselines — no empirical results exist yet.
Loading leaderboard…
TargetSpace: Benchmarking Predictive Understanding in Personal AI — a benchmark protocol and validity stack for measuring target-specific predictive lift from longitudinal observation. Public preprint — manuscript and benchmark release V1.0.
@misc{sylvester2026targetspace,
author={Sylvester, Yuri Andrade},
title={TargetSpace: Benchmarking Predictive
Understanding in Personal AI},
year={2026},
note={Public preprint --- manuscript and
benchmark release V1.0},
url={https://targetspace.org/paper}}