OUR RESEARCH

From predictive
world models to
grounded reasoning.

We study how machines can learn compressed models of the world, reason and plan over what those models represent, and hold every capability to the standard of held-out, real-world evidence.

NEW — FROM THE LAB

On measurement and honest growth in AI systems.

A four-part series of research articles on evaluating AI capability honestly — two-sided error discipline, compute allocation, compounding without retraining, and what parametric and non-parametric systems owe each other.

The Coin Economy: Seeding Towards Generality

AI evaluation polices its false positives and is structurally blind to its false negatives. I pay my research system in a currency only reality can mint — and the hard part, it turns out, is not stopping the system from printing money; it is catching the mint when it silently refuses to pay.

READ ARTICLE

Knowing What You're Good At Beats Being Curious

Compute allocation across a portfolio of tasks is a layer of intelligence that has been measured piecemeal but never put through a controlled head-to-head. On real forecasting data, a memory of past competence allocated a fixed budget 1.22× better than a uniform policy and 1.67× better than a curiosity policy — beating the bandit algorithms, and the field's repaired form of curiosity, along the way.

READ ARTICLE

Measurable Compounding Invention Towards Generality

We built the skill library that is supposed to make an agent compound without retraining. Measured honestly, it barely compounded — a cheap learned policy did. The depth of a domain's solutions predicts which mechanism wins — and the only headline worth trusting is a slope, not a count.

READ ARTICLE

Mutual Training: What Parametric and Non-Parametric Systems Owe Each Other

Weights and ledgers are argued as rival camps. The scarce resource in both is the same — verification — and in the system we run, the two sides train each other across one verified border.

READ ARTICLE

THE RESEARCH PROGRAM

Eight directions, one architecture of intelligence.

OUR RESEARCH PROCESS

Pre-registered.
Ablated. Held out.

Methodology is the product. Every claimed capability survives the same gauntlet before we believe it ourselves.

Pre-registered predictions

Hypotheses and success criteria are written down before the experiment runs.

Ablation-driven architecture

Every component must justify its place with measured contribution.

Held-out by construction

Skill is graded on domains and compositions the system has never seen, against strong baselines.

Calibrated claims

Confidence is scored with proper scoring rules — in our systems and in our own statements.

FOUNDATIONS

The literature we build on.

Our program stands on a specific lineage of ideas — learned world models, self-supervised predictive architectures, intrinsic motivation, hybrid neurosymbolic systems, and spatial intelligence. These are the load-bearing references.