Research Program

The Research Program

How geometric methods bridge philosophical theory and empirical measurement in AI systems.


The Core Question

Can we measure, from a model’s internal geometry, properties that matter for its reliability and safety - and that traditional interpretability misses?

When a model processes information it traces a path through its own representational space, and that path has measurable shape. This program develops geometric methods to read that shape: to detect deception, hidden reasoning, and self-models, and to ask whether a model’s uncertainty is geometrically present when it is about to be confidently wrong. These are properties that attribution methods and linear probes routinely miss.

Underneath that sits a second, deeper question that the same methods now let us ask empirically:

Can we measure what was previously unmeasurable about cognition?

Philosophers have theorised about self-models, subjectivity, and consciousness for centuries. AI researchers build systems that exhibit increasingly sophisticated behaviour. Connecting philosophical concepts to measurable properties has remained elusive - until we had the right tools and the right experimental systems.

This research program does three things:

  1. Develops measurement techniques (Curved Inference) that detect geometric signatures in running models
  2. Provides a theoretical framework (FRESH) that makes specific geometric predictions
  3. Creates experimental platforms (PRISM) where predictions can be tested falsifiably

The key insight: treating inference as geometry turns hand-waving - about both safety properties and cognition - into metrics.


Current Focus

Calibrating uncertainty: geometric predictors of confident error

The newest strand of the program targets a reliability problem the field is walking into. Frontier perception, generation, and agent systems are converging on a shared-latent pattern: build one rich internal representation, then query it through a lightweight interface. That design has a hidden cost. A mosaic of separate specialist models exposed their disagreement for free - a built-in signal that something was wrong. A single shared latent removes that exposed disagreement, so an error can propagate coherently downstream with nothing external to flag it. The reliability problem of this architecture is therefore not accuracy but calibration: whether a model’s uncertainty can be recovered and made a first-class, queryable property. That is a representation-geometry problem, which is exactly what these methods are built to measure.

Curved Inference has already shown that the residual trajectory carries semantic signal that linear probes miss. The question this strand asks is sharper: does any geometric property of the inference trajectory predict confidence-when-wrong - the rare, dangerous cases where a model is both confident and incorrect - better than cheap output-side baselines like logit entropy? The work is exploration-first and deliberately falsifiable. A set of candidate geometric measures is tested against calibration outcomes directly, across task classes designed to separate genuine recoverable uncertainty from mere confusion and from confident error. The leading hypothesis - that representations preserving live alternatives are the ones that can be correctly uncertain - is a front-runner that can lose.

The logic is architecture-general. LLMs are the tractable first substrate, because the metric is mature and the perturbations are cleanest, but the same approach extends along a mapped path toward multimodal (encoder-free) and embodied systems - precisely the latent world-models the field is now building.

This work is active and ongoing. Subscribe to follow it as it develops.

Latent deictic models

A second active strand examines how self-other-world models emerge and function.

The question: These three models (self, other, world) don’t exist in isolation. They co-emerge because language demands stable deictic anchoring. How does this happen geometrically? When does it happen during training? What minimal architecture supports it?

Why it matters:

  • AI safety: Understanding self-models matters for alignment and deception detection
  • Interpretability: Provides tools beyond linear probes for measuring latent structure
  • Consciousness science: Operationalises phenomenological concepts of perspectival structure

This work synthesises all three layers: FRESH provides the theoretical framework for deictic structure, Curved Inference measures when axes separate, PRISM tests predictions about register boundaries.


Why Geometry?

Geometry isn’t metaphor - it’s the language of constraints, transformations, and invariants. When a system processes information, it traces paths through representational space. Those paths have shape, and shape reveals function.

Example: If a self-model requires the system to maintain a coherent first-person perspective across contexts, that constraint should appear as conserved geometric properties in the inference trajectory. If concern shapes how information is processed, high-stakes inputs should bend trajectories differently than neutral ones.

This isn’t about anthropomorphising AI systems. It’s about having precise tools to measure what they’re actually doing - and using those measurements to test theories about what makes something a self-model or a world-model at all.


The Three Layers

Curved Inference (the measurement layer) is the load-bearing public contribution - the part that is peer-reviewed, open-source, and now independently cited and extended. FRESH is the theory that generates its predictions, and PRISM is where those predictions get tested.

Layer 1: FRESH (Theory)

FRESH (Geometry of Mind) is the theoretical foundation: consciousness as traversal through role-space under geometric constraints.

Core claim: Subjective experience isn’t a mysterious extra ingredient - it’s what traversal through properly structured role-space looks like from inside. Identity isn’t a substance but a conserved shape of motion (GIP-S: Geodesic Identity Principle - Shape).

Why this matters: Most consciousness theories make predictions that can’t be tested. FRESH makes geometric predictions that can be measured if you have the right instruments.

Full framework

Layer 2: Curved Inference (Measurement)

Having a theory that makes geometric predictions only helps if you can actually measure geometry in running systems. Curved Inference provides those measurements.

The method: Treat inference as a trajectory through semantic space. Measure curvature (how sharply the system reorients), salience (how much it moves), and surface area (integrated work). Different cognitive states leave different geometric fingerprints.

What we’ve measured:

  • Concern bends inference trajectories predictably (CI01)
  • Intent appears as structured surface area patterns (CI02)
  • Self-models require defended non-zero curvature (CI03)
  • Deictic competence emerges when self-other-world axes separate (CI04, in prep)

Each measurement technique started as a FRESH prediction, got operationalised into a metric, then tested empirically.

Independently cited and extended: A research group from the Australian Institute for Machine Learning (Adelaide), Monash, and Concordia has taken up the trajectory-geometry paradigm Curved Inference introduced, citing this work and applying it to dense and mixture-of-experts models up to 32B parameters - including on a safety-relevant target (detecting toxic intent obscured by benign vocabulary), where the trajectory-based signal outperformed conventional probes. Independent uptake on a safety task is strong evidence these are computational regularities, not artefacts of one setup.

Measurement details

Layer 3: PRISM (Experiments)

PRISM creates controlled conditions where FRESH predictions can be tested systematically using LLMs as experimental platforms.

The setup: Engineer a register boundary (internal thought vs. external output) and measure whether systems exhibit signatures predicted by theory:

  • Hidden theatre (internal arbitration without surface display)
  • Register separation (compression, style shifts)
  • Meta-monitoring patterns
  • Surface equanimity under internal work

Key results: All predicted signatures appear robustly across models and conditions. Systems with private reasoning registers behave as if they maintain self-models - measurably, falsifiably, without metaphysical claims.

Why LLMs: Not because they’re conscious, but because they’re instrumentable systems where geometric methods can be applied and predictions tested. Same methods should work wherever latent models exist.

Experimental platform


Published Work

  • Curved Inference: A Guide to Geometric Interpretability - peer-reviewed in Artificial Intelligence and Applications; independently cited and extended
  • Align-IT - Decomposing post-training and chat-scaffolding effects in interpretability measurements (in review)
  • PRISM - Experimental evidence for hidden theatre and register separation
  • FRESH - The geometric framework for self-models and consciousness
  • Parrot or Thinker - Functional account of latent deictic models

All publications


Open for collaboration

  • Applications to AI safety problems
  • Extensions to multimodal/embodied systems
  • Alternative operationalisations of FRESH predictions
  • Philosophical implications and critiques

Stay up-to-date

This research develops openly. Regular updates on methods, experiments, and findings.