A model’s parameter count need not capture its statistical complexity: distinct parameters may describe the same function, and different solutions may fit the data equally well. I study how this geometry shapes learning and what it permits us to infer, using algebraic geometry, singular learning theory, and algebraic statistics.

Geometry & statistics

What does a model’s structure make possible?

Before asking how a model learns, I characterize the functions it can represent and the information its representations retain.

Algebraic Statistics · 2024

Geometry of polynomial neural networks

For polynomial networks with monomial activations, we study the geometry of representable functions, with explicit semialgebraic descriptions for selected architectures. We obtain dimension results that quantify expressivity and study the learning degree, an algebraic measure of training complexity. Computations and experiments accompany the geometric analysis.

With Kaie Kubjas and Maximilian Wiesmann

PaperCode & computations

Transactions on Machine Learning Research · 2024

Pull-back geometry of persistent homology encodings

Which changes in a dataset does a topological representation detect? We use the pull-back metric induced by an encoding to measure its sensitivity to input perturbations. This lets us examine which features it captures and compare encoding choices before training a downstream model.

With Shuang Liang, Renata Turkeš, Nina Otter, and Guido Montúfar

PaperCode & data

Learning

How does this structure influence what is learned?

Within a model’s function space, solutions with similar training error can have different statistical complexity. Singular learning theory relates their local geometry to Bayesian learning and generalization.

Preprint · 2026

A basin-selection perspective on grokking

Grokking is the delayed onset of generalization after a network has fit its training data. In shallow quadratic networks, we derive formulas for the local learning coefficient in both lazy and feature-learning regimes. These describe the statistical preference among competing solution basins; the timing of a transition also depends on optimization dynamics.

In experiments, estimates of the learning coefficient track the onset of generalization, connecting the geometric analysis to quantities measurable during training.

With Ben Cullen, Sergio Estan-Ruiz, and Riya Danait, from a project I mentored at LOGML 2025

Paper

Ongoing: estimating local complexity

I am investigating how reliably local learning coefficients can be estimated in neural networks, including variational alternatives to sampling-based methods. This is a step toward using complexity estimates to study changes in learning and representation.

Guarantees

What can the data identify?

A good fit alone does not establish that a model’s parameters are recoverable. Identifiability asks which conclusions the observations support, and under what assumptions.

Preprint · 2026

Identifiability in graphical discrete Lyapunov models

We study parameter recovery from the steady-state distribution of a graphical dynamical model. With non-Gaussian noise, higher-order cumulants provide information beyond covariance.

We establish generic identifiability for directed acyclic graphs with a self-loop at every vertex, and local identifiability for directed graphs with self-loops at every vertex and no isolated vertices.

With Cecilie Olesen Recke, Sarah Lumpp, Nataliia Kushnerchuk, Janike Oldekop, Jane Ivy Coons, and Elina Robeva

Paper

Ongoing: identifiability and interpretation

Parameter symmetries can give different internal descriptions of the same learned function. I am exploring which interpretability claims remain meaningful across these descriptions, and how to formulate such claims using invariant quantities. This question motivates my developing work in AI alignment.