A model’s parameter count need not capture its statistical complexity: distinct parameters may describe the same function, and different solutions may fit the data equally well. I study how this geometry shapes learning and what it permits us to infer, using algebraic geometry, singular learning theory, and algebraic statistics.
Geometry & statistics
What does a model’s structure make possible?
Before asking how a model learns, I characterize the functions it can represent and the information its representations retain.
Algebraic Statistics · 2024
Geometry of polynomial neural networks
For polynomial networks with monomial activations, we study the geometry of representable functions, with explicit semialgebraic descriptions for selected architectures. We obtain dimension results that quantify expressivity and study the learning degree, an algebraic measure of training complexity. Computations and experiments accompany the geometric analysis.
Pull-back geometry of persistent homology encodings
Which changes in a dataset does a topological representation detect? We use the pull-back metric induced by an encoding to measure its sensitivity to input perturbations. This lets us examine which features it captures and compare encoding choices before training a downstream model.
With Shuang Liang, Renata Turkeš, Nina Otter, and Guido Montúfar
How does this structure influence what is learned?
Within a model’s function space, solutions with similar training error can have different statistical complexity. Singular learning theory relates their local geometry to Bayesian learning and generalization.
Preprint · 2026
A basin-selection perspective on grokking
Grokking is the delayed onset of generalization after a network has fit its training data. In shallow quadratic networks, we derive formulas for the local learning coefficient in both lazy and feature-learning regimes. These describe the statistical preference among competing solution basins; the timing of a transition also depends on optimization dynamics.
In experiments, estimates of the learning coefficient track the onset of generalization, connecting the geometric analysis to quantities measurable during training.
With Ben Cullen, Sergio Estan-Ruiz, and Riya Danait, from a project I mentored at LOGML 2025
I am investigating how reliably local learning coefficients can be estimated in neural networks, including variational alternatives to sampling-based methods. This is a step toward using complexity estimates to study changes in learning and representation.
Guarantees
What can the data identify?
A good fit alone does not establish that a model’s parameters are recoverable. Identifiability asks which conclusions the observations support, and under what assumptions.
Preprint · 2026
Identifiability in graphical discrete Lyapunov models
We study parameter recovery from the steady-state distribution of a graphical dynamical model. With non-Gaussian noise, higher-order cumulants provide information beyond covariance.
We establish generic identifiability for directed acyclic graphs with a self-loop at every vertex, and local identifiability for directed graphs with self-loops at every vertex and no isolated vertices.
With Cecilie Olesen Recke, Sarah Lumpp, Nataliia Kushnerchuk, Janike Oldekop, Jane Ivy Coons, and Elina Robeva
Parameter symmetries can give different internal descriptions of the same learned function. I am exploring which interpretability claims remain meaningful across these descriptions, and how to formulate such claims using invariant quantities. This question motivates my developing work in AI alignment.