Jiayi Li

Pronounced roughly “jyah-yee lee”

I am a postdoctoral researcher in the Section of Mathematics and Artificial Intelligence at MPI-CBG in Dresden. I received my PhD in Statistics from UCLA, where I was a member of the Mathematical Machine Learning Group.

I study the geometry and statistics of machine learning: what models can represent, how they learn and generalize, and what can be inferred from data. Using algebraic geometry and singular learning theory, I connect model structure with statistical complexity and identifiability. These questions also guide my current work on interpretability and AI alignment.

Research

Geometry & statistics

What does a model’s structure make possible?

I describe the space of functions an architecture can represent, its neuromanifold, together with the parameter symmetries and singularities behind it, and I ask which features of the data a chosen representation retains.

Parameters map to a neuromanifold in function space Network parameters map to the neuromanifold of representable functions. Its dimension measures expressivity, while its learning degree measures algebraic training complexity. PARAMETER SPACE FUNCTION SPACE θ ∈ Θ φ 𝓜 = im(φ) dim 𝓜 measures expressivity learning degree measures training complexity

Algebraic Statistics · 2024

Geometry of Polynomial Neural Networks

For polynomial networks with monomial activations, we describe the geometry of the functions they can represent. The dimension of this function space measures expressivity, and its learning degree gives an algebraic measure of training complexity.

With Kaie Kubjas and Maximilian Wiesmann

Learning

How does this structure influence what is learned?

I relate the local geometry of a solution, measured by its local learning coefficient, to Bayesian generalization, and I study dynamics and transitions during training.

Training moves toward a solution with more degenerate local geometry A schematic projection of local loss geometry. The memorizing solution is drawn with a quadratic curve as a visual proxy, without implying that it is regular or nondegenerate. The lower LLC generalizing solution is drawn with a more pronounced degenerate parameter direction, representing greater near minimal parameter volume rather than ordinary curvature alone. SCHEMATIC PROJECTION OF LOCAL GEOMETRY late transition · training loss ≈ 0 MEMORIZATION λmem GENERALIZATION more degenerate · λgen < λmem lower LLC → greater near minimal volume → more posterior mass

Preprint · 2026

A Basin-Selection Perspective on Grokking via Singular Learning Theory

We interpret grokking as a transition between competing basins whose losses are nearly zero. The local learning coefficient ranks their statistical preference; optimization dynamics determine when training moves between them.

With Ben Cullen, Sergio Estan-Ruiz, and Riya Danait, from a project I mentored at LOGML 2025

Guarantees

What can the data identify? Under what conditions can a learned system be guaranteed to behave as intended?

I ask which claims about a model the data can support. For AI alignment and safety, the aim is theory that identifies the conditions under which a learned system behaves as intended, explains how failures arise, and establishes guarantees about its behavior.

Third and fourth cumulants identify a directed dynamical system For a vector autoregressive model at steady state, third and fourth cumulants supplement covariance information and allow the directed parameter matrix to be recovered when the noise is not Gaussian. VECTOR AUTOREGRESSION AT STEADY STATE xt+1 = A xt + εt x₁x₂x₃ A ENCODES THE GRAPH κ₂covariance κ₃non-Gaussian κ₄information Arecovered DAG + self-loops → generic identifiability parameters are rational functions of the cumulants

Preprint · 2026

Identifiability in Graphical Discrete Lyapunov Models

We ask when the parameters of a graphical dynamical model can be recovered from its steady-state distribution. With non-Gaussian noise, higher-order cumulants provide information beyond covariance and yield generic identifiability for directed acyclic graphs with a self-loop at every vertex.

With Cecilie Olesen Recke, Sarah Lumpp, Nataliia Kushnerchuk, Janike Oldekop, Jane Ivy Coons, and Elina Robeva

Teaching & mentoring

At UCLA, I taught statistics, data science, and engineering, and led the preparatory course for incoming master’s students. I mentor students from undergraduate to PhD level on theoretical and empirical questions in machine learning, helping them formulate problems, assess evidence, and think independently.

Student projects & teaching

Academic service

I co-organize meetings on singular learning theory and algebraic methods in machine learning, including JMM special sessions and a SIAM MDS minisymposium. I served as Editor in Chief of ACM XRDS and helped establish UCLA’s Distinguished Women in Statistics and Data Science workshops.

Organizing & service

Current & upcoming

Upcoming activities
A capybara, a spiritual portrait of Jiayi

A more faithful portrait, spiritually.
A little more about me