Independent research project

Formally verified autoresearch for theoretical mechanistic interpretability

LLM agents collaborate to discover and formally verify theorems about the internal computations of transformers.

Verified Mechanisms is an independent, grant-funded research project developing theoretical foundations for mechanistic interpretability.

We aim to produce theoretical and foundational results in mechanistic interpretability by leveraging the recent advances in the mathematical abilities of frontier models. To make sure that the theorems provided by the frontier models do not contain logical flaws, we formally verify them in Lean 4.

Preliminary analyses suggest that multi-head attention may admit a powerful mathematical abstraction, akin to the role of Feynman diagrams in particle physics and twistors in the study of cosmological correlators. Such an abstraction could open avenues for knowledge discovery in mechanistic interpretability, providing a foundation for tracking features across the layers of a transformer, eventually informing the development of tools for feature discovery in transformers.

We start with a simple pilot question: how many attention heads are required to represent a Boolean function? The setting is simple enough for frontier models to make progress while requiring a human expert to ask the high level questions.

Autoresearch

An autoresearch pipeline, consisting of multiple frontier models collaborating to generate theorems and formally verify them, with the harness tailored to the problem at hand.

Formal verification

Frontier models make mistakes, long derivations accumulate subtle errors, and human evaluation is expensive. Requiring every proof to pass the Lean kernel removes these failure modes.

Empirical grounding

The research questions are selected to address concrete mathematical problems arising in empirical mechanistic interpretability. For instance, the pilot problem emerged from work investigating how features are represented after an attention update.

A precedent for foundational theory shaping mechanistic interpretability

Theoretical work in mechanistic interpretability (Elhage et al., 2021; Elhage et al., 2022) has shaped empirical interpretability by providing useful abstractions, and principled starting points for experimentation. However, developing this kind of theory has been demanding, requiring sustained expert effort to work through the underlying mathematics. Yet recent advances in language models for mathematical reasoning and formal theorem proving may make it possible to produce rigorous results at a lower cost.

Our aim is more modest than that of Elhage et al. (2021) but similar in spirit: to come up with mathematical abstractions that inform empirical mechanistic interpretability.

Mech Interp Dojo: a growing body of machine-verified theorems about the internal mechanisms of LLMs.

Mechanistic interpretability

Our aim is a growing library of kernel-checked theorems that applied researchers can use as reference points, produced faster than hand-derivation alone. We start from attention heads and work towards composing component level theorems into the expressivity of multilayer transformers.

Harness design

An open source multi agent orchestration harness: a system for coordinating frontier models for conjecture generation, proof search and Lean verification.

We have found that several frontier models collaborating produce novel theorems that none of them reaches alone. We want to understand which orchestration choices unlock mathematical progress, and to publish the harness together with evaluations across different choices of harness.

Researchers in mechanistic interpretability and Lean formalization.

The current team consists of 4 researchers with expertise in mechanistic interpretability and Lean formalization. We plan to onboard two additional researchers over the coming months to help with the multi agent orchestration and mechanistic interpretability aspects of this project.

Karthik Viswanathan

Research lead

Sets the research direction, does the mathematics alongside the agents, and runs the autoresearch pipeline. Has been a part of MARS 3.0 and MATS 9, and has published papers on interpretability in ICML and TMLR.

Collaborators

Berkeley Research Academy

Maintain the repository with the formally verified proofs in Lean 4. They also guide the autoresearch pipeline by suggesting the mathematical abstractions that can be used to characterize multihead attention.

Help us build formally verified autoresearch for mechanistic interpretability.

We are preparing to recruit two researchers to collaborate with our autoresearch framework on a three-month project (September to November 2026): a research scientist for mechanistic interpretability theory, and a research engineer for harness design who will set up the autoresearch experiments.