Mini Papers¶
Reproductions of key research papers using nnsight.
- Function Vectors - Todd et al. on function vectors
- Geometry of Truth - Marks et al. on truth representations
- LLM Depth - Csordás et al. on LLM computational depth
- Dual Route Induction - Feucht et al. on induction heads
- Demystifying Memorization - Huang et al. on memorization
- Jacobian Lens - Gurnee et al. on reading the residual stream through its causal effect on the output
- Patchscopes - Ghandeharioun et al. on decoding hidden states with the model itself
- IOI Path Patching - Wang et al. on the indirect object identification circuit, at the edge level
- Progress Measures for Grokking - Nanda et al. on the circuit that forms during grokking, trained from scratch in the notebook
- Emergent World Models - Li et al. and Nanda et al. on the board Othello-GPT keeps in its head
- Concept Attention - Helbling et al. on per-concept heatmaps from a diffusion transformer's own attention
- The Logit Lens Over Image Tokens - Neo et al. on reading a VLM's image patches as language
- Base-10 Arithmetic Behind Cyclic Reasoning - Feucht et al. on how Llama does month and weekday arithmetic in base 10
- Gaze Heads - Gandikota & Bau on the attention heads a VLM looks through, and steering what it describes
- Massive Activations - Sun et al. on the handful of residual-stream dimensions that carry enormous magnitude, and what breaks without them
- Induction Heads and the ICL Phase Change - Olsson et al. on induction heads forming in a single interval of training, and the in-context learning jump that accompanies it
- Under-trained Tokens - Land & Bartolo on finding the tokens training never touched, from the unembedding alone
- Chain-of-Thought Faithfulness - Lanham et al. on whether a reasoning model's stated reasoning actually drives its answer
- Belief Lookbacks - Prakash et al. on how a model tracks a character's belief when it diverges from reality, reproduced at 8B