Back to Home

Influential Papers

A curated collection of academic papers and articles that influence my research and thinking across various domains.

Language Models Use Trigonometry to Do Addition
Interpretability

Language Models Use Trigonometry to Do Addition

Subhash Kantamneni, Max Tegmark

The title is all you need! Super cool!

Prompting as Scientific Inquiry
Positions and Visions

Prompting as Scientific Inquiry

Ari Holtzman, Chenhao Tan

Makes a really interesting case for disambiguating prompt 'engineering' from a possible 'prompt science'. The writing is very compelling. A helpful analogy presented in the paper is the idea that plant breeders were able to infer a lot of the internal structure of plants before genetic theory explained it. Not sure where I stand on this still but it's given me a lot to think about, especially w.r.t. 'aesthetic' concerns in ML research that bar 'prompting' from being seen as a legitimate research method. I do agree with the paper's claim that many important works in NLP are basically interfaces/structures upon prompting, and we shouldn't be afraid to more closely associate them with a 'prompt science'.

AI as Governance
Positions and Visions

AI as Governance

Henry Farrell

Really useful description of AI as a governance system, akin to markets, democracy, and bureaucracy -- a lens that lets political scientists to usefully contribute towards thinking on AI.

AI and the Demise of College Writing
Positions and Visions

AI and the Demise of College Writing

Adam Walker

Advocates for rhetoric over composition as the methodology for writing pedagogy in the AI era.

Why Chatbots Are Not the Future
Human-AI Interaction

Why Chatbots Are Not the Future

Amelia Wattenberger

Really nice argumentative piece on why we can build much better AI interfaces than chat interfaces.

Tools for Conviviality
Philosophy

Tools for Conviviality

Ivan Illich

An ambitious yet informed vision for what it would mean and cost for us to have *convivial* tools -- tools that we can make and shape our own lives for joy and purpose.

Reflections on Qualitative Research
Interpretability

Reflections on Qualitative Research

Chris Olah, Adam Jermyn

Interesting thoughts on what kind of methodology suits interpretability as a growing, immature field.

Nietzsche and the Virtues of Mature Egoism
Philosophy

Nietzsche and the Virtues of Mature Egoism

Christine Swanton

Swanton reads Nietzsche's 'immoralism' and 'egoism' as articulating virtues of the 'mature egoist' -- someone who has overcome immature egoism (ressentiment, self-deception, reactive self-assertion) in favor of genuine self-affirmation and creative engagement with the world. Setting aside questions of fidelity to Nietzsche's texts, I find this a compelling moral outlook: a rejection of both slavish self-denial and petty self-aggrandizement in favor of something like joyful, active, honest self-cultivation.

Deep Learning is Not So Mysterious or Different
Representation Learning

Deep Learning is Not So Mysterious or Different

Andrew Gordon Wilson

Super interesting and illuminating perspective explaining why supposedly deep-learning-unique phenomena like deep double descent, overparametrization, etc. can be explained using soft inductive biases and existing generalization frameworks. The references are a treasure trove!

Toward cultural interpretability: A linguistic anthropological framework for describing and evaluating large language models
Positions and Visions

Toward cultural interpretability: A linguistic anthropological framework for describing and evaluating large language models

Graham M Jones, Shai Satran, Arvind Satyanarayan

Advocates for understanding LLM behavior as indicative of nuances in human social behavior.

The Shape of Math To Come
Philosophy and History of Math

The Shape of Math To Come

Alex Kontorovich

Kontorovich reflects on how AI and formal verification systems like Lean are reshaping mathematical practice, intended for ICM 2026. What I find compelling is that the examples come from someone deeply embedded in both traditional research mathematics and these new technologies. The discussion of the 'Bitter Lesson' applied to mathematics is sobering, and the concrete examples of what formal verification looks like in practice are clarifying. A useful document for thinking about what 'doing math' might mean going forward.

Cognitive Behaviors that Enable Self-Improving Reasoners
Representation Learning

Cognitive Behaviors that Enable Self-Improving Reasoners

Kanishk Gandhi, Ayush Chakravarthy, Anikait Singh, Nathan Lile, Noah D. Goodman

Identifies four cognitive behaviors (verification, backtracking, subgoal setting, backward chaining) that predict whether a model can self-improve via RL. The key finding is striking: it's the presence of reasoning behaviors, not answer correctness, that matters. Models exposed to training data with proper reasoning patterns -- even incorrect answers -- matched the improvement of models that had these behaviors naturally. A useful framing for thinking about what 'reasoning' actually is in these systems.

Discovering Latent Knowledge in Language Models Without Supervision
Interpretability

Discovering Latent Knowledge in Language Models Without Supervision

Collin Burns, Haotian Ye, Dan Klein, Jacob Steinhardt

A method to probe structure in language models without any notion of ground truth, relying instead on the consistency property of tru statements.

HCI for AGI
Human-AI Interaction

HCI for AGI

Meredith Ringel Morris

Useful outline of what HCI researchers can contribute to 'AGI'. It's not obvious (and people may fear that) interaction problems will be solved by AGI. Perhaps not?

Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions
Human-AI Interaction

Towards Bidirectional Human-AI Alignment: A Systematic Review for Clarifications, Framework, and Future Directions

Hua Shen, Tiffany Knearem, Reshmi Ghosh, Kenan Alkiek, Kundan Krishna, et al.

A recognition that not only must AI 'align to' human (values, behavior, knowledge, etc.) (-- whatever this means), but we also need to think about how humans might 'align' to AI by working with AI-structured systems. This paper recognizes social 'looping effects' brought about by AI and its behavior.

Unsupervised Elicitation of Language Models
Open-ended Modeling

Unsupervised Elicitation of Language Models

Jiaxin Wen, Zachary Ankner, Arushi Somani, Peter Hase, Samuel Marks, et al.

Interesting way to automatically label datasets using mutual predictability and logical consistency.

We Can't Understand AI Using our Existing Vocabulary
Positions and Visions

We Can't Understand AI Using our Existing Vocabulary

John Hewitt, Robert Geirhos, Been Kim

A compelling articulation of what human-AI communication could look like. Proposes neologism learning.

Jury Learning: Integrating Dissenting Voices into Machine Learning Models
Concept-structured AI

Jury Learning: Integrating Dissenting Voices into Machine Learning Models

Mitchell L. Gordon, Michelle S. Lam, Joon Sung Park, Kayur Patel, Jeffrey T. Hancock, Tatsunori Hashimoto, Michael S. Bernstein

By modeling individual views rather than an aggregated 'view', we can explicitly define the voices 'heard' in making a decision and consider counterfactuals.

Loading more articles...