Raj et al. · arXiv preprint · 2026
Forty-one failure modes, each pinned to the interaction that produced it and the side responsible for fixing it. Turns 'the agent failed' into a repair assignment.
Papers (and one essay) I suggest to anyone getting into data science, across machine learning, statistics, systems, and the occasional heresy, plus whatever I am currently reading for my own work. Titles link out to the source.
Raj et al. · arXiv preprint · 2026
Forty-one failure modes, each pinned to the interaction that produced it and the side responsible for fixing it. Turns 'the agent failed' into a repair assignment.
Vasundra Srinivasan · arXiv preprint · 2026
Across three agent benchmarks the agent itself explains under 3% of the variance, while agent-by-task interaction explains 7 to 23%. Leaderboards rank specialization, not capability.
Enomoto et al. · arXiv preprint · 2026
232 hours of end-to-end benchmarking, replaced by a proxy metric that needs neither a browser nor a model and still predicts the outcome. The cheap measurement is the contribution.
Debjyoti Paul · arXiv preprint · 2026
Freeze the model and make prompt assembly the thing you tune. A control-theory case for why the scaffolding around the model is where the engineering actually lives.
Tom Zahavy · ICML · 2026
The abductive leap, inventing the premises rather than deriving the proof, is the one move LLMs still can't make. Read it right after The Bitter Lesson and let the two argue.
Rich Sutton · essay · 2019
1,100 words explaining seventy years of AI research regret: general methods plus compute beat human cleverness, every time.
Claude Shannon · Bell System Technical Journal · 1948
One paper invents the bit, entropy, and the ceiling on every channel. Everything else on this list is measured in its units.
Brown et al. · NeurIPS · 2020
Scale as a research result in its own right, and the paper where the current era of AI actually begins.
Krizhevsky, Sutskever & Hinton · NeurIPS · 2012
The starting gun for modern deep learning. Eight pages, one GPU trick, and a decade of consequences.
John Ioannidis · PLoS Medicine · 2005
The cheapest inoculation available against taking p-values at face value: uncomfortable, famous, and worth the discomfort.
Hadley Wickham · Journal of Statistical Software · 2014
The paper behind why every dataframe you've ever liked felt likeable. You already follow its rules; this is where they come from.
Halevy, Norvig & Pereira · IEEE Intelligent Systems · 2009
Why more data beats a cleverer model more often than anyone's pride would like. Four pages that predicted the next fifteen years.
Sculley et al. · NeurIPS · 2015
The model is the smallest box in the diagram; everything around it is the actual job. Read before your first ML role, re-read during it.
Leo Breiman · Statistical Science · 2001
The stats-versus-ML worldview split, named in 2001, and still the argument underneath every modeling debate you'll ever be in.