1 The Probable Error of a Mean | William S. Gosset (as “Student”) | 1908 | Statistics / Inference | Derives the t-distribution so you can do inference with tiny samples instead of needing a mountain of data. | Any time I’m A/B testing with laughably small n, this is the ghost in the machine bailing me out. | The whole idea that you can get sane uncertainty estimates from like a dozen observations still feels like cheating. | Medium if you’ve survived one real stats class; otherwise bring coffee. | Biometrika archives, most stats textbooks, or free scans online. | It’s the OG “do more with less data” paper—basically Moneyball for sample sizes. |
|---|
2 Regression Discontinuity Designs in Economics | David S. Lee and Thomas Lemieux | 2010 | Econometrics / Causal inference | A tour of regression discontinuity designs and how to treat arbitrary cutoffs like nature’s randomized trials. | If you work with messy policy data, this is the closest you get to sorcery without p-hacking. | Those jump-at-the-cutoff plots where one side of the line looks like it’s living in a different universe. | Medium-hard; the intuition lands fast, the assumptions take a few re-reads. | Journal of Economic Literature; also floating around as a PDF from various universities. | This is the paper that made me see every score cutoff—test scores, credit scores, you name it—as a potential natural experiment. |
|---|
3 The Use of Multiple Measurements in Taxonomic Problems | R. A. Fisher | 1936 | Statistics / Pattern recognition | Introduces linear discriminant analysis using iris flowers, basically the grandparent of half the classifiers we still use. | It’s a reminder that a simple linear boundary, done right, can hang with the fancy deep models on the right problem. | The separation of iris species in low-dimensional space—still wild how clean it looks for a real dataset. | Medium; algebra-heavy but conceptually pretty clean. | Annals of Eugenics archives or any ML history reader; PDFs are everywhere. | Feels like watching black-and-white game tape of a legend and realizing the fundamentals still work in today’s league. |
|---|
4 Random Forests | Leo Breiman | 2001 | Machine learning / Ensemble methods | Shows how averaging lots of noisy decision trees with randomness baked in gives you a shockingly strong predictor. | Any time I need a baseline model that punches above its weight, this is still the first jersey off the bench. | The error vs. number-of-trees plots that flatten out like a good defensive rotation—diminishing returns, but steady. | Medium; readable for practitioners, the theory parts get spicy but skimmable. | Machine Learning journal or the author’s reprints page online. | It’s the paper behind half the Kaggle gold medals and more than a few production systems nobody brags about but everyone trusts. |
|---|
5 Object-Level Representation of Images | Jitendra Malik, Serge Belongie, Thomas Leung, Jianbo Shi | 1999 | Computer vision | Pushes beyond raw pixels to represent images in terms of objects and segments, laying groundwork for modern vision models. | It’s a reminder that all the fancy deep nets are still chasing the core idea: see scenes as objects, not noise. | The segmentation examples where messy real-world scenes suddenly break into clean, meaningful regions. | Medium-hard; you’ll need some comfort with both stats and vision basics. | Papers from the IEEE or CVPR proceedings; PDFs are widely mirrored. | Reading it now feels like seeing the early sketches for the stuff your phone casually does every time you open the camera. |
|---|