← All writing

Case study

What I learned building an analytics app for my ESPN fantasy league

Luck, bench points, draft grades and trade values for a 12-team league, and the lesson I relearned every week: the dangerous bugs don't crash, they return a confident wrong number.

7 min read
  • Python
  • React
  • Statistics
  • Fantasy football

ESPN tells you who won. My league argues about why. Was that loss bad luck or a bad lineup? Who left the most points on the bench? Did anyone actually win the draft, or did they just stay healthy? Is the trade in the group chat fair, or is somebody getting fleeced?

So I built a place to settle it: a Flask API and a React app that read our 12-team league through ESPN's fantasy API, pull in eleven seasons of public NFL box scores, and answer those questions with numbers. It started as a dashboard. It is now a couple dozen pages, about 36,000 lines of Python and TypeScript, and 626 automated tests.

This is a tour of what it does. But the part worth writing down is a lesson I relearned about once a week: the dangerous bugs don't crash. They return a confident number that is wrong.

What's in it

  • Standings and luck. Records rebuilt from the scoreboard, plus an all-play record that scores every team against every other team, every week.
  • Season audit. Where a team's record sat among all the records the same scores could have produced, and how many wins each manager left on the bench.
  • Playoff odds by simulation, and bench efficiency against the optimal lineup every week.
  • Draft tools. A board built from thirteen ranking and ADP lists in three scorings, a comparison of what two lists actually disagree about, a strategy lab that drafts a whole room of bots, a draft recap graded against the preseason market, and an ADP heat map (more below).
  • Trade analyzer. Grades a deal by what it does to each team's best legal lineup and playoff odds, not by the points changing hands.
  • Weekly projections built from NFL data and scored against ESPN's, and a stats lab that runs textbook statistics on the league's own scores.

The numbers were quietly wrong

The first version of the dashboard looked great. Several of the numbers on it were wrong, and nothing ever raised an error. A few of my favorites:

  • "Total points" was one week's score. The column ESPN's sync filled was a single week's points, repeated on a row per week. The page summed nothing and labeled it a season.
  • The champion was seeded third. Standings sorted on raw win count. A 9-4-1 season has the same win count as a 9-5 one, so the tie the team had earned counted for nothing. Records are now rebuilt from who actually outscored whom, and seeded by win percentage, then points for.
  • The leaderboard was empty forever. Our weekly pick'em leaderboard read from a table that no code had ever written to. It rendered an empty state, politely, for every league.
  • The wrong team was crowned. The season page called the top of the regular-season table the champion. That was the wrong team in both seasons our league has played: the top seed lost the final. Worse, the sync stopped at week 14, so not a single playoff game had ever been stored.
  • Everything returned 403. ESPN retired the host an old version of the API library used, then our league went private and every call needed two browser cookies.

The fix for this category isn't cleverness. It's a habit: derive numbers from the rawest data you have (the scoreboard, not a summary column), and for every bug, write the test that reproduces the original wrong number. Those tests fail with messages like 120.0 == 220.0, which is exactly the bug a user saw, and they stay in the suite as tripwires.

Measuring luck

Head-to-head fantasy has a coin flip built in: you can post the second-highest score in the league and lose, because your opponent posted the highest. The all-play record removes the schedule. Each week, every team plays every other team's score. An example:

WeekYour scoreTeams beaten (of 11)All-playActual result
A151.2100.91Loss (you drew the top score)
B88.410.09Win (you drew the only worse one)

Average the all-play rate over a season and multiply by games played, and you get expected wins. The gap between that and the real record is luck. The season audit goes one step further. Redrawing the schedule makes your opponent each week uniform over the league, so your chance of winning that week is exactly your all-play rate, and the season's win total follows a Poisson binomial distribution. That can be computed exactly, with no simulation, and ties count as half a win.

The hardest part was names

The draft board lines up eight ranking lists player by player. The ranks are trivial; deciding that two rows are the same human being is not. One source writes "Marvin Harrison Jr.", another "M. Harrison". Those are one player. But "Bijan Robinson" and "Brian Robinson Jr." differ only by the suffix, and an early version merged them, reporting the consensus RB2 as one analyst's RB57.

The rule that fixed it: suffix variants merge only if no single list contains both. A ranking list doesn't rank the same player twice, so a list holding both spellings proves they are different people. Most of the draft board's tests are about identity, not arithmetic, because a board that renders the wrong two players side by side fails silently.

Projections: ESPN still wins, and the page says so

I wanted a weekly PPR projection that beats ESPN's. Before writing a single feature, I built the scoreboard it would be judged on. The harness hands a model only what was knowable before kickoff. It hides gameday inactives, because letting them through tells a model exactly who scored zero. It fixes the set of scored players so a model can't pick its own test, splits by season (tune on 2018–2023, validate on 2024, test once on 2025), and bootstraps differences over weeks, not player-games.

SeasonRoleMAE behind ESPN
2024Validation+0.56 pts
2025Test (touched once)+0.39 pts
2026, weeks 1–2Unseen (this season)+0.48 pts

ESPN wins on error every season so far. The projections page shows that scorecard right next to the numbers, including the weeks ESPN wins, which is all of them. There is one small bright spot. In week 4, the first week published before a single game was played, the model's ordering of players within each position beat ESPN's (rank correlation 0.40 vs 0.36), even though its point estimates were further off. A model that isn't allowed to look at the answer is humbling.

How much of fantasy is skill?

The stats lab asks the question everyone in a league has an opinion on. Treating each team-season as a group and each week as an observation, an analysis of variance on our scores says about 9% of a weekly score's variance belongs to the team. The F-test is significant (p < 0.001), so skill is real. It's just small next to the week-to-week noise. The same lab found that teams' weekly volatilities are statistically indistinguishable, and that the shrinkage constant the playoff odds use may be too small, though two seasons can't pin it down. It turns out the honest output of most of these tests is "you need more seasons".

The ADP heat map

The newest page borrows an idea from Hayden Winks: take the ADP board as a 12-team snake draft, and color every pick by the win rate of the fantasy teams that drafted him. Yahoo can compute that from millions of real leagues. I can see one, so the rooms are simulated and the points are real. The app drafts 2,000 leagues off the ADP, each player going somewhere between the earliest and latest pick real drafts took him at. Then it scores every roster's best lineup each week on actual NFL points, and records the share of the other eleven teams it outscored.

2025 ADP board · rounds 1–3 12-team PPR · weeks 1–17
  1. 1.0149
  2. 1.0256
  3. 1.0346
  4. 1.0442
  5. 1.0544
  6. 1.0655
  7. 1.0740
  8. 1.0861
  9. 1.0957
  10. 1.1048
  11. 1.1152
  12. 1.1253
  13. 2.1253
  14. 2.1145
  15. 2.1047
  16. 2.0952
  17. 2.0846
  18. 2.0759
  19. 2.0649
  20. 2.0551
  21. 2.0449
  22. 2.0344
  23. 2.0257
  24. 2.0150
  25. 3.0147
  26. 3.0253
  27. 3.0344
  28. 3.0450
  29. 3.0552
  30. 3.0646
  31. 3.0745
  32. 3.0845
  33. 3.0945
  34. 3.1059
  35. 3.1160
  36. 3.1258
≤40%≥60% Hover a pick: win % of the teams that drafted him
The first three rounds of 2025 PPR ADP, colored by how the teams that drafted each player did across weeks 1–17. Hover or tab through for the pick.

Christian McCaffrey, the 1.08 and the fourth back off the board, finished RB1, and the teams that took him won 61% of their weeks. Malik Nabers went one pick earlier and finished WR98; his teams won 40%. Jaxon Smith-Njigba at 3.11, the sixteenth receiver taken, finished WR2, and his teams won 60%. The rooms are seeded from the ADP list alone, so when a new week of games lands, the same 2,000 leagues are re-scored. A number moves only because of football.

What I'd tell myself at the start

  1. Build the scoreboard before the model. Every rule decided before tuning is one you can't bend to flatter yourself later.
  2. Rebuild derived numbers from raw data. Summary columns lag, repeat and drop ties.
  3. A regression test is only worth having if it fails when the bug comes back. I check by putting the bug back and watching the right test fail.
  4. "Not enough data" is a valid answer. Most pages here say so somewhere, and that's what makes the other numbers believable.

Under the hood

Python 3.12, Flask, SQLAlchemy with Alembic migrations, pandas and NumPy, and a hand-written statistics module (no SciPy). The front end is React with TypeScript, Vite, Tailwind and TanStack Query. NFL data comes from nflverse's public releases, and league data from ESPN through the espn-api library. I'm cleaning the code up to open-source it, and I'll link the repo here when it's public.