← Back to work
LiveModel

F1 Race Predictor

Three forecasters, one honest scorecard. The deliverable is the evaluation, not the prediction.

What this shows: I can test an AI system honestly: no peeking at future data, probabilities that mean what they say, and the result published even when the AI loses.

Outcome

A calibrated probability for every driver, every race, graded against the real finishing order once it exists.

Proof

Every pick is committed to git before the race, so the timestamp proves it was called in advance. Across nine graded 2026 rounds neither AI has beaten the simple baseline, and at that sample size the gap is still inside the noise.

Takeaway

Most of F1’s predictable signal is just where you start. The scorecard mattered more than the model.

Predicted order vs actual

Monaco GP · projected order vs actual finishLIVE

Monaco 2026, as it crossed the line

real result

At Monaco, overtaking is nearly impossible, so grid position usually decides the finish; the model leaned on that and called Antonelli's pole-to-win at 76%. Six cars retired, three our own podium picks (orange): externalities no pre-race data can see.

FINISHcasino climbhairpintunnelpiscineP4 Piastri: we said P7P5 Lawson: we said P10P6 Lindblad: we said P15P7 Gasly: we said P9P8 Albon: we said P11P9 Ocon: we said P17P10 Alonso: we said P21Antonelli: we said P1, finished P11ANTHamilton: we said P2, finished P22HAMHadjar: we said P6, finished P33HADVER: we said P3, then anti-stall at lights out forced him to retireVER · DNFmechanicalLEC: we said P4, then crashed at the final corner, bringing out the safety car and red flagLEC · DNFcrash → safety carNOR: we said P8, then a power-unit failure ended his raceNOR · DNFmechanicalTOP 3 · OUR PRE-RACE ODDSANT finished P1: 76% to win1ANT76% to winHAM finished P2: 54% podium odds2HAM54% podium oddsHAD finished P3: 13% · our P6 pick3HAD13% · our P6 pick6 DNFs · 2 crashes → safety car& red flag · 4 mechanical
ANT · MercedesHAM · FerrariHAD · Red Bulldnf · externality we can't predictrest of the field

Next race

projection · pre-qualifying

Round 11 · 26 Jul 2026

Hungarian Grand Prix

Hungaroring, Budapest

PROJECTED PODIUM

How often each slot delivers

P1

Russell

Mercedes

46% to win
P2

Hamilton

Ferrari

50% podium
P3

Norris

McLaren

47% podium

These are the model’s history, not Russell’s odds. Across 74 scored races, its pre-qualifying P1 pick won 46% of the time, and its P2 and P3 picks reached the podium 50% and 47%. The real per-driver probabilities arrive after qualifying.

Why these three

Russell won at the Red Bull Ring and has been the quickest Mercedes over one lap lately, so the form-only model makes him the pick. Antonelli still leads the championship on 158 points from Hamilton on 129 and Russell on 128, but he has not won since Monaco, and the model reads recent form rather than the table.

What changes Saturday

Grid position is the model's strongest feature, and qualifying has not happened yet. Once the real grid exists, all three forecasters rerun and the win and podium probabilities update before lights out. The Hungaroring is hard to overtake on, so the grid should matter more here than it did at Silverstone.

The form window stops at Silverstone

Spa ran on 19 July, but the results feed had not published them when this projection was built, so the model has not seen that race. Everything here is based on the season through round 9. That gap is worth stating rather than hiding: the pick could look different once Spa lands.

ON THE HORIZONR12Netherlands · 23 Aug·Russellform pickR13Italy · 6 Sep·Russellform pick

The season so far

stats model · graded

Every 2026 round, the model’s predicted podium against what actually happened. Green means a podium pick landed. Locked in after qualifying, before the race.

RoundOur podiumActual podiumWinner
R1 AustraliaRUSANTLECRUSANTLEC
R2 ChinaRUSANTLECANTRUSHAM×
R3 JapanANTRUSLECANTPIALEC
R4 MiamiANTRUSLECANTNORPIA
R5 CanadaANTRUSNORANTHAMVER
R6 MonacoANTHAMVERANTHAMGAS
R7 BarcelonaANTHAMRUSHAMRUSNOR×
R8 AustriaHAMRUSANTRUSVERANT×
R9 BritainANTHAMRUSLECRUSHAM×

Called the winner in 5 of 9 · landed 17 of 27 podium picks. The first half of the season was the easy half: the last three rounds all went to a driver we had further down the order.

At a glance

79
Races tested (2023-26)
3
Forecasters compared
22% worse
Claude's odds vs. just using the grid order. Nine races, so still inside the noise
2.79
Places off per driver, on average

How it works

  1. 1

    Gather

    Race and qualifying results for 2023 to 2026 become nine pre-race clues per driver: grid slot, quali gap, recent form, team pace, track history. Strictly nothing from the race being predicted.

  2. 2

    Predict

    Three forecasters fill in the same form: a naive baseline (you finish where you start), a statistical model trained only on past races, and Claude reasoning over a written pre-race brief.

  3. 3

    Grade

    Proper scoring rules (Brier score, log loss, skill vs the baseline) plus calibration curves: when it says 70%, does that happen 70% of the time?

  4. 4

    Track

    Every forecaster is graded race by race across the season, and the next race is always called before lights out, so the prediction is locked in before the result exists.

Honest about it

A prediction is only worth the eval behind it. So I keep score: three forecasters, every race, graded against what actually happened.

How often do we call the winner?

2026 · 9 rounds
naive baselineour model
0123AUSRUSCHNANTJPNANTMIAANTCANANTMONANTBARHAMAUTRUSSILLECAUS: naive baseline called the winner (RUS won)CHN: naive baseline called the winner (ANT won)JPN: naive baseline called the winner (ANT won)MIA: naive baseline called the winner (ANT won)CAN: naive baseline 1 place off (ANT won)MON: naive baseline called the winner (ANT won)BAR: naive baseline 1 place off (HAM won)AUT: naive baseline called the winner (RUS won)SIL: naive baseline 1 place off (LEC won)AUS: our model called the winner (RUS won)CHN: our model 1 place off (ANT won)JPN: our model called the winner (ANT won)MIA: our model called the winner (ANT won)CAN: our model called the winner (ANT won)MON: our model called the winner (ANT won)BAR: our model 1 place off (HAM won)AUT: our model 1 place off (RUS won)SIL: our model 3 places off (LEC won)we picked RussellLeclerc won, our P4places off the winner2026 round · who actually won it
Places off our winner pick, race by race, so zero means the winner was our top pick. The winner is usually the pole-sitter, so the baseline calls it six times in nine and our model five. The worst miss is Silverstone, where Leclerc won from our P4 and neither model saw it coming. Calling the winner is the easy part.
real output

Problem

Prediction posts are easy to fake after the fact, and LLMs make it worse: past seasons sit in their training data, so a strong backtest proves memory, not skill. I wanted calls put on the record before each race, and an evaluation I could actually trust.

Approach

Three forecasters emit the same output, so they compete like for like. The statistical model only sees earlier races, automated tests prove no future data leaks in, and Claude is graded only on races after its training cutoff. The rest of 2026 is the live test: every pick is locked in after qualifying, before the race.

Eval results

The headline is a negative result, the kind most write-ups quietly drop. Over the nine graded 2026 rounds, neither AI beat the baseline that just predicts the starting grid order: Claude came in 22% worse on win probability, the stats model 11% worse. Then the same honesty applied to my own headline. Nine races is thin, and bootstrapped error bars put Claude's win skill anywhere between 77% worse and 5% better. Every 2026 gap straddles zero, so the defensible claim is the direction and the uncertainty, not the number. The stats model shows why that matters: it led on wins earlier in the year and gave it all back once three different drivers won in a row. The second finding is the sturdy one. Almost all the predictable signal is a single feature: where you start. Remove grid position and podium error jumps about 19%. Remove any other feature and nothing moves.

What broke

A free data API silently returned four empty races after rate-limiting, caught by validation, not an error. Grid position 0 means a pit-lane start, which a model reads as better than pole. And the LLM sometimes returns duplicate finishing positions, so the schema rejects loudly and a deterministic repair re-ranks. The lesson that stuck: the eval design mattered more than the model. Most of the work was keeping the test fair.

Curious how it scores? Just ask.

🥬Kale Bot

Ask me anything about Cael and his projects. I answer from real sources, and I will tell you if I do not know.