Product

GW3: Our Model vs FPL's Own Expected Points

Two gameweeks ago we said we could not tell you whether our predictions beat FPL's own numbers, because we were not saving them. We are now. Here is the first head-to-head, plus the week's full scorecard.

The short version: GW3 was our most accurate week so far, and the model beat FPL's published expected points on every measure we track. It also lost to a price ranking at the very top of the board for the second week running, and it is still below its own benchmark at the hardest job in FPL forecasting. All of that is below.

The head-to-head we promised

FPL publishes its own expected-points figure for every player before each deadline, then overwrites it. From GW3 we save a copy alongside our own prediction, so the two can be scored on exactly the same 652 players. Lower is better on error; higher is better on ranking.

GW3, same playersOur modelFPL's own xP
Average error, all players1.051.29
Ranking score, all players0.7750.706
Average error, players who played1.892.43
Ranking score, players who played0.3270.210
Points from its top ten picks4227

Four for four, and the two ranking gaps hold up when we resample the week thousands of times to see whether they could be chance - the margins are +0.068 across all players and +0.118 among those who played, and neither range crosses zero.

The caveat that matters: this is one gameweek. FPL's xP had a genuinely poor week - among players who started, it ranked them worse than simply guessing the average would have on error. Beating a benchmark on its bad week is not the same as beating it in general. Our historical testing across 19 gameweeks also has us ahead of xP, which is why we expect this to hold, but one week does not demonstrate it.

The season so far

The benchmark column is how the model performed across 19 gameweeks of historical data it never saw during development. That is the standard we expect it to hold week to week.

MeasureGW1GW2GW3Benchmark
Average error, all players1.471.181.051.01
Ranking score, all players0.5270.7150.7750.703
Ranking score, players who played0.2320.4300.3270.400
Highest-predicted player scored2 (Haaland)23 (B.Fernandes)9 (Haaland)8.9 avg

GW3 was the best week yet on the two whole-field measures, and the first to essentially match the benchmark on average error. The third row is the honest weak spot: 0.327 against a 0.400 benchmark. Telling apart players who all actually started is the hardest thing the model does, and in GW3 it did it worse than its long-run average - while still doing it better than FPL's own numbers did.

Taking the two weeks that were measured properly - GW2 and GW3, the same model version saved just before each deadline - the combined picture is 1.11 average error and 0.745 ranking across all players. We are not averaging GW1 into that. That forecast was 13 days stale and from an older build, and as the next section shows, staleness alone explains the whole gap.

When you check matters - but not in the way we expected

Last time we reported that the same model, run 1.5 days before a deadline instead of 6 days before, ranked players 0.109 better. GW3 gave us a longer gap to test and a shorter one, in the same week.

Prediction madeRanking score, allRanking score, played
13 days before deadline0.5730.135
4 hours before deadline0.773-
1 hour before deadline0.7750.327

Thirteen days of staleness cost 0.208 - twice what six days cost in GW2, in the same direction, which is about what you would expect. A stale forecast is not a slightly worse forecast; among players who actually played, the 13-day-old numbers scored 0.135, which is close to useless.

The genuinely useful finding is the last two rows. Between four hours out and one hour out, the ranking moved by 0.002. Nothing. Once the week's team news has landed, the picture stops changing. So the practical advice we gave last time needs correcting: check the predictions after Friday's press conferences, not frantically at T-minus-one-hour. The extra runs we schedule near each deadline are insurance against a missed run, not extra accuracy.

Where we got beaten again

In GW2, ranking players by price alone picked a better top ten than our model did. It happened again in GW3.

Points returned by the top...10 picks20 picks
Our model4292
Ranking by price4878
FPL's own xP2763

Two things are true at once here. The gap closed sharply - 31 points in GW2, 6 in GW3 - and we won at twenty picks and beat FPL's xP at both. But we are not going to declare that fixed off one narrowing, because at ten players a six-point difference is statistical noise (the formal test returns p = 0.56, nowhere near meaningful).

The GW2 diagnosis was that the model rates each defender separately, so when it likes a defence it likes all of that team's defenders, and a single goal wrecks several picks at once. GW3 looks better on that specific point - three of its top ten were defenders or keepers, against five from just two defences in GW2. That is what we would want to see, but one week is not evidence of a fix, and our advice is unchanged: do not stack three defenders from one team just because they all appear near the top.

The individual misses are worth naming too. Five of our top ten scored 2 or fewer. Enciso, our second-highest pick, played 90 minutes and returned 1. Foden and Collins were both in our top five and managed 24 and 16 minutes on the pitch, for 1 point each. In the other direction we badly under-rated Bogle (predicted 1.8, scored 14) and Mitchell (predicted 3.5, scored 15) - two cheap defenders, which is the same blind spot GW1 exposed.

The thing we are still bad at

Every week we run the same diagnostic: what would the model score if someone handed it the team sheets in advance, and it only had to rank the players it knew would appear?

In GW3 that takes the ranking score from 0.775 to 0.896, with no change to the model whatsoever. Knowing who plays is worth more than every statistical improvement we have tested in months of work put together. The encouraging part: players who never took the pitch carried 15% of our total error in GW3, down from 19% in GW2 and 25% in GW2's stale forecast. Running the model late is what buys that, which is the same lesson as the timing table.

The model's own FPL team

We run a real FPL entry picked entirely by the model - no chips, no hits, and no human overriding it when we disagree.

GameweekModel teamGame averageDifference
GW16250+12
GW29781+16
GW35651+5
Total215182+33

That is an overall rank of about 961,000 out of 10.26 million - roughly the top 9%, improved from the top 11% after GW2. Ahead of the average in all three weeks, including both weeks its captain blanked. Being above average three times running is also roughly what a coin does, so treat it as a fun public commitment rather than a result.

GW3 brought its first transfers, both chosen by the model with 2 free transfers and no points hit: Gabriel out for Collins, and van Ewijk out for N.Williams. Those returned 1 and 7 against the 2 and 3 they replaced - a net gain of 3 points and £1.5m banked. It captained Szoboszlai, who scored 3.

One detail worth admitting, because it is the timing finding showing up in advice rather than in statistics: an earlier run that week recommended two different transfers and a different captain. The run we acted on was the one made hours before the deadline. Same model, same week - the stale version wanted a different team.

The honest summary

  • GW3 was our best week, and we beat FPL's own xP on every measure. First time we could check it. Both ranking margins survive a statistical test.
  • xP had a bad week too. Beating a benchmark once, on its bad week, is not the same as beating it in general. Ask us again at GW10.
  • Stale predictions are much worse than we thought - and last-minute ones are not better than same-day. Thirteen days of staleness cost 0.208. The last four hours before a deadline were worth 0.002.
  • A price ranking beat our top ten again, by 6 points after 31. Still not statistically meaningful, still being published, still being watched.
  • Our weak spot is unchanged. Separating players who all started scored 0.327 against a 0.400 benchmark. Knowing the team sheets would be worth more than every modelling improvement we have found this year.
  • Three gameweeks still proves very little. Two of them were measured fairly. That is two data points.

See this week's predictions

Predicted Points is part of Pulse Pro - £2.99/month or £24.99/year, with a 7-day free trial. Every player, five gameweeks ahead, refreshed each morning and again in the final hours before every deadline.

FAQs

Did FPL Pulse's model beat FPL's own expected points in GW3?
Yes, on all four measures we track, and GW3 is the first week we could check. Across all 652 players our average error was 1.05 points against xP's 1.29, and our ranking score was 0.775 against xP's 0.706. Among only the players who actually took the pitch - the harder test - we scored 0.327 against xP's 0.210. Both ranking gaps survive a statistical check, so this is a real result for one gameweek rather than a rounding difference. It is still only one gameweek.
How accurate was the model in GW3?
It was our best week yet on two of the four measures. Average error was 1.05 points per player, against the 1.01 our long-run historical testing says to expect, and the ranking score across all players was 0.775 against a 0.703 benchmark. The weak spot is unchanged: separating players who all actually started scored 0.327, below the 0.400 benchmark.
Did anything beat the model in GW3?
Yes - ranking players by price picked a better top ten than we did for the second week running, 48 points against our 42. The gap narrowed a lot (it was 97 to 66 in GW2) and at ten players it is nowhere near statistically meaningful, but we said we would watch it and publish it either way. Our top twenty beat price ranking 92 to 78, and beat FPL's xP on both.
Does it matter what time of day the predictions are refreshed?
Enormously versus days-old numbers, barely at all in the last few hours. A GW3 forecast made 13 days early scored 0.573; the same model an hour before the deadline scored 0.775. But between four hours out and one hour out the ranking moved by 0.002 - nothing. In practice: check the predictions once team news has landed, and do not worry about refreshing again right before the deadline.
How is the model's own FPL team doing?
It scored 56 in GW3 against a game average of 51, taking it to 215 points and an overall rank of about 961,000 out of 10.26 million - roughly the top 9%, up from the top 11% after GW2. It also made its first transfers, both suggested by the model with no human input: Gabriel out for Collins, and van Ewijk out for N.Williams. Those two returned 8 points against the 5 they replaced.
What is the model still bad at?
Knowing who will play, and by extension separating two players who both start. If we magically knew the team sheets in advance, GW3's ranking score would jump from 0.775 to 0.896 with no change to the model at all. That single unknown is a bigger lever than any new statistic we have tested in months of trying.
© 2026 - Not affiliated with the Premier League.