Product

How Our FPL Model Did in GW1 and GW2 - The Full Scorecard

A bad first week, a strong second week, one result that beat us, and the change we made because of what we found. Every number, including the unflattering ones.

Most FPL prediction services tell you how good they are. Very few tell you how they actually did last week. We think that is backwards - so from now on, every gameweek gets scored in public, win or lose. This is the first instalment, covering the opening two gameweeks of 2026/27.

First, the jargon - in plain English

There are only three technical terms in this post. Here they are, once, in normal words.

Average error ("MAE")

On average, how many points our prediction missed by. If we said 5 and a player scored 8, that is an error of 3. Lower is better. An average error of 1.2 means our typical prediction was about 1.2 points off.

Ranking score ("rank correlation")

Sort every player by what we predicted. Now sort them again by what they actually scored. How similar are the two lists? 1.0 means we got the order perfectly right. 0 means our order was no better than random. This is the number that matters most for FPL, because you are choosing between players, not trying to guess exact scores.

"Players who played"

We report every number twice - once for all players, once for only the players who actually got on the pitch. The second one is the hard test. Correctly predicting zero for an injured reserve is easy and flatters the score; separating a 2-point start from an 11-point start is the actual job.

The scorecard

The right-hand column is our benchmark: how the model performed across 19 gameweeks of historical data it had never seen during development. That is the standard we expect it to hold in the real world.

MeasureGW1GW2BothBenchmark
Average error, all players1.471.181.321.01
Ranking score, all players0.5270.7150.6330.703
Ranking score, players who played0.2320.4300.3190.400
Top captain pick scored2 (Haaland)23 (B.Fernandes)-8.9 avg

GW2 matched or beat the benchmark on both ranking measures. GW1 fell well short of it on everything.

A warning about that "Both" column: it is the least meaningful number in the table, and we have included it only because leaving it out would look like hiding. Averaging the two weeks blends a forecast made 13 days early by an older version of the model with one made properly, so the result describes neither. The two weeks are not measuring the same thing and should be read separately.

GW1: a bad week, and an honest one

GW1 was our worst measured gameweek. Our top captain pick, Haaland, scored 2. Our ten highest-predicted players returned 44 points between them, against the 134 that the ten actual top scorers made.

The gameweek itself was chaos. The top scorers were De Cuyper (£4.5m, 17 points), Hinshelwood (16), two £4.0m Hull defenders on 15 and 14, and Kayode (13). Almost every premium player blanked. As a sanity check we asked how well you would have done by simply ranking every player by price - the crudest possible model. It scored 0.128 among players who played. Our 0.232 was roughly twice as good, but both numbers are poor, and that tells you the week was close to unpredictable rather than that we were unusually clever.

There is a second, less comfortable caveat. The GW1 prediction we scored was generated 13 days before the deadline, and by an older build of the model. It was the freshest copy we had kept. That is a record-keeping failure on our side rather than a modelling one, and it is fixed now - which brings us to the most useful thing we learned.

GW2: what a normal week looks like

GW2 is the first gameweek this model was measured on fairly - a current build, saved properly before the deadline. It performed at exactly the level its historical testing said it should.

Ranking players who actually playedScore
Our model, GW20.430
Ranking by price, GW20.261
Our model, long-run benchmark0.400
FPL's own xP, long-run benchmark0.315

Telling apart players who all actually started is the hardest thing this model does, and the number that matters most for picking between two players you are torn over. In GW2 it did that better than its own long-run average, and comfortably better than ranking by price - the market's own verdict on who is good.

One thing that table does not show: a head-to-head against FPL's xP in GW2 itself. The bottom two rows are long-run averages from historical testing, not GW2 results. We could not compute the GW2 comparison for the reason explained further down, and we are not going to imply one by putting numbers from different tests side by side without saying so.

The model's top pick, Bruno Fernandes, scored 23 - the highest score of any player in the game that week.

The big finding: when we run matters more than anything else

For GW2 we happened to keep two saved predictions from the exact same model - one made six days before the deadline, one made a day and a half before. Same code, same gameweek, same players. Only the timing differed.

Prediction madeRanking score, allRanking score, played
6 days before deadline0.6060.351
1.5 days before deadline0.7150.430

That gap is large and it holds up statistically. To put it in context: we have spent months testing new information sources to add to this model, and the best of them moved the ranking score by less than 0.001. Simply running the model later moved it by 0.109.

The reason is not mysterious. Most of what separates a good FPL prediction from a bad one is knowing who is actually going to start. Injury news, press conferences and rotation hints all land in the last few days before a deadline. A model run on Sunday is guessing about things that become public on Friday.

So we changed the schedule. Predictions used to refresh every 6 hours on a fixed clock. They now run once each morning, plus three more times in the final three hours before every deadline. Fewer runs overall, far better placed. The Predicted Points page always shows when the model last ran - and the practical advice that follows from this is simple: check it again just before the deadline.

Where we got beaten in GW2

GW2 was a good week overall, but not everywhere. Our ten highest-predicted players returned 66 points. Ranking the same field by price alone returned 97. The dumbest possible baseline picked a better top ten than we did.

We can see exactly why. The model liked Chelsea and Liverpool defenders that week - five of them made our top 20, on the expectation of clean sheets. Chelsea v Brighton finished 4-3. Liverpool v Nottingham Forest finished 2-2. Eleven goals across the two matches we most fancied to keep things tight, and our nine highest-rated defenders and goalkeepers returned 19 points between them.

This is a real structural weakness rather than bad luck, and it is worth understanding if you use the predictions: the model rates each defender separately, so when it likes a defence it tends to like all of that team's defenders at once. If the clean sheet does not arrive, several of its top picks fail together. Our practical advice, which we follow ourselves: do not stack three defenders from the same team just because they all appear near the top of the list.

The statistical honesty note: with only ten players involved, this gap is not yet statistically meaningful, and in GW1 the comparison went the other way (89 for us, 79 for price). We are flagging it, not concluding from it. If it persists across more gameweeks, it is a real problem and we will say so.

What we cannot tell you yet

The obvious question is whether we beat FPL's own published expected points - the "xPts" figure in the game. For GW1 and GW2, we honestly cannot say. FPL overwrites that number every gameweek, and we were not keeping a copy of it alongside our own predictions, so the comparison cannot be rebuilt after the fact.

We have fixed it: from GW3, every saved prediction stores FPL's figure next to ours, on exactly the same players, and every future gameweek review will publish the head-to-head whichever way it falls. GW1 and GW2 stay permanently un-benchmarked, and we would rather leave that gap visible than fill it with an estimate.

The model's own FPL team

Accuracy statistics are abstract, so we also run a real FPL team picked entirely by the model. It chose its own opening squad and makes its own transfers under the normal rules - no chips, no points hits, and no human overriding it when we disagree.

GameweekModel teamGame averageDifference
GW16250+12
GW29781+16
Total159131+28

That puts it at an overall rank of about 1.13 million out of 10.26 million managers - roughly the top 11% - after two gameweeks. It captained Bruno Fernandes both weeks, for 2 points and then 23.

It has not made a transfer yet. From here we ask it for its transfers the day before each deadline and then make exactly what it says - no overriding it when we disagree. The day before, rather than now, for the reason above: a prediction made days early is measurably worse than one made late, so asking it this far out would just be acting on the weaker forecast.

The honest summary

The verdict, stated plainly so you do not have to infer it from a pile of caveats: in the one gameweek this model was measured fairly, it performed at the level its testing said it should, and the team it picks itself sits in the top 11%. That is a good start. It is also only a start - the caveats below are about how much can be concluded from two weeks, not about the results being weak.

  • GW2 was good. It matched the model's historical benchmark, beat a price-ranking baseline clearly, and its top pick was the highest scorer in the game.
  • GW1 was poor, but it was not a fair test. That forecast was 13 days stale and from an older build. Given we now know that six days of staleness alone costs about 0.109 of ranking score, a 13-day-old forecast scoring badly needs no other explanation. It was also a genuinely freak week in which ranking by price scored 0.128 - close to useless - so the market had no idea either.
  • A naive baseline beat our top ten in GW2. Not yet statistically meaningful, and we are watching it.
  • Timing beats cleverness. The single biggest improvement available to us was running the model later, not making it smarter.
  • Two gameweeks proves very little. Anyone claiming a verdict on a prediction model after two weeks - including us - is over-reading the data. Check back at GW10.

See this week's predictions

Predicted Points is part of Pulse Pro - £2.99/month or £24.99/year, with a 7-day free trial. Every player, five gameweeks ahead, refreshed each morning and again in the final hours before every deadline.

FAQs

How accurate were FPL Pulse's predictions in GW1 and GW2?
The two weeks are best read separately, because the GW1 forecast we have on record was 13 days stale and from an older build. GW2, the first week measured fairly, had an average error of 1.18 points per player and a rank correlation of 0.715 across all players and 0.430 among players who actually played - matching or beating the benchmark the model set across 19 gameweeks of held-out historical data. GW1 scored 0.527 and 0.232 on the same measures, on a freak week where ranking players by price scored just 0.128, so the market was close to useless too.
Did the model beat FPL's own expected points (xP)?
We genuinely do not know for GW1 and GW2, and we would rather say that than guess. FPL overwrites its published expected-points figure every gameweek and we were not saving a copy at the time, so the comparison cannot be reconstructed after the fact. We started archiving it from GW3 onward, and every gameweek from GW3 will be published with the head-to-head included.
Did anything beat the model?
Yes. In GW2, simply ranking players by price picked a better top ten than our model did - 97 points versus 66. Our model ranked the overall field far better, but its very top picks leaned on Chelsea and Liverpool defenders keeping clean sheets, and those two matches finished 4-3 and 2-2. With only ten players involved this is not yet a statistically meaningful result, but we are publishing it and watching it.
How is the model's own FPL team doing?
It scored 62 in GW1 against a game average of 50, and 97 in GW2 against an average of 81. That is 159 points and an overall rank of about 1.13 million out of 10.26 million managers, so roughly the top 11%. It is a real FPL entry picked entirely by the model, with no human overrides.
How often do the predictions update?
Once every morning as a baseline, plus three extra runs in the final three hours before each deadline. We changed to this after measuring that a prediction made 1.5 days before a deadline is significantly more accurate than the same model's prediction 6 days out. Running late matters more than running often.
Why publish results that make the model look bad?
Because a prediction model you cannot audit is just a number on a screen. Anyone can publish their good weeks. We publish every week, including the ones where a naive baseline beat us, so you can judge for yourself how much weight to put on the predictions.
© 2026 - Not affiliated with the Premier League.