GW4: The Week We Drew With FPL's Own Numbers
Last week our model beat FPL's expected points on every measure and we said so. This week it finished ahead on every measure again - and we are calling it a tie. Here is why.
The short version: GW4 was a perfectly normal week for the model, a better week for the benchmark it is measured against, and a bad week for the FPL team it picks. All three of those things are in here, with the numbers.
The head-to-head
Both forecasts scored on exactly the same 656 players. Lower is better on error, higher is better on ranking.
| GW4, same players | Our model | FPL's own xP | Call |
|---|---|---|---|
| Average error, all players | 1.17 | 1.23 | too close |
| Ranking score, all players | 0.725 | 0.706 | too close |
| Average error, players who played | 2.12 | 2.30 | too close |
| Ranking score, players who played | 0.389 | 0.353 | too close |
Four numbers in our favour and not one of them worth anything. The test we run is simple: resample the gameweek a couple of thousand times and see whether the gap ever disappears. In GW3 it did not - the margins were +0.068 and +0.118 and both ranges stayed clear of zero. In GW4 they are +0.018 and +0.036, and both ranges cross zero. That is a draw, and we would rather write "draw" than print four green ticks.
For contrast, on the same players with the same test, our margin over a simple price ranking was +0.360 across all players. That is what a gap that means something looks like.
What changed between GW3 and GW4 was not us
| Ranking players who started | GW3 | GW4 |
|---|---|---|
| Our model | 0.327 | 0.389 |
| FPL's own xP | 0.210 | 0.353 |
We got better at the hardest measure between the two weeks. The gap closed because FPL's numbers improved more. Last week's post said our GW3 result was partly xP having a bad week; GW4 is the evidence for that, arriving one week later.
The general lesson is worth more than either week: when you score two forecasts against a single gameweek, you are measuring the gameweek as much as the forecasts. Anyone showing you one week of accuracy figures - us included - is showing you something with enormous error bars. This is why we keep publishing them in a row rather than picking one.
Three weeks in, a pattern
GW2, GW3 and GW4 were all measured the same way: current model, saved just before the deadline. GW1 is excluded because that forecast was 13 days stale and from an older build. Three weeks is the first time anything here has looked like a pattern rather than a sample.
| Ranking score | GW2 | GW3 | GW4 | Benchmark |
|---|---|---|---|---|
| All players | 0.715 | 0.775 | 0.725 | 0.703 |
| Players who played | 0.430 | 0.327 | 0.389 | 0.400 |
Above the benchmark on the full field in all three weeks; slightly below it at separating players who all started, in two of three. Pooled across the three weeks that is 0.738 and 0.385 against a 0.703 and 0.400 benchmark, with an average error of 1.13 points per player. In plain terms: the model is doing a bit better than expected at sorting the whole player pool, and a shade worse than expected at the specific job of choosing between two players who will both definitely start.
The model's team had its worst week
| Gameweek | Model team | Game average | Difference |
|---|---|---|---|
| GW1 | 62 | 50 | +12 |
| GW2 | 97 | 81 | +16 |
| GW3 | 56 | 51 | +5 |
| GW4 | 57 | 69 | -12 |
| Total | 272 | 251 | +21 |
Its first below-average week dropped it from an overall rank of about 961,000 to about 2.6 million - 1.6 million places for 12 points, which tells you how tightly packed the middle of the game is more than it tells you anything about the model. Bruno Fernandes was captained and scored 2. Mbeumo, Szoboszlai, Tzolis and Isak returned 2, 3, 3 and 2. The one thing that worked was its transfer: Hughes out, Tavernier in, 1 point for 8.
Here is the part worth sitting with, because it cuts against us as often as for us: GW4 was a good forecasting week and a bad team week at the same time. The predictions edged FPL's own numbers on all four measures while the squad built from them finished 12 below average. Ranking 656 players and picking 11 are different problems, and at eleven players the second is mostly luck. If you take one thing from this series, take that - and apply it to our good weeks too.
The timing finding, third measurement
We had a 21-day-old forecast for GW4 sitting in the archive alongside the one made three hours before the deadline. Same model, same players, only the timing different.
| Prediction made | Ranking score, all | Ranking score, played |
|---|---|---|
| 21 days before deadline | 0.581 | 0.314 |
| 3 hours before deadline | 0.725 | 0.389 |
That is the third week running that a fresher forecast has ranked the full field better, now measured at six, thirteen and twenty-one days of staleness. It is the most reliable finding in this series.
One correction to last week, in the direction that costs us something: for the harder measure - separating players who all actually started - GW4's gain was small enough that it could be chance. So we are no longer claiming the freshness effect for that specific case, only for ranking the field as a whole. The advice does not change: check the predictions once the week's team news has landed.
The honest summary
- GW4 was a draw with FPL's own xP. We were ahead on all four measures and none of the margins survived a statistical check. One win, one draw.
- The change from last week was the benchmark, not us. Our hardest measure improved, 0.327 to 0.389. FPL's improved more.
- Three clean weeks now show a consistent shape. Better than benchmark at ranking the whole field, slightly worse at separating confirmed starters.
- The model's team had its worst week in the same gameweek its predictions were good. Both are true; neither proves anything about the other.
- Fresher is better, on the full field. Third measurement, same direction. We have dropped the claim for confirmed starters specifically.
See this week's predictions
Predicted Points is part of Pulse Pro - £2.99/month or £24.99/year, with a 7-day free trial. Every player, five gameweeks ahead, refreshed each morning and again in the final hours before every deadline.
FAQs
- Did FPL Pulse's model beat FPL's expected points in GW4?
- It finished ahead on all four measures, but by margins too small to call a win. Our ranking score was 0.725 against xP's 0.706 across all players, and 0.389 against 0.353 among players who played. When we resample the gameweek thousands of times to test whether those gaps could be chance, both ranges include zero. So GW4 goes in the record as a tie. Last week's GW3 gaps were about four times larger and did clear that test.
- Why was GW3 a clear win and GW4 a draw?
- Mostly because FPL's xP had a better week, not because our model had a worse one. Our own numbers barely moved between the two weeks - a ranking score of 0.775 then 0.725. FPL's xP went from ranking players who started at 0.146 to 0.297. That is a useful warning about single-week comparisons in general: when you score two forecasts against one gameweek, you are measuring the gameweek as much as the forecasts.
- How accurate was the model in GW4?
- Average error of 1.17 points per player, and a ranking score of 0.725 across all players against a long-run benchmark of 0.703. Among players who actually took the pitch it scored 0.389 against a 0.400 benchmark. Across the three gameweeks measured properly so far - GW2, GW3 and GW4 - it sits at 0.738 on the full field and 0.385 among players who played.
- How is the model's own FPL team doing?
- It had its worst week: 57 points against a game average of 69, dropping it from an overall rank of about 961,000 to about 2.6 million. Its captain, Bruno Fernandes, scored 2, and four of its five most expensive outfielders returned 2 or 3 points each. Its one transfer worked - Hughes out for Tavernier, 1 point for 8 - but that was the only part of the week that did.
- How can the predictions be good in a week the model's team was bad?
- Because they are different jobs. Ranking 656 players well and picking the right 11 are not the same task, and eleven players is a small enough sample that one bad week says almost nothing. GW4 is a clean example: the forecast beat FPL's own numbers on every measure while the squad built from it finished 12 points below average. We publish both precisely so neither gets read as proof of the other.
- Does refreshing the predictions later still help?
- On the full field, consistently yes - it is now measured three times, at six days, thirteen days and twenty-one days of staleness, and running later has helped every time. For separating players who all started, GW4's gain was not distinguishable from zero, so we are no longer claiming the effect for that specific case. The practical advice is unchanged: check once the week's team news has landed.