We develop a data-driven model that estimates the number of chances created (CCs) by a typical player, based on various opportunities presented to him, such as minutes on the pitch, crosses he can make, and passing opportunities in the final third. This model establishes a sliding scale for benchmarking all playmakers in terms of their chance creation capabilities. Players who excel in making incisive passes or successful crosses will outperform the average, generating more CCs than expected. This surplus in opportunity-adjusted chance creation serves as our measure of playmaking talent, with the most exceptional playmakers identified as outliers in excess CC.
We utilize public domain data from Opta Analyst to replicate these analyses across eight different European leagues for the 2025-26 season, along with a legacy dataset from the previous season for one of the leagues. Our model employs robust regression techniques, which effectively down-weight the few observations that deviate from the trend. This approach yields a more refined understanding of football statistics than traditional OLS regression. Importantly, robust regression helps pinpoint outliers that diverge from expected trends. Typically, these outliers are dismissed as erroneous data caused by coding mistakes or unidentified anomalies. However, when evaluating playmakers, positive outliers are crucial and warrant significant attention.
One of the subplots of this year’s EFL was Bruno Fernandes’ incredible achievement of 21 assists for the season, breaking the previous records set by playmakers like de Bruyne (20), Henry (20), and Özil (19). To accomplish such a feat in chance creation, you need to be an outstanding playmaker, but luck also plays a role. BF’s expected assists for the season were only 12.32, suggesting that fortune, excellent finishing, and/or indifferent goalkeeping contributed to boosting his football statistics.
Opportunity was also a key factor: BF played 90% of the season’s minutes, which is substantial for players in forward positions, and he took most free kicks and 90 corners. When set pieces are factored out, his xA drops to just 9.12 in open play. Moreover, his teammates significantly contributed by positioning BF to create those chances, which Manchester United certainly did. For example, BF made 871 passes in the final third—10% more than the next highest playmaker in the league, Dominik Szoboszlai.
Given these boosts, how remarkable was BF’s underlying performance when adjusted for opportunity? And which other playmakers are noteworthy, particularly those who may not have enjoyed the same opportunities?

Set pieces are undoubtedly important in football statistics. Arsenal won this year's EPL largely thanks to their effectiveness, and having a dead ball specialist like James Ward-Prowse (marked in the right scattergram with an “X”) can significantly enhance one's chance creation abilities. With a tally of 5 in open play, his contributions are notably augmented by 18 from set pieces. However, most players don’t have the opportunity to take free kicks or corners, which can skew the statistics in favor of those who do. It's worth noting that chance creation through set pieces requires a different skill set compared to what playmakers exhibit in open play. Therefore, we will focus our analysis on CC in open play, leaving set pieces for another discussion.
The left scatterplot gained attention on social media and TV, but reworking the data as in the right scatterplot proves to be more instructive. BF led with the most CCs in open play, followed closely by Enzo Fernandez, who benefitted from more minutes than BF. Additionally, BF also accumulated more CCs from set pieces than anyone else, with Dominic Szoboszlai being the next closest player.
We build our model of the many using robust regression (specifically, Tukey’s biweight regression with c = 4.685, with hc3 correction for heteroskedasticity) for each league in turn to estimate the chance creation we should expect a player to typically achieve, given his number of minutes, number of passes in the final third, number of crosses, and number of passes not in the final third – all publicly available from Opta Analyst as part of football statistics. All players played at least 450 minutes. The playmaker’s job is to turn these opportunities into chances created. If a player performs this task better than average, he will have a positive regression residual (an excess of chances above expected). We use this residual to measure players’ creativity in each of the 9 leagues: The Big Five (E1, F1, G1, I1, S1 for EPL, Ligue 1, Bundesliga, Serie A, and La Liga) and the other English Football Leagues below E1 (= EPL). These include E2, E3, E4, plus E3 in the 2024-25 season, noted as E3a.

This graph illustrates the smoothed distributions for all 9 leagues, showcasing various football statistics. Far outliers, as defined by Tukey, are marked as filled black circles; these outliers lie greater than 3 x IQR above Q3 (approximately 3.5 x IQR above the median) and similarly for those below the median. These far outliers are significantly more extreme than the typical outliers found in Tukey boxplots, which are defined as lying 1.5 x IQR above (or below) Q3 (Q1). Given that leagues have different numbers of games (with a maximum of 46 and a minimum of 34), the raw, expected, and excess chance creation values will naturally differ in magnitude. By measuring Excess CC in league-specific IQR units, as shown in the graph, we can compare across leagues. Evidently, Bruno Fernandes (highlighted in red) stands out as the most extreme playmaker in all leagues. Although his team certainly affords him more opportunities than most to create chances, when we adjust for this boost, BF is still remarkably efficient at converting those opportunities into tangible chances. In the forthcoming section, we will present the top 10 players in each league, ranked by Excess CC, with far outliers highlighted in green.


The top half of the table (green) displays the regression coefficients, while the bottom half (blue) shows the corresponding t-values. These coefficients have been scaled for easier viewing and are to be interpreted as follows: if a player competing in the EPL (English1) plays for 1000 minutes, makes 100 passes in the final third, 100 crosses, and 100 crosses NOT in the final third, we expect him to achieve significant chance creation of: 3.601 + 5.259 + 4.500 – 1.251 + 0.807 = 12.915 chances.
If instead he makes 200 passes in the final third, the potential for chance creation increases by an additional 5.259 chances, and so forth. It is important to note that minutes played do not substantially affect chance creation, and an increase in passes not in the final third correlates with a decrease in expected chances. The R2 goodness of fit for the football statistics across the 9 leagues models were: 0.816, 0.859, 0.799, 0.785, 0.822, 0.859, 0.826, 0.824, 0.837, respectively, highlighting the impact of playmakers in generating scoring opportunities.
The coefficients (highlighted in green) across the leagues are remarkably similar, especially when examining chance creation in football statistics. To appreciate just how similar, consider the implications of mistakenly applying F1’s regression coefficients to E2 players. How much difference would it make to the chance creation estimates for each player? The table below shows all the Pearson correlation coefficients of expected chance creations when a ‘wrong’ set is applied, correlated with expected chance creations using the correct set. Specifically, row 2 column 6 answers our question: when the F1 coefficients are improperly applied to E2 players, the resulting estimates show an extremely high correlation of 0.9903 with the correct estimates (E2 weight applied to E2 players). The other boxed correlation in the table illustrates the opposite scenario: wrongly applying E2 weight to F1 players. Unlike in a standard correlation matrix, these figures differ slightly.

The conclusion is that model coefficients are highly interchangeable between leagues, with the possible exception of F1 with E4 (highlighted in yellow). The highest coherence in football statistics is between E3 and E3a (pink highlights), as we might expect, since E3 and E3a represent the same league in consecutive seasons; there is also notable coherence between I1 and G1 (blue highlights). However, these are mere nuances. The very high correlation throughout the table suggests we could construct a single ‘super model’ of how chances are created through chance creation that would serve almost as well as any of the 9 league-specific sub-models. A quick and dirty way to achieve this would be to average, column-wise, the coefficients in the table above. The great similarity of E3 with E3a coefficients indicates that a one-size-fits-all set of coefficients might also work effectively, if not perpetually, at least for a few seasons.
This stable and highly portable model of football statistics allows us to clearly identify those who overperform in chance creation. Some playmakers, such as Bruno Fernandes, are given every opportunity to express their creativity, resulting in a happy synergy. Others, however, are not afforded the opportunities their talents deserve. Yet others, hidden within the negatives, are players who do not make the best use of their opportunities – even though they may create chances, they do not create enough. All these dynamics are revealed by the model.
John Doyle, Football Analytica.