Data Scouting Football: How Recruitment Analytics Works

Data scouting football is often described as if a model finds players and a club signs them. The reality is narrower: a database removes most of the world from consideration, and a human decides what remains. Understanding that division of labour is the difference between a useful tool and an expensive one.
Departments that get it right treat the model as a filter with a known error rate, not as an oracle. They also spend as much time on data quality as on the analysis itself.

Recruitment analytics: building the database
Recruitment analytics starts with coverage. A database is only as useful as the leagues it observes, and the gap between well-covered divisions and poorly covered ones is still enormous. A model built on five European leagues will happily agree with the eye about those leagues and say nothing about the rest of the world.
The second layer is identity resolution: making sure the right player is attached to the right record across competitions, and that minutes in a reserve or cup fixture are not mixed into league totals. This unglamorous work consumes more analyst time than any modelling.
The third layer is the metric set itself. Most departments begin with a standard package of per-90 rates and add a few custom measures for the roles they recruit most often, rather than building everything from scratch.
Football data analysis: the filters that remove names
Football data analysis in recruitment is mostly subtraction. A typical search starts with a position, a minimum number of minutes, an age band and a league band, and finishes with a list of perhaps forty names. Each filter has to be defensible, because a careless minutes threshold quietly deletes every player who was injured or rotated.
The order of filters matters too. Applying a league restriction first removes entire markets before style is considered; applying a style filter first usually produces a list of interesting players who are not available.
- Minutes: at least 1,500 league minutes in the last 18 months, split by competition.
- Age: a band tied to the squad plan, not a default value.
- League: a band with a translation factor, not a hard border.
- Role: position plus a behavioural filter, such as receptions in the half space.
Recruitment analytics: league translation
Recruitment analytics has to answer the standard question: what does a good season in one division mean in another? Simple translation factors based on how players have historically moved between the two leagues are the usual approach, and they are crude but better than nothing.
The honest position is that translation factors describe an average and say nothing about an individual. They also struggle with style: a league widely regarded as physical may be a good environment for one type of defender and a bad one for another.
Football data analysis: the eye test still votes
Every model output should be verified on video before a name reaches a decision maker. The reason is not that the data is wrong; it is that data cannot see intent. A defender with a poor duel win rate may be covering for a team-mate who cannot defend, and only a viewing will show that.
| Stage | Owner | Typical error |
|---|---|---|
| Coverage and cleaning | Data team | Missing leagues, mixed minutes |
| Filtering and shortlisting | Analyst and scout | Thresholds nobody can justify |
| Video verification | Scout | Confirming the model instead of testing it |
| Live viewing | Senior scout | Too few matches to conclude anything |
| Recommendation | Sporting director | Weighing the model above the evidence |
What a model cannot price
Models handle events well and situations badly. They cannot price a player's ability to hold a defensive structure through a bad run, or his effect on a dressing room, or the way a crowd lifts when he presses. Those factors decide whether a signing works.
They also cannot price availability risk properly. A brilliant player who has missed 40 percent of the last three seasons is a different asset from one who has never been injured, which is why the football performance metrics that scouts end up trusting most are usually the boring ones about minutes and availability.
Building the loop that keeps improving
The last piece is feedback. A department that records what it predicted about a signing and what actually happened can measure its own filters, which is the only way the process improves. That record also shows which template fields were actually used in a decision and which were never read.
Over time the loop changes behaviour in predictable ways: fewer filters, better data, and much more time spent on the players already on the shortlist. The output of that loop is what a transfer target list should look like, short and argued rather than long and unread.


