Skip to content
Prospect Housethe scouting desk · talent and data

Data Scouting Football: How Recruitment Analytics Works

Data & Metrics · Models · 2026-09-16
Recruitment meeting room with a long table, closed laptops and empty chairs

Data scouting football is often described as if a model finds players and a club signs them. The reality is narrower: a database removes most of the world from consideration, and a human decides what remains. Understanding that division of labour is the difference between a useful tool and an expensive one.

Departments that get it right treat the model as a filter with a known error rate, not as an oracle. They also spend as much time on data quality as on the analysis itself.

Recruitment meeting room with a plain board, closed laptops and empty chairs seen from the doorway
The database narrows the field. It does not tell anyone which of the remaining names can actually play.

Recruitment analytics: building the database

Recruitment analytics starts with coverage. A database is only as useful as the leagues it observes, and the gap between well-covered divisions and poorly covered ones is still enormous. A model built on five European leagues will happily agree with the eye about those leagues and say nothing about the rest of the world.

The second layer is identity resolution: making sure the right player is attached to the right record across competitions, and that minutes in a reserve or cup fixture are not mixed into league totals. This unglamorous work consumes more analyst time than any modelling.

The third layer is the metric set itself. Most departments begin with a standard package of per-90 rates and add a few custom measures for the roles they recruit most often, rather than building everything from scratch.

Football data analysis: the filters that remove names

Football data analysis in recruitment is mostly subtraction. A typical search starts with a position, a minimum number of minutes, an age band and a league band, and finishes with a list of perhaps forty names. Each filter has to be defensible, because a careless minutes threshold quietly deletes every player who was injured or rotated.

The order of filters matters too. Applying a league restriction first removes entire markets before style is considered; applying a style filter first usually produces a list of interesting players who are not available.

  • Minutes: at least 1,500 league minutes in the last 18 months, split by competition.
  • Age: a band tied to the squad plan, not a default value.
  • League: a band with a translation factor, not a hard border.
  • Role: position plus a behavioural filter, such as receptions in the half space.

Recruitment analytics: league translation

Recruitment analytics has to answer the standard question: what does a good season in one division mean in another? Simple translation factors based on how players have historically moved between the two leagues are the usual approach, and they are crude but better than nothing.

The honest position is that translation factors describe an average and say nothing about an individual. They also struggle with style: a league widely regarded as physical may be a good environment for one type of defender and a bad one for another.

Football data analysis: the eye test still votes

Every model output should be verified on video before a name reaches a decision maker. The reason is not that the data is wrong; it is that data cannot see intent. A defender with a poor duel win rate may be covering for a team-mate who cannot defend, and only a viewing will show that.

Where data and viewing each carry the decision
StageOwnerTypical error
Coverage and cleaningData teamMissing leagues, mixed minutes
Filtering and shortlistingAnalyst and scoutThresholds nobody can justify
Video verificationScoutConfirming the model instead of testing it
Live viewingSenior scoutToo few matches to conclude anything
RecommendationSporting directorWeighing the model above the evidence

What a model cannot price

Models handle events well and situations badly. They cannot price a player's ability to hold a defensive structure through a bad run, or his effect on a dressing room, or the way a crowd lifts when he presses. Those factors decide whether a signing works.

They also cannot price availability risk properly. A brilliant player who has missed 40 percent of the last three seasons is a different asset from one who has never been injured, which is why the football performance metrics that scouts end up trusting most are usually the boring ones about minutes and availability.

Building the loop that keeps improving

The last piece is feedback. A department that records what it predicted about a signing and what actually happened can measure its own filters, which is the only way the process improves. That record also shows which template fields were actually used in a decision and which were never read.

Over time the loop changes behaviour in predictable ways: fewer filters, better data, and much more time spent on the players already on the shortlist. The output of that loop is what a transfer target list should look like, short and argued rather than long and unread.