Science
How to validate a sports knowledge test
The four methodological problems named by Bergkamp et al. (2019) and the principles of Williams, Ford and Drust (2020), as the standard Insporta holds to.
By Insporta · updated 6 September 2026
A sports knowledge test is validated the same way any talent-identification instrument is, and the field has written down what that requires. Bergkamp et al. (2019) identified four methodological problems that undermine most football talent research: the success criterion is defined loosely, single indicators are studied in isolation, samples are restricted to already-selected players, and accuracy is reported without the base rate. Williams, Ford and Drust (2020) add that credible work must be multidisciplinary and longitudinal. Insporta treats these two papers as the standard for any claim it makes about its own tests.
Bergkamp et al. 2019: four problems to solve before claiming validity
A group of psychometricians at the University of Groningen reviewed the football talent-identification literature in Sports Medicine and found the same four weaknesses recurring.
1. Operationalisation of the criterion
What counts as success? Reaching professional level? Minutes played? Goals scored? Being retained by an academy for another season? Each definition produces a different result from the same predictor, and studies that do not fix the criterion in advance can pick the one that looks best afterwards. This is the main reason two reviews can disagree about the size of an effect, as Murr (2018) and Ivarsson (2020) do in does sport knowledge predict future level.
For a knowledge test the discipline is to state the criterion before collecting data. A natural criterion for Insporta is progress through the difficulty levels of a sport over a fixed period, six, twelve or twenty-four months, declared in advance and not changed once results are in.
2. Focus on isolated performance indicators
Most studies take a single metric. Bergkamp and colleagues argue that a multidimensional assessment across five or more variables gives far more stable results, and Höner et al. (2021) showed on 13,869 players that no single predictor beat a combination. Insporta records several dimensions per player: accuracy, response time, the profile of topics answered correctly and incorrectly, and retention over time. Any validation study should use them together rather than reporting one and ignoring the rest.
3. Restriction of range
If a study measures only academy players, as Reinhard, Mann and Höner (2025) and most others do, it sees a narrow slice of the population. Everyone in the sample has already passed a selection filter, so the spread of ability is compressed and correlations can shrink, vanish or even reverse relative to the full range. Effects estimated on academy players may not describe amateurs, and vice versa.
A public test that anyone can take has a structural advantage here: its population runs from casual fans to professionals. That is a strength only if the full range is kept in the analysis rather than trimmed to the top.
4. Base rate
Verburgh et al. (2014) classified talented and amateur players with 89% accuracy. That sounds decisive. But suppose the true success rate in a population is 5%. Out of 1,000 players, 50 will succeed. A test that is right 89% of the time will correctly flag most of those 50 and also wrongly flag about 11% of the remaining 950, which is roughly 105 players. The false positives outnumber the true positives two to one. Höner et al. (2021) found a 9% success rate in a national talent pool, so this is not a hypothetical.
Any accuracy figure Insporta reports will come with the base-rate calculation beside it.
Williams, Ford and Drust 2020: the design principles
Williams, Ford and Drust reviewed two decades of talent research in the Journal of Sports Sciences against four principles first set out in 2000, and found the field had only partly lived up to them.
| Principle | What it asks of a test like Insporta |
|---|---|
| Multidisciplinary, not single-measure | Combine knowledge with time, retention and context rather than reporting one score |
| Longitudinal, not cross-sectional | Follow the same players over time; a single snapshot cannot show development or predict it |
| Extend research to women’s football | Do not assume male findings transfer; Reinhard et al. (2025) list the absence of female data as a limitation of their own study |
| Study the subjective criteria of scouts and coaches | Record a coach’s judgment as its own variable and compare it with test results, since Höner (2021) showed it adds information |
The longitudinal principle is the one the literature most often fails. Prospective studies are expensive and slow, which is why there are so few and why Reinhard et al., with 110 players, ask for larger samples. A platform that records every test a player takes, over months and years, accumulates exactly the kind of data these authors say is missing. That is an opportunity, and also an obligation to analyse it properly.
The standard Insporta holds itself to
Put together, the two papers give a checklist that Insporta applies to any validity claim about its tests, in line with its research principles:
- The criterion is declared in advance and does not change once data are collected.
- Several dimensions are analysed together, not one headline score.
- The full range of players is kept, from beginner to professional, and results are reported by range where they differ.
- Accuracy is never reported without the base rate.
- Reliability is measured, not assumed: a longer assessment needs a test-retest check on the same group before a reliability figure is quoted, with r = .78 from Höner et al. (2023) as the benchmark.
- Age is controlled for in any comparison between players, following Vestberg (2017) and Heisler (2023).
- Data are segmented by sex so that findings for women’s sport can be reported separately rather than assumed.
None of this has been demonstrated for Insporta yet. It is the standard by which the tests will be judged when there is enough data to judge them, and the reason the science section describes what the literature supports rather than what Insporta has proven. What the evidence already says about the design is set out in how Insporta applies the research.
Sources
- Bergkamp, T. L. G., Niessen, A. S. M., den Hartigh, R. J. R., Frencken, W. G. P., & Meijer, R. R. (2019). Methodological issues in soccer talent identification research. Sports Medicine, 49(9), 1317-1335
- Williams, A. M., Ford, P. R., & Drust, B. (2020). Talent identification and development in soccer since the millennium. Journal of Sports Sciences, 38(11-12), 1199-1210
- Verburgh, L., et al. (2014). PLoS One, 9(3), e91254
- Reinhard, M. L., Mann, D. L., & Höner, O. (2025). Journal of Science and Medicine in Sport, 28(7), 587-593
Questions people ask
- What are the four methodological problems in talent identification research?
- According to Bergkamp et al. (2019): vague operationalisation of the success criterion, reliance on isolated performance indicators, restriction of range from studying only selected players, and ignoring the base rate when reporting accuracy.
- Why does an 89% accurate test still produce many false positives?
- Because of the base rate. If only 5% of a population succeeds, a classifier with 89% accuracy can label far more non-successes as successes than actual successes. Accuracy figures must be reported with the base rate.
- What is restriction of range?
- Studying only a narrow slice of the population, for example only academy players. Correlations inside that slice can be much smaller or different from those in the full range from amateur to professional.
- What do Williams, Ford and Drust recommend?
- That talent research be multidisciplinary rather than single-measure, longitudinal rather than cross-sectional, extend to women's football, and study the subjective criteria scouts and coaches actually use.