← Back to ABA Research

Measurement and TMM

Validity Before Tradition: Does Your Qualification Measure the Skill It Claims?

A qualification target viewed through a diagnostic lens that separates score, skill, and operational claim.

A qualification course is often treated as self-validating. It uses
firearms, targets, time limits, and formal scoring; therefore, the score
is assumed to represent armed competence. That conclusion does not
follow. A test can be carefully administered, difficult to pass, and
highly repeatable while measuring only a narrow subset of the capability
invoked by its title. If a course is called a “combat shooting
qualification,” the word combat creates claims about
perception, decision, contextual adaptation, and performance under
pressure that a predictable target sequence may never observe.

Validity concerns the interpretation and use of a score. The question
is not whether the test looks professional or whether experienced
instructors respect it. The question is whether available evidence
supports the claim made from the result. Morrow et al. (2011/2014)
distinguish content, criterion, and construct evidence as related ways
of asking whether a measurement reflects what it is supposed to reflect.
Modern validity thinking also examines the consequences of score use,
because even a technically sound measure can be misapplied to a decision
it was never designed to support.

The process begins by defining the construct. “Firearms skill” is too
broad to test directly. It may include safe handling, mechanical
operation, accuracy, time-constrained movement, target discrimination,
judgment, communication, use of cover, adaptation to equipment, and
regulation under stress. A single course cannot sample every component
with equal depth. The test designer must state which components are
included, which are excluded, and what inference a passing score
permits. Ambiguity at this stage cannot be corrected by more elaborate
scoring later.

Content evidence asks whether the tasks adequately sample the defined
domain. If the stated construct includes decision-making but every
target is known in advance, the content is incomplete. If the construct
includes concealed access but the test begins with the pistol already in
hand, access is absent. If safe handling is essential but errors are
ignored unless a shot misses, the scoring system underrepresents safety.
Content validity is partly a logical judgment informed by job or task
analysis; it is not established because the course contains many rounds
or multiple distances.

Criterion evidence asks how the score relates to an external measure
that has defensible relevance. A new compact assessment might be
compared with a longer established protocol, instructor ratings, or
later performance on a representative task. Correlation alone does not
prove interchangeability, and an established test is not automatically a
gold standard. Bland and Altman (1986) warned that association and
agreement are different questions. If both tests share the same blind
spot, their correlation merely shows that they repeat the same
limitation.

Construct evidence asks whether the score behaves as theory predicts.
Experienced performers should generally outperform novices on a test of
developed skill, but expertise differences alone are insufficient.
Manipulations that increase relevant difficulty should affect
performance in explainable ways. Training that targets the construct
should change the score more than unrelated training. Measures intended
to capture distinct capabilities should not collapse into one
indistinguishable total. Construct validation accumulates through
multiple observations; it is not a stamp applied once by the test
author.

Target size and distance illustrate how hidden design choices shape
the construct. Fitts (1954) demonstrated that movement time changes with
target width and movement amplitude. A generous target under a
permissive time limit may emphasize safe completion; a small zone under
a severe limit may emphasize speed-accuracy calibration. Neither is
inherently superior. The validity problem appears when the test designer
labels one configuration as universal competence without explaining
which performance tradeoff the configuration rewards.

Predictability is another critical variable. A known sequence can
assess execution under standardized conditions with strong reliability.
It is less suited to measuring stimulus discrimination, response
selection, or adaptation to uncertainty. Hick (1952) and Hyman (1953)
showed that response time is affected by the amount and probability of
stimulus information. A draw to a known target after a known signal is
not equivalent to deciding whether a response is required. The two tasks
can share the same firearm and still measure materially different
processes.

Stress cannot be added as theatrical decoration to rescue weak
validity. Noise, shouting, physical exertion, or punitive coaching may
increase arousal without reproducing the information demands of the
intended environment. Nieuwenhuys and Oudejans (2010) found that anxiety
altered accuracy, movement, orientation, and blink behavior in a small
police sample, while Oudejans (2008) showed that practice under
representative pressure could reduce performance degradation. These
studies support carefully controlled pressure exposure, not the
assumption that any unpleasant drill measures operational readiness.

The distinction between qualification and diagnosis is also
essential. A qualification asks whether a minimum requirement was met. A
diagnostic assessment asks why performance appears as it does and which
variable should be trained next. Pass-fail scoring can be appropriate
for the first purpose while being too coarse for the second. Conversely,
a detailed technical scorecard may guide coaching without justifying a
professional certification. One course can support multiple uses only
when evidence exists for each interpretation.

Consequences matter because misclassification has costs. A false pass
can place an underprepared person in a role or encourage unwarranted
confidence. A false fail can remove a capable person, damage employment,
or direct training toward the wrong limitation. Borderline scores
deserve particular care when measurement error is substantial.
High-stakes decisions may require repeated testing, multiple components,
trained evaluators, and an appeals process rather than one aggregate
number obtained on one day.

Transfer is frequently claimed and rarely measured. A performer may
improve on the exact qualification because of sequence memory, pacing,
and course-specific strategy. That improvement is real performance on
the test, but it is not necessarily generalized learning. Transfer
requires a novel task that preserves the relevant information-movement
relationship while changing superficial details. The more ambitious the
operational claim, the more representative and independent the transfer
test must be.

Introduction to Combat Shooting treats technique as embedded
in application rather than as an isolated choreography (Silveira, 2023).
The TMM Triad extends that position by requiring the metric to match the
capability and the method to respond to diagnosed evidence (Bearare
& Silveira, 2026). A qualification that measures only accuracy from
a ready position may still have value, but TMM requires honest naming:
it is evidence about that task, not proof of comprehensive
readiness.

Test designers should resist the prestige of tradition. A long-used
course may possess valuable longitudinal data and administrative
familiarity. It may also reflect outdated equipment, old task
assumptions, ceiling effects, or inherited scoring decisions no one can
defend. Longevity is evidence of use, not evidence of validity. The
correct response is not automatic rejection but periodic validation:
task analysis, score review, reliability study, subgroup analysis, and
comparison with representative performance.

Instructors can improve validity by writing an explicit claim before
designing the drill. Define the intended performer and context. Break
the construct into observable components. Select tasks that sample those
components. Specify what the score will and will not mean. Pilot the
protocol, examine reliability, and test whether known contrasts behave
as expected. Add a transfer measure. Review errors, not only totals.
This sequence is slower than copying a familiar course, but it creates
an assessment whose authority comes from evidence rather than
costume.

The final question should be uncomfortable: what result would
demonstrate that this qualification does not measure what we say it
measures? If the answer is “none,” the course is protected doctrine, not
a test. Nullius in Verba requires a measurement system capable
of disappointing its designer. Validity is the discipline of limiting a
claim to the evidence that supports it—and then improving both the test
and the training when the evidence is not enough.

References

Bearare, S. C., & Silveira, L. (2026). Technique-Method-Metric
Triad in firearms training under extreme stress. RECIMA21 – Revista
Científica Multidisciplinar, 7
(7), e778536.
https://doi.org/10.47820/recima21.v7i7.8536

Bland, J. M., & Altman, D. G. (1986). Statistical methods for
assessing agreement between two methods of clinical measurement. The
Lancet, 1
(8476), 307–310.
https://doi.org/10.1016/S0140-6736(86)90837-8

Fitts, P. M. (1954). The information capacity of the human motor
system in controlling the amplitude of movement. Journal of
Experimental Psychology, 47
(6), 381–391.
https://doi.org/10.1037/h0055392

Hick, W. E. (1952). On the rate of gain of information. Quarterly
Journal of Experimental Psychology, 4
(1), 11–26.
https://doi.org/10.1080/17470215208416600

Hyman, R. (1953). Stimulus information as a determinant of reaction
time. Journal of Experimental Psychology, 45(3), 188–196.
https://doi.org/10.1037/h0056940

Morrow, J. R., Jr., Jackson, A. W., Disch, J. G., & Mood, D. P.
(2011). Measurement and evaluation in human performance (4th
ed.). Human Kinetics. [Portuguese edition: Artmed, 2014.]

Nieuwenhuys, A., & Oudejans, R. R. D. (2010). Effects of anxiety
on handgun shooting behavior of police officers: A pilot study.
Anxiety, Stress, & Coping, 23(2), 225–233.
https://doi.org/10.1080/10615800902977494

Oudejans, R. R. D. (2008). Reality-based practice under pressure
improves handgun shooting performance of police officers.
Ergonomics, 51(3), 261–273.
https://doi.org/10.1080/00140130701577435

Silveira, L. M. da. (2023). Introduction to combat shooting:
Scientific foundations, training, and application for instructors and
trainees
. Editora CRV. https://doi.org/10.24824/978652514835.9

Article-specific visual synthesis. Consult the article for context, limitations, and complete references.

Continue from research to practice

Knowledge is useful when it changes what you do next.

Use this evidence to identify the next capability you need to build and the ABA route designed for it.

Find your ABA training path