
Tactical training culture is fond of rankings. Shooters are sorted by
time, instructors publish standards, agencies define pass-fail
thresholds, and equipment is compared by tenths or hundredths of a
second. Yet ranking is the last step, not the first. Before a score can
separate people or detect improvement, the test must demonstrate that
repeated observations remain interpretable. If the same performer can
receive materially different results because the protocol drifts, the
evaluator changes, the device behaves differently, or ordinary
biological variation is large, the ranking may describe noise with
impressive numerical formatting.
Reliability is the degree to which a measurement is consistent under
specified conditions. It does not mean that every repeated score must be
identical; human performance is inherently variable. It means that the
amount and structure of variation are understood well enough for the
score to support the intended decision. Morrow et al. (2011/2014)
describe reproducibility as a foundational property of useful
human-performance measurement, while Hopkins (2000) emphasizes that
sports tests must quantify within-subject variation rather than assume
that observed differences represent real change.
Several forms of reliability matter in firearms assessment.
Within-session reliability concerns how trials behave during one test.
Test-retest reliability concerns agreement across days or sessions.
Inter-rater reliability concerns whether different evaluators score the
same performance similarly. Device reliability concerns timer detection,
target measurement, and any sensing technology. Protocol reliability
concerns whether the instructions, start condition, distances, and
exclusion rules are applied consistently. A test can perform well in one
domain and poorly in another; a precise timer cannot rescue ambiguous
target scoring, and a clear scoring zone cannot rescue an improvised
start procedure.
Reliability is often confused with correlation. Two test sessions can
be highly correlated because the faster performers remain generally
faster, while individual scores differ enough to make small improvement
claims unsafe. Bland and Altman (1986) showed that association is not
the same as agreement. For coaching, agreement is usually the more
relevant question: how far apart can repeated scores for the same person
be when no meaningful change has occurred? If that natural test-retest
band is ±0.12 seconds, celebrating a 0.03-second improvement as proof of
a superior method is not justified.
The raw material of reliability is repeated data. One repetition
cannot reveal within-person variability. A five-run series begins to
expose it; a longer series estimates it more comfortably. The practical
design must balance statistical stability against fatigue, ammunition
cost, and learning during the test itself. A compact baseline can be
repeated frequently, while a larger validation session can be scheduled
periodically. What matters is that the chosen design be acknowledged as
a compromise rather than presented as an error-free measurement of an
invisible fixed ability.
Warm-up is a major source of uncontrolled variation. A performer who
begins cold may improve across the first several attempts because of
task familiarization rather than because the underlying capability
changes during the session. Conversely, an excessive warm-up can
introduce fatigue or permit rehearsal that is unavailable in the
intended operational context. A reliable protocol defines what happens
before scoring: the number and type of preparatory repetitions, whether
they are timed, and how long the performer rests. Without that rule, two
“identical” tests may estimate different states.
Start procedures also matter. An audible timer signal, visual cue,
verbal command, self-initiated start, and decision stimulus impose
different perceptual demands. Even within an audible protocol,
inconsistent delay settings or anticipatory rhythm can alter response
time. Hick (1952) and Hyman (1953) demonstrated that reaction time
depends on stimulus information and response alternatives, so a simple
known response cannot be treated as equivalent to a discrimination task.
Reliability requires holding the cue architecture constant when the test
is meant to be repeated.
Equipment should be treated as part of the measurement system.
Holster retention, belt stiffness, firearm dimensions, garment
properties, footwear, and target material can change movement and
scoring. If the question is whether the performer improved, equipment
must remain stable or changes must be documented and analyzed as
interventions. If the question is whether equipment A outperforms
equipment B, order effects must be controlled because the second
condition may benefit from practice or suffer from fatigue. Alternating
or counterbalancing conditions is more defensible than always testing
the preferred equipment last.
Timer data require particular discipline. Acoustic detection can be
affected by suppressors, adjacent shooters, echoes, sensitivity
settings, and the difference between live fire and dry-fire surrogate
signals. Video-derived timing depends on frame rate and on the
evaluator’s definition of movement onset and task completion. A
30-frame-per-second recording resolves time in steps of roughly 0.033
seconds before judgment error is considered; reporting a video-derived
result to the thousandth of a second would be numerical theater. Device
resolution, detection rule, and analysis method should be recorded with
the score.
Target scoring introduces evaluator variance. A hit that clearly
falls inside a zone is easy to classify; a line-breaking or partially
obscured hit may not be. The protocol should define boundary treatment,
target replacement, overlays, and how multiple shots are identified.
When scoring includes qualitative technical ratings—such as grip quality
or movement efficiency—anchors and examples are needed so that different
instructors apply the categories similarly. Objectivity is not achieved
by using professional vocabulary. It is achieved when another qualified
evaluator can reproduce the decision from the same evidence.
Environmental conditions can be either controlled or intentionally
varied, but they cannot be ignored. Lighting, wind, temperature, surface
traction, noise, range configuration, and social observation affect
performance. A stable indoor baseline may provide strong reliability but
limited ecological breadth. A variable field test may be more
representative but less repeatable. The solution is to state the
purpose. Controlled tests isolate change; representative tests examine
adaptation. Trying to make one score perform both functions usually
produces weak measurement and overconfident interpretation.
The reliability of a test is population- and context-specific. A
protocol that is stable for experienced shooters may be unreliable for
novices whose strategy changes quickly across trials. A course with
generous time limits may compress skilled performers near a ceiling,
making meaningful differences hard to detect. A highly complex scenario
may overwhelm beginners and create floor effects. Guadagnoli and Lee’s
(2004) challenge-point framework explains why functional difficulty
depends on both task demands and performer level. Reliability cannot be
assumed to transfer unchanged from one group to another.
Thresholds should account for measurement error. If a qualification
cut score is placed near the width of ordinary day-to-day variation, a
performer may pass or fail because of noise rather than a meaningful
difference in capability. High-stakes tests therefore deserve stronger
reliability evidence than informal coaching drills. They may require
multiple stages, repeated opportunities, standardized evaluator
training, and review procedures. The cost of measurement rigor should be
proportional to the consequence of the decision supported by the
score.
For individual training, the practical question is the smallest
change worth acting on. Statistical language sometimes calls this a
minimal detectable change or smallest worthwhile change, but the
operational definition must be tied to the task. A 0.05-second change
may matter in one competitive context and be irrelevant in another; a
small increase in valid rate may matter greatly when the initial failure
rate is unacceptable. The instructor should define improvement before
seeing the post-test, then compare the observed change with both
ordinary variation and task relevance.
The TMM Triad turns reliability from an academic ornament into a
training safeguard. Technique identifies the behavior to be expressed.
Metrics require a protocol capable of observing that behavior
consistently. Method changes the selected constraint and returns the
performer to the same measurement. If the metric is unstable, the loop
cannot distinguish an effective method from favorable noise. Bearare and
Silveira (2026) frame TMM as an operational bridge between scientific
knowledge and field pedagogy; reliability is one of the bridge’s
load-bearing elements.
An instructor can improve reliability without building a laboratory.
Write the protocol. Preserve all trials. Fix the target and scoring
rule. Standardize warm-up and rest. Record equipment and environmental
conditions. Use the same timer configuration. Train evaluators with
common examples. Repeat the test on another day. Report the raw values,
not only the best score. These steps do not eliminate uncertainty, but
they make uncertainty visible and reduce the opportunity for expectation
to rewrite the result.
Only after a test survives repetition should it be used to rank
people, compare methods, or advertise improvement. Reliability does not
prove that a test measures the right construct; that is validity’s
question. It does establish that the score is stable enough for the
validity question to be meaningful. A test that cannot agree with itself
has no authority to judge a performer. In evidence-led training,
reproducibility is not bureaucratic caution. It is the minimum respect
owed to every person whose competence, progress, or professional
standing will be represented by a number.
References
Bearare, S. C., & Silveira, L. (2026). Technique-Method-Metric
Triad in firearms training under extreme stress. RECIMA21 – Revista
Científica Multidisciplinar, 7(7), e778536.
https://doi.org/10.47820/recima21.v7i7.8536
Bland, J. M., & Altman, D. G. (1986). Statistical methods for
assessing agreement between two methods of clinical measurement. The
Lancet, 1(8476), 307–310.
https://doi.org/10.1016/S0140-6736(86)90837-8
Guadagnoli, M. A., & Lee, T. D. (2004). Challenge point: A
framework for conceptualizing the effects of various practice conditions
in motor learning. Journal of Motor Behavior, 36(2), 212–224.
https://doi.org/10.3200/JMBR.36.2.212-224
Hick, W. E. (1952). On the rate of gain of information. Quarterly
Journal of Experimental Psychology, 4(1), 11–26.
https://doi.org/10.1080/17470215208416600
Hopkins, W. G. (2000). Measures of reliability in sports medicine and
science. Sports Medicine, 30(1), 1–15.
https://doi.org/10.2165/00007256-200030010-00001
Hyman, R. (1953). Stimulus information as a determinant of reaction
time. Journal of Experimental Psychology, 45(3), 188–196.
https://doi.org/10.1037/h0056940
Morrow, J. R., Jr., Jackson, A. W., Disch, J. G., & Mood, D. P.
(2011). Measurement and evaluation in human performance (4th
ed.). Human Kinetics. [Portuguese edition: Artmed, 2014.]

Continue from research to practice
Knowledge is useful when it changes what you do next.
Use this evidence to identify the next capability you need to build and the ABA route designed for it.
Find your ABA training path

