PofoliaShared via Pofolia

The Journal of Strength and Conditioning Research· 2026Q1

Assessing Inter-rater Reliability and Intra-rater Reliability During the 2-Minute Hand-Release Push-Up Event in the Army Fitness Test

Drew J. Van Dam, Rachel Gidaro, Sarah Ferreira, Jenna Morogiello et al.

Short summary

Raters assessing the Army's 2-minute hand-release push-up test showed inconsistent scoring, with significant fluctuations observed between and within individual raters.

AI-generated from the title and abstract; the full text is not read.

Abstract

ABSTRACT: Van Dam, D, Gidaro, R, Ferreira, S, Morogiello, J, Jaffe, D, and Semmel, A. Assessing inter-rater reliability and intra-rater reliability during the 2-minute hand-release push-up event in the army fitness test. J Strength Cond Res XX(X): 000-000, 2026-The Army Fitness Test, implemented in 2025, assesses soldier readiness via 5 physically challenging events (3 repetition maximum deadlift, hand-release push-up, sprint-drag-carry, plank, and 2-mile run). Each event is graded by individuals with varying levels of experience, thus posing a threat to the reliability of the physical performance scores. The purpose of this study was to examine the inter-rater and intra-rater reliability of the two-minute hand-release push-up test to facilitate learning and improvement of assessing physical performance. Seventeen raters (5 female, 12 male) graded 10 hand-release push-up (HRPU) videos over 3 separate weeks (30 total videos, 60 total minutes). Each video contained a recording of a Cadet performing the two-minute HRPU test. Raters were required to provide a response after each repetition by stating "yes" for a proper repetition and "no" for an improper repetition. Researchers recorded all responses from raters in each of the 3 settings. Videos were randomized for each week of viewing with no rater viewing the videos in the same order. Using a logistic model, we were able to measure inter-rater and intra-rater reliability. Researchers calculated the intraclass correlation coefficient to evaluate overall reliability and agreement and Prevalence-Adjusted, Bias-Adjusted Kappa to further understand inter-rater and intra-rater reliability. Ultimately, we found that raters were not consistent as a group, with some showing significant fluctuations in their responses between tests. This finding is crucial because even small differences in scoring can negatively impact assessment on the HRPU event influencing a subjects's total score.

The authors' abstract, as published at the source. The Journal of Strength and Conditioning Research, 2026 · DOI ↗

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Occupational TherapyHealth Professions