Bone and Joint Research· 2026Q1
Improving patient-reported outcome measure scoring in shoulder trials
- 0citations
- Q1SCImago
- 2026year
Short summary
Modern psychometric methods (IRT and CAT) produced trial conclusions identical to traditional CTT scoring for the Oxford Shoulder Score, while simulated CAT reduced questionnaire length by up to 75% at early follow-up.
AI-generated from the title and abstract; the full text is not read.
Key points
- Treatment effect estimates and statistical significance were highly consistent between CTT, IRT, and CAT scoring methods for the Oxford Shoulder Score in two large UK shoulder trials (n=753).
- Simulated computerized adaptive testing (CAT) reduced the median questionnaire length from the original 12 items to 4-5 items at early follow-up.
- At later timepoints, simulated CAT still reduced the median questionnaire length to 8-9 items.
- Modern psychometric methods (IRT and CAT) demonstrated potential efficiency gains and lower respondent burden compared to traditional CTT scoring.
AI-generated from the title and abstract; the full text is not read.
Abstract
Aims Patient-reported outcome measures (PROMs) underpin orthopaedic trials by capturing pain and function from the patient’s perspective. Traditionally, PROMs such as the Oxford Shoulder Score (OSS) are scored using classical test theory (CTT), which assumes equal item contribution and uniform measurement error. Item response theory (IRT) and computerized adaptive testing (CAT) offer modern alternatives that may improve precision and reduce respondent burden, but their impact on trial outcomes is unclear. Methods We reanalyzed patient-level OSS data from two major UK randomized controlled trials, UK FROST (n = 503) and PROFHER (n = 250). We used three scoring approaches: CTT sum scores, IRT-based expected a posteriori (EAP) scores, and a simulated CAT. Original statistical analysis plans were replicated to compare treatment effect estimates, CIs, and p-values across methods. CAT simulations assessed potential reductions in questionnaire length. Results Across both trials, treatment effect estimates and statistical significance were highly consistent between CTT, IRT, and CAT. Differences in magnitude were minimal, and overall trial conclusions were unchanged. Simulated CAT reduced median questionnaire length to four to five items at early follow-up, increasing to eight to nine items at later timepoints. Conclusion In these large orthopaedic trials, modern psychometric methods produced results consistent with traditional scoring, but demonstrated potential efficiency gains. IRT and CAT may facilitate lower-burden, patient-centred outcome measurement and support innovative trial designs, without compromising the interpretation of treatment effects. Cite this article: Bone Joint Res 2026;15(10):1206–1212.
The authors' abstract, as published at the source. Bone and Joint Research, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Management Science and Operations Research
Management Science and Operations ResearchDecision Sciences