PofoliaShared via Pofolia

Scientific Reports· 2026Q1

Systematic benchmarking of evaluation paradigms, safety boundaries, and clinical reasoning gaps for multimodal large language models in rehabilitation

Ping Ye, Yu Li, Mengjian Qu, Xuan Zhang et al.

Short summary

A new benchmark, ActionEval-MedBench, reveals leading multimodal large language models (MLLMs) struggle with fine-grained clinical reasoning in rehabilitation videos, scoring high (91.1% F1) on action recognition but poorly (below 50% F1) on identifying pathological compensation and clinical decision-making.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract Motor rehabilitation and clinical biomechanical assessment rely on expert judgement to evaluate movement quality, pathological compensation, and spatiotemporal motor dynamics. Existing automated approaches mainly depend on conventional computer vision methods, such as skeletal keypoint-based analysis, but they provide limited clinical semantic interpretation and mechanistic reasoning. With the emergence of multimodal large language models (MLLMs), movement assessment may shift from geometric pose tracking toward end-to-end visual clinical reasoning. However, the ability of current MLLMs to identify fine-grained kinematic abnormalities and suppress visually unsupported hallucinations in rehabilitation videos remains insufficiently evaluated. This study aimed to develop a multidimensional video benchmark, ActionEval-MedBench , to systematically evaluate the clinical reasoning capability, safety boundaries, and failure modes of MLLMs in motor rehabilitation and clinical biomechanics. ActionEval-MedBench included 522 rigorously screened and de-identified videos, comprising 328 clinical upper-limb rehabilitation videos and 194 general functional biomechanics videos. Each video was paired with six multi-select multiple-choice questions, yielding 3132 video—MCQ assessment items. We designed a six-dimensional ActionEval protocol covering semantic action recognition, kinematic feature alignment, pathological compensation identification, spatiotemporal dynamics analysis, comprehensive movement quality and clinical decision-making, and visual grounding for anti-hallucination assessment. Expert consensus labels were established through a dual-track Delphi-style procedure. Twenty-six leading MLLMs were evaluated under standardized zero-shot prompts and structured response-parsing rules. The primary outcome was the cohort-weighted ActionEval composite score, while the D6 false-positive hallucination rate ( $$FP_{D6}$$ ) was used as an independent safety endpoint. Among 81,432 expected dimension-level response opportunities, 81,006 responses were successfully returned, whereas 426 were unavailable because of missing outputs or system-level failures. Among the returned responses, 80,931 were parsable and valid, while 75 were invalid or unparsable, yielding an overall valid-response rate of 99.38%. Gemini-3.1-Pro-Preview, Doubao-Seed-2.0-Lite-260428, Gemini-3.5-Flash, Claude-Opus-4-7, GLM-5V-Turbo, and GPT-5.5 formed a numerically high-performing cluster, with weighted ActionEval scores ranging from 66.17 to 65.46. The Friedman test showed significant overall performance differences among the 26 models ( $$\chi ^2 = 2671.34$$ , $$P < 0.001$$ ), whereas Holm–Bonferroni-adjusted pairwise comparisons within the high-performing cluster were not statistically significant (all $$P > 0.05$$ ). Dimension-wise analysis revealed a marked semantic-to-clinical reasoning gap: models achieved strong performance in D1 semantic action recognition, with a mean F1-score of 91.1%, but declined substantially in higher-order clinical dimensions involving pathological compensation, spatiotemporal dynamics, and clinical decision-making, with mean F1-scores below 50% across D3–D5. Performance was lower in the clinical rehabilitation cohort than in the general functional biomechanics cohort, with an average drop of 9.44 percentage points; this difference was interpreted as a combined domain-shift burden between standardized functional videos and real-world fine-grained clinical rehabilitation videos rather than the isolated effect of pathology. In D6 negative-probe testing, nine models produced visually unsupported affirmative selections, with the highest $$FP_{D6}$$ reaching 6.51%. ActionEval-MedBench quantifies the cognitive boundaries and safety risks of current MLLMs in dynamic rehabilitation video understanding. Although MLLMs show promise for structured preliminary screening and clinical semantic reasoning, their limitations in detecting subtle pathological compensation and modeling spatiotemporal dynamics preclude their use as independent diagnostic tools. Future digital rehabilitation AI should follow a human-in-the-loop paradigm and incorporate domain-specific causal reasoning, kinematic-chain modeling, and rigorous visual-grounding safety probes.

The authors' abstract, as published at the source. Scientific Reports, 2026 · DOI ↗

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Rehabilitation

RehabilitationMedicine