PofoliaShared via Pofolia

BMC Medical Research Methodology· 2026Q1

Optimization of coarsened exact matching through random forest classification

Ryo Mishima, Daisuke Koide

Short summary

Random forest classification can inform data-driven coarsening rules for Coarsened Exact Matching (CEM), preserving sample size and improving interpretability compared to standard CEM.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract Background Coarsened exact matching (CEM) was developed to resolve the “matching paradox” in propensity score matching (PSM), where narrower caliper widths can paradoxically worsen covariate balance. However, its application to complex medical data faces practical challenges, including the need for subjective decisions when categorizing continuous variables and substantial reductions in sample size when multiple variables are handled. Methods We investigated several data-driven strategies for determining coarsening rules using classification-tree models. Specifically, we explored coarsening informed by classification and regression trees (CART), random forests (RF), and the most representative tree (MRT). The MRT provides a single representative tree that can be used to summarize and interpret the splitting structure learned by RF, thereby addressing the limited interpretability of ensemble partitions. Simulation studies comparing different matching procedures were conducted across several scenarios, with varying degrees of similarity between the propensity score and outcome models, as well as varying complexities in each modeling structure. To demonstrate the practical utility of this framework, we applied it to the Framingham Heart Study dataset. Results For the tree-based CEM methods, covariate imbalance and bias generally decreased as matching became more stringent, reflecting the monotonic behavior expected from CEM-based matching. PSM often achieved the smallest bias when a large proportion of subjects was retained, but stricter propensity-score-based matching did not necessarily improve estimation performance. In the Framingham study, PSM and the CEM-based approaches produced broadly similar effect estimates; tree-based CEM preserved the full sample size, whereas standard CEM incurred sample loss. Conclusions Integrating random forest with CEM may help mitigate key limitations of traditional CEM, particularly those related to subjective coarsening choices and sample loss. By providing an interpretable data-driven coarsening strategy, CEM+MRT may offer a practical option for implementing CEM when transparent coarsening rules are desired.

The authors' abstract, as published at the source. BMC Medical Research Methodology, 2026 · DOI ↗

TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways, key points and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Statistics and Probability

Statistics and ProbabilityMathematics