Key papers in Statistics and Probability

Pofolia’s corpus holds 79 papers from the Statistics and Probability subfield (2012–2025). The list below starts with the most cited.

Most cited

Ranked by citation count. Because citations accumulate over time, this list naturally leans towards work published a few years ago; for where the field is now, see “recently added”.

  • <b>lmerTest</b> Package: Tests in Linear Mixed Effects Models

    Journal of Statistical Software · 2017 · Q1 · SJR 3.00 · FWCI 1076.75 · 23,710 citations · Open access

    The lmerTest R package now provides p-values for F and t tests in linear mixed effects models, addressing a common user need for the lme4 package's lmer function.

    Go to source

  • Applied Multiple Regression/Correlation Analysis for the Behavioral Sciences

    Psychology Press eBooks · 2014 · FWCI 232.26 · 20,892 citations

    This classic text offers a non-mathematical, applied approach to multiple regression analysis, focusing on verbal-conceptual exposition and practical data analysis for behavioral sciences.

    Go to source

  • <tt>emcee</tt>: The MCMC Hammer

    Publications of the Astronomical Society of the Pacific · 2013 · Q1 · SJR 2.00 · FWCI 219.66 · 11,975 citations · Open access

    A new Python implementation of the affine-invariant ensemble sampler for Markov chain Monte Carlo (MCMC) called emcee is introduced, offering improved performance and requiring fewer parameters to tune than traditional methods.

    Go to source

  • A Mathematical Theory of Evidence

    Princeton University Press eBooks · 2020 · FWCI 114.25 · 11,892 citations

    This book constructs a new theory of epistemic probability by building on an abstract understanding of how facts are combined, diverging from Dempster's viewpoint by identifying his 'upper and lower probabilities' as fundamental.

    Go to source

  • Applied Multivariate Statistics for the Social Sciences

    2012 · FWCI 147.54 · 10,705 citations

    This is a textbook designed for social science students and researchers who need to apply advanced statistical methods without necessarily developing them.

    Go to source

  • Correlation Coefficients: Appropriate Use and Interpretation

    Anesthesia & Analgesia · 2018 · Q1 · SJR 1.00 · FWCI 341.12 · 10,519 citations

    This tutorial guides researchers and clinicians on the appropriate use and interpretation of correlation coefficients, specifically Pearson and Spearman, to measure associations between variables.

    Go to source

  • <b>brms</b> : An <i>R</i> Package for Bayesian Multilevel Models Using <i>Stan</i>

    Journal of Statistical Software · 2017 · Q1 · SJR 3.00 · FWCI 319.75 · 9,492 citations · Open access

    The brms package in R provides a flexible framework for fitting complex Bayesian multilevel models using the Stan probabilistic programming language.

    Go to source

  • emmeans: Estimated Marginal Means, aka Least-Squares Means

    2017 · 8,335 citations

    The 'emmeans' R package provides a unified framework for obtaining estimated marginal means (EMMs) and performing post-hoc analyses for a wide range of linear, generalized linear, and mixed models.

    Go to source

  • Generalized Additive Models

    2017 · FWCI 110.84 · 8,008 citations

    This book serves as a leading, introductory reference for Generalized Additive Models (GAMs), offering practical examples and software implementation.

    Go to source

  • <i>Stan</i> : A Probabilistic Programming Language

    Journal of Statistical Software · 2017 · Q1 · SJR 3.00 · FWCI 540.93 · 7,425 citations · Open access

    Stan is a new probabilistic programming language designed for specifying statistical models and performing Bayesian inference. It offers full Bayesian inference via advanced Markov chain Monte Carlo methods and penalized maximum likelihood estimates through optimization algorithms.

    Go to source

  • Root mean square error (RMSE) or mean absolute error (MAE)? – Arguments against avoiding RMSE in the literature

    Geoscientific model development · 2014 · Q1 · SJR 2.00 · FWCI 133.29 · 6,160 citations · Open access

    This paper argues against the proposed avoidance of Root Mean Square Error (RMSE) in favor of Mean Absolute Error (MAE) for model evaluation.

    Go to source

  • Sensitivity Analysis in Observational Research: Introducing the E-Value

    Annals of Internal Medicine · 2017 · Q1 · SJR 2.00 · FWCI 200.14 · 5,818 citations

    The E-value quantifies the minimum strength an unmeasured confounder needs to explain away an observed association in observational studies, offering a standardized way to assess robustness to confounding.

    Go to source

  • Testing Statistical Hypotheses

    Wiley series in probability and statistics · 2021 · 5,240 citations

    This chapter reviews established concepts and results in statistical hypothesis testing, focusing on generalized likelihood ratio tests and conditional tests for situations with nuisance parameters.

    Go to source

  • The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation

    PeerJ Computer Science · 2021 · Q2 · FWCI 531.97 · 4,961 citations

    The coefficient of determination (R-squared) is a more informative and truthful metric for evaluating regression analysis than SMAPE, MSE, RMSE, MAE, and MAPE.

    Go to source

  • Regression Models for Categorical Dependent Variables Using Stata

    2014 · FWCI 375.41 · 4,670 citations

    This book serves as a comprehensive guide to regression models for categorical dependent variables, specifically within the Stata software environment.

  • Two-Way Fixed Effects Estimators with Heterogeneous Treatment Effects

    American Economic Review · 2020 · Q1 · SJR 23.00 · FWCI 322.07 · 4,646 citations

    Standard linear regressions with period and group fixed effects can produce misleading treatment effect estimates, even showing negative coefficients when all underlying effects are positive.

    Go to source

  • A Practitioner’s Guide to Cluster-Robust Inference

    The Journal of Human Resources · 2015 · Q1 · SJR 6.00 · FWCI 344.55 · 4,589 citations

    When data is grouped into clusters, standard statistical error calculations can overestimate precision. This guide explains how to use cluster-robust standard errors for more accurate inference, especially with a large number of clusters.

    Go to source

  • Moving towards best practice when using inverse probability of treatment weighting (IPTW) using the propensity score to estimate causal treatment effects in observational studies

    Statistics in Medicine · 2015 · Q1 · SJR 1.00 · FWCI 121.94 · 4,349 citations · Open access

    A review of recent studies using inverse probability of treatment weighting (IPTW) reveals that most do not formally check for covariate balance after weighting, a critical step for valid causal effect estimation from observational data.

    Go to source

  • STAMP: statistical analysis of taxonomic and functional profiles

    Bioinformatics · 2014 · Q1 · SJR 2.00 · FWCI 110.97 · 4,341 citations · Open access

    STAMP is a new graphical software package designed for statistical analysis of taxonomic and functional profiles, offering hypothesis testing and exploratory plots.

    Go to source

  • Adding It Up: Helping Children Learn Mathematics

    DSpace Biblioteca Universidad de Talca (Universidad de Talca) · 2013 · FWCI 73.30 · 4,124 citations

    This report outlines key components of mathematical proficiency and how students develop it, recommending changes in teaching, curricula, and teacher education for pre-K through 8th grade.

    Go to source

Recently added

  • Transparent Reporting of Observational Studies Emulating a Target Trial—The TARGET Statement

    JAMA · 2025 · FWCI 239.49 · 183 citations

    The 21-item TARGET checklist provides standardized guidance for reporting observational studies that emulate a target trial, aiming to improve transparency and interpretation of causal effect estimates.

    Go to source

  • GetDist: a Python package for analysing Monte Carlo samples

    Journal of Cosmology and Astroparticle Physics · 2025 · 231 citations

    The GetDist Python package offers new tools for analyzing Monte Carlo samples, including automatic bandwidth selection for Kernel Density Estimation (KDE) and bias correction for correlated and weighted samples.

    Go to source

  • Root-mean-square error (RMSE) or mean absolute error (MAE): when to use them or not

    Geoscientific model development · 2022 · Q1 · SJR 2.00 · FWCI 315.22 · 1,795 citations · Open access

    Neither RMSE nor MAE is universally superior; RMSE is optimal for Gaussian errors, while MAE is optimal for Laplacian errors, with other metrics preferred for different error distributions.

    Go to source

  • How much should we trust staggered difference-in-differences estimates?

    Journal of Financial Economics · 2022 · Q1 · SJR 18.00 · FWCI 519.41 · 3,002 citations

    Staggered difference-in-differences (DiD) regression estimators, widely used for policy impact analysis, are often biased, leading to incorrect conclusions. This paper explains these biases and reviews three alternative estimators that address them.

    Go to source

  • Why 90% of clinical drug development fails and how to improve it?

    Acta Pharmaceutica Sinica B · 2022 · Q1 · SJR 3.00 · FWCI 295.09 · 1,627 citations

    A new framework, Structure-Tissue Exposure/Selectivity-Activity Relationship (STAR), is proposed to address the 90% failure rate in clinical drug development by rebalancing optimization from solely potency/specificity (SAR) to include tissue exposure/selectivity (STR).

    Go to source

Add this field to your daily feed

Pick your interests and new work in your area arrives every day, summarised. Full summaries live in the app.

Open the app

Other subfields in the same field

All fields