PofoliaShared via Pofolia

PLoS ONE· 2026Q1

Predicting student academic performance using TabPFN and SHAP: An interpretable machine learning approach

Qiang Yin

Short summary

TabPFN combined with SHAP achieves superior predictive performance (RMSE 4.8454, R² 0.9216) for student academic performance on small datasets, outperforming traditional models and maintaining high accuracy even with only 500 samples.

AI-generated from the title and abstract; the full text is not read.

Key points

  • TabPFN combined with SHAP provides an interpretable framework for predicting student academic performance.
  • TabPFN achieved superior performance (RMSE 4.8454, R² 0.9216, MAE 3.7950) on the SAP-4000 dataset compared to traditional models.
  • TabPFN maintains high accuracy with as few as 500 samples, outperforming traditional models at 3,500 samples.
  • SHAP analysis identified weekly study hours, attendance rate, and tutoring status as crucial predictors.

AI-generated from the title and abstract; the full text is not read.

Abstract

Predicting student academic performance is of significant importance for enabling timely educational interventions and improving educational outcomes in secondary education. However, existing predictive models in educational data mining often face the dual challenges of limited interpretability and suboptimal performance on small to medium-sized datasets. This study introduces the Prior-Data Fitted Network (TabPFN) combined with SHapley Additive exPlanations (SHAP) to construct an interpretable machine learning framework for predicting student achievement. Utilizing a publicly available dataset of 4,000 Spanish high school students (SAP-4000), we employed 10-fold cross-validation to train and compare eight machine learning models. The results demonstrate that TabPFN achieves superior predictive performance, with an RMSE of 4.8454, an R 2 of 0.9216, and an MAE of 3.7950, outperforming traditional models. Comparative analyses across varying data scales indicate that TabPFN maintains high accuracy even when sample sizes are reduced to 500 samples, achieving performance comparable to or better than what traditional models attained at 3,500 samples, exhibiting stronger robustness than commonly used ensemble learning methods such as XGBoost. SHAP analysis reveals that learning behavior and study habit features—specifically weekly study hours, attendance rate, and tutoring status—play a decisive role in predicting academic performance. The proposed framework combines strong predictive performance with interpretability, overcoming the “black-box” limitation prevalent in existing algorithms. It is well-suited for small to medium-scale data modeling tasks in the educational domain, providing technical support for personalized instruction and intervention, while also offering methodological references for future research. An additional case study on a distinct student performance dataset further confirms the model’s robustness and generalizability across different educational prediction tasks.

The authors' abstract, as published at the source. PLoS ONE, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Computer Science Applications

Computer Science ApplicationsComputer Science