PofoliaShared via Pofolia

Discover Artificial Intelligence· 2026Q1

Uncovering the drivers of IPO underpricing in Malaysia through machine learning benchmarking

Ali Albada, Rabie Α. Ramadan, Eimad Eldin Abusham, Boon Heng Teh et al.

Short summary

Support Vector Regression (SVR-Linear) achieved the best predictive performance (test R² = 0.512) for Malaysian IPO underpricing, outperforming ensemble methods prone to overfitting. SHAP analysis revealed the over-subscription ratio (OSR) as the strongest predictor (40.7%), followed by a knowledge gap (15.6%).

AI-generated from the title and abstract; the full text is not read.

Key points

  • SVR-Linear model achieved the highest predictive accuracy for IPO underpricing (test R² = 0.512) among seven machine learning algorithms.
  • The over-subscription ratio (OSR) was the most significant predictor of underpricing, accounting for 40.7% of the effect, followed by a knowledge gap (15.6%).
  • Intense investor demand (OSR > 30x) significantly moderates the relationship between the knowledge gap and underpricing, reducing the latter's effect by 73%.
  • SHAP analysis identified non-linear moderation effects, such as the impact of investor demand on information asymmetry, which are missed by conventional regression.

AI-generated from the title and abstract; the full text is not read.

Abstract

Abstract Purpose In this study, IPO underpricing in a fixed-price market is examined through two distinct objectives: (1) a predictive objective identifying the machine learning model that best generalizes out-of-sample, and (2) an economic interpretation objective using model-agnostic interpretability tools (SHAP) and formal statistical tests (bootstrapped hierarchical regression) to assess which economic mechanisms are consistent with the observed predictive patterns. Conventional linear models with additive feature effects are inefficient for understanding moderation effects involving interactions. This is done through two methodological innovations: (1) explicit non-linear feature engineering (interaction terms to describe moderation effects), allowing linear models to learn complex relationships, and (2) model-agnostic interpretability of tree-based ensembles through SHAP analysis to numerically quantify interaction strengths. It combines the predictive stability of regularized linear models (SVR-Linear: R 2 = 0.512) with the interpretability of non-linear models (Random Forest SHAP) for hypothesis testing. Design/methodology The data employed in this paper include 350 Malaysian IPOs of the respective years 2004–2021. It uses seven machine learning algorithms: Support Vector Regression (SVR), Random Forest, XGBoost, LightGBM, k-Nearest Neighbors, Multilayer Perceptron, and Linear Regression. The evaluated algorithms are systematically assessed using 5-fold cross-validation. In terms of feature importance, SHAP (Shapley Additive exPlanations) values are used to quantify feature importance and characterize associations regarding investor demand, offer price, and listing board. Formal hypothesis testing is conducted via bootstrapped hierarchical regression. Findings SVR-Linear offers better predictive performance (test R 2 = 0.512, CV R 2 = 0.486 ± 0.087), whereas ensemble techniques are prone to overfitting. SHAP analysis shows that the over-subscription ratio (OSR) contributes 40.7%, followed by a knowledge gap (15.6%). Hierarchical regression is used to establish strong moderating roles: investor demand is a significant moderator of the knowledge gap-underpricing relationship (β = 0.0284, p < 0.01), and intense demand (OSR > 30x) dilutes the effect of information asymmetry by 73%. There is marginal moderation in the offer price (β = 0.0198, p = 0.073) and no significant effect by listing board differences (β = 0.0089, p = 0.496). The threshold analysis shows that when investor demand is high, cascading demand effects lead to fundamental valuation concerns and patterns consistent with the predictions of Welch’s (1992) bandwagon theory. Practical implications The findings from this research can benefit a variety of stakeholders. The stockholders can benefit from this research and minimize underpricing. Furthermore, regulatory bodies will consider book-building processes for large IPOs when there is a signal of demand. Additionally, the institutional investors should avoid oversubscription. Furthermore, it can be concluded that precise targeting policies can be achieved by measuring interaction effects. Originality/value The study contributes in three ways, including (1) the first application of SHAP-based interpretability to IPO underpricing, which allows the identification of non-linear moderation effects that are not observed by conventional regression, (2) a rigorous overfitting diagnostic can be evaluated using hierarchical regression with bootstrapped confidence intervals, a feature that identifies the gap between black-box prediction and inferential rigor; and (3) formal statistical validation of SHAP results through bootstrapped hierarchical regression, a property that has not been previously available in financial machine learning studies. Explainable AI, coupled with causal inference, offers a transparent and replicable analytical framework for financial research that requires predictive validity and theoretical elucidation.

The authors' abstract, as published at the source. Discover Artificial Intelligence, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

AccountingBusiness, Management and Accounting