Bioinformatics· 2026Q1
Information Leakage in Enzyme Substrate Prediction
- 1citations
- Q1SCImago
- 2026year
Short summary
Four popular enzyme-substrate prediction models, trained on a widely used dataset, show inflated performance due to similarity-induced information leakage, performing no better than statistical baselines once leakage is controlled.
AI-generated from the title and abstract; the full text is not read.
Key points
- Popular enzyme-substrate prediction models exhibit similarity-induced information leakage.
- Controlling for leakage reveals models perform no better than statistical baselines.
- Performance degrades substantially when common or structurally similar ligands are separated between training and test sets.
- Models generalize better across enzyme sequence space than across ligand chemical space.
AI-generated from the title and abstract; the full text is not read.
Abstract
MOTIVATION: Enzymes are essential catalysts in many cellular processes. Understanding their interactions with small molecules, such as regulators, cofactors, and most importantly, substrates, is crucial for understanding the biochemical processes that occur in cells. Correctly interpreting the roles of small molecules that interact with enzymes is key to elucidating enzyme function. Recently, enzyme-small molecule interaction prediction has attracted growing interest from computational methods, especially deep learning. As a result, researchers have published several datasets and numerous models with remarkable performance. RESULTS: In this work, we critically examine one of the most popular datasets and four models trained on it, identifying similarity-induced information leakage that may overinflate reported model performance in out-of-distribution applications. We show that the inspected models are susceptible to information leakage, and their performance is not better than statistical baselines when the leakage is removed. Furthermore, controlling for leakage due to protein sequence similarity alone still yields high predictive performance, whereas performance decreases substantially when common substrates shared between training and test sets are removed from one of them, and approaches random prediction when structurally similar ligands are separated between splits. Thus, the investigated models generalize considerably better across enzyme sequence space than across ligand chemical space, indicating that their reported performance depends strongly on the small-molecule similarity between training and evaluation data. AVAILABILITY AND IMPLEMENTATION: All data splits we calculated in this study are available on Zenodo (DOI: 10.5281/zenodo.18786394). The code is available on GitHub and backed up on Zenodo (DOI: 10.5281/zenodo.18788609). SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online.
The authors' abstract, as published at the source. Bioinformatics, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Molecular Biology
Molecular BiologyBiochemistry, Genetics and Molecular Biology