PofoliaPofolia ile paylaşıldı

Bioinformatics· 2026Q1

Enzim Substrat Tahmininde Bilgi Sızıntısı

Information Leakage in Enzyme Substrate Prediction

Vahid Atabaigi Elmi, Roman Joeres, Olga V. Kalinina

Kısa özet

Yaygın olarak kullanılan bir veri kümesi üzerinde eğitilmiş dört popüler enzim-substrat tahmin modeli, sızıntı kontrol edildiğinde istatistiksel temellere göre daha iyi performans göstermeyen, benzerlik kaynaklı bilgi sızıntısı nedeniyle şişirilmiş performans göstermektedir.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Ana noktalar

  • Popüler enzim-substrat tahmin modelleri, benzerlik kaynaklı bilgi sızıntısı sergilemektedir.
  • Sızıntı kontrol edildiğinde modellerin istatistiksel temellere göre daha iyi performans göstermediği ortaya çıkmaktadır.
  • Eğitim ve test kümeleri arasında ortak veya yapısal olarak benzer ligandlar ayrıldığında performans önemli ölçüde düşmektedir.
  • Modeller, enzim dizisi uzayı boyunca ligand kimyasal uzayından daha iyi genelleme yapmaktadır.

Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.

Özet (abstract)

MOTIVATION: Enzymes are essential catalysts in many cellular processes. Understanding their interactions with small molecules, such as regulators, cofactors, and most importantly, substrates, is crucial for understanding the biochemical processes that occur in cells. Correctly interpreting the roles of small molecules that interact with enzymes is key to elucidating enzyme function. Recently, enzyme-small molecule interaction prediction has attracted growing interest from computational methods, especially deep learning. As a result, researchers have published several datasets and numerous models with remarkable performance. RESULTS: In this work, we critically examine one of the most popular datasets and four models trained on it, identifying similarity-induced information leakage that may overinflate reported model performance in out-of-distribution applications. We show that the inspected models are susceptible to information leakage, and their performance is not better than statistical baselines when the leakage is removed. Furthermore, controlling for leakage due to protein sequence similarity alone still yields high predictive performance, whereas performance decreases substantially when common substrates shared between training and test sets are removed from one of them, and approaches random prediction when structurally similar ligands are separated between splits. Thus, the investigated models generalize considerably better across enzyme sequence space than across ligand chemical space, indicating that their reported performance depends strongly on the small-molecule similarity between training and evaluation data. AVAILABILITY AND IMPLEMENTATION: All data splits we calculated in this study are available on Zenodo (DOI: 10.5281/zenodo.18786394). The code is available on GitHub and backed up on Zenodo (DOI: 10.5281/zenodo.18788609). SUPPLEMENTARY INFORMATION: Supplementary information is available at Bioinformatics online.

Yazarların özeti; kaynağından alınmıştır. Bioinformatics, 2026 · DOI ↗

ÇıkarımlarPremium
Makaleye SorÜcretsiz hesapla

Ücretsiz hesapla devam et

Makaleye Sor ile bu makaleye günde 3 soru ücretsiz; makaleyi kaydet, kaynakçasını al, ilgi alanına göre her gün yeni özetler. Çıkarımlar Premium.

Web'de ücretsiz devam et

Google ya da Apple hesabınla giriş; kart istemez. Bu makaleye geri dönersin.

Telefonda:

Alan: Moleküler Biyoloji

Molecular BiologyBiochemistry, Genetics and Molecular Biology