Journal of Chemical Information and Modeling· 2026Q1
Language of Toxicity: An eXplainable Artificial Intelligence Approach
- 0citations
- Q1SCImago
- 2026year
Short summary
A descriptor-free, language-inspired AI model using CNN, GRU, and attention mechanisms accurately predicts chemical toxicity with an average AUC of 0.83 on small, imbalanced datasets.
AI-generated from the title and abstract; the full text is not read.
Key points
- Developed a descriptor-free AI model for toxicity prediction using CNN, GRU, and attention mechanisms.
- Treated canonical SMILES strings of toxic/nontoxic chemicals as 'words' in two distinct 'languages'.
- Achieved an average AUC of 0.83 (range 0.70-0.94) across eight toxicity endpoints.
- Model demonstrates interpretability by highlighting relevant molecular substructures influencing predictions.
AI-generated from the title and abstract; the full text is not read.
Abstract
Abstract Toxicity prediction in small molecules represents a fundamental challenge in drug development and chemical safety assessment. Traditional approaches heavily rely on predefined molecular descriptors or fingerprints, potentially limiting the ability to capture complex and nonlinear structure–activity relationships. Here, we present a descriptor-free, language-inspired framework that can be applied to different toxicity prediction tasks within a unified architecture. The model proposed combines a multiscale Convolutional Neural Network (CNN) layer to capture chemical patterns at different scales and a Gated Recurrent Unit (GRU) layer to capture the sequential nature of these patterns. This architecture also exploits an attention mechanism that computes attention weights across the sequence, enabling the model to focus on the most relevant molecular substructures for toxicity prediction. Toxic and nontoxic chemicals, represented by canonical SMILES, are investigated as the words of two languages which have to be discriminated; using eight different end points, the model provided an accurate description of toxicity patterns, with an average Area under the ROC curve (AUC) of 0.83 (min: 0.70, max: 0.94) under repeated cross-validation. The models were trained on relatively small data sets (∼1000 samples) and often strongly imbalanced, two important challenges that highlight the opportunities for future improvement; moreover, the proposed attention-based framework offers a representation of the molecular regions influencing model predictions, providing a basis for future investigations into toxicity-related structural patterns and potentially supporting hypothesis generation in drug design or drug repurposing applications.
The authors' abstract, as published at the source. Journal of Chemical Information and Modeling, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Computational Theory and Mathematics
Computational Theory and MathematicsComputer Science