PofoliaShared via Pofolia

Nature Communications· 2026Q1

Scalable molecular representations enabled by multimodal fusion and sequence distillation

Suvendu Kumar, Saveena Solanki, Mudit Gupta, Sonam Chauhan et al.

Short summary

A new 'Chemical Dice Integrator' unifies six molecular views into a single latent space, distilled into a sequence-based model for direct embedding from molecular strings, outperforming classical fusion and established descriptors in prediction benchmarks.

AI-generated from the title and abstract; the full text is not read.

Key points

  • The Chemical Dice Integrator fuses six distinct molecular views into a unified latent space.
  • This multimodal representation is distilled into a sequence-based model for direct embedding from molecular strings.
  • The new representation outperformed classical fusion methods and matched/improved established molecular descriptors on benchmarks.
  • It showed stable generalization and improved utility under limited data conditions.
  • The framework successfully prioritized compounds (isoeugenol, eugenyl acetate) that experimentally reduced genome instability in yeast.

AI-generated from the title and abstract; the full text is not read.

Abstract

Molecular prediction depends on how chemical structures are represented, yet descriptors capture only partial aspects of chemical information. Here, we show that the Chemical Dice Integrator combines six complementary molecular views spanning physicochemical properties, molecular topology, two-dimensional structural images, bioactivity profiles, quantum properties, and molecular language into a unified latent space. This multimodal representation is distilled into a sequence-based model that generates the embedding directly from molecular strings. Across classification and regression benchmarks, the representation outperformed classical fusion methods, matched or improved established molecular descriptors, retained complete embedding coverage when individual feature-generation pipelines failed, and supported efficient inference. Scaffold-based and low-data evaluations showed stable generalization across chemical space and improved utility under limited data. The framework also prioritized compounds predicted to protect genome stability, leading to experimental validation of isoeugenol and eugenyl acetate in a yeast damage-response assay. These findings establish a scalable framework for molecular prediction and discovery. This study integrates six complementary molecular representations into a distilled molecular string embedding, enabling scalable property prediction and prioritization of compounds that reduce DNA-damage phenotypes in yeast.

The authors' abstract, as published at the source. Nature Communications, 2026 · DOI ↗

TakeawaysIn the app
Ask the paperIn the app

The rest is in the Pofolia app

Takeaways and questions to the paper; new summaries every day for your field. Free.

Sign in on the web to open

Field: Computational Theory and Mathematics

Computational Theory and MathematicsComputer Science