Scientific Reports· 2026Q1
Explainable multimodal deep learning fusion of tongue and facial images for noninvasive coronary heart disease screening
- 0citations
- Q1SCImago
- 2026year
Short summary
A multimodal deep learning model fusing tongue and facial images achieved an AUC of 0.9992 and F1 score of 0.9756 for non-invasive coronary heart disease (CHD) screening, significantly outperforming unimodal tongue (F1=0.9524) or face (F1=0.9250) models.
AI-generated from the title and abstract; the full text is not read.
Key points
- A multimodal deep learning model fusing tongue and facial images achieved an AUC of 0.9992 and F1 score of 0.9756 for CHD screening.
- The multimodal model significantly outperformed unimodal tongue (F1=0.9524) and face (F1=0.9250) models.
- Deep learning backbones (ResNet50, EfficientNet-B1) were used with late concatenation for feature fusion.
- Grad-CAM visualization confirmed decision-making regions aligned with traditional Chinese medicine tenets.
AI-generated from the title and abstract; the full text is not read.
Abstract
Abstract There is an urgent need for convenient, non‑invasive and low‑cost tools for early screening of coronary heart disease (CHD). While tongue and facial features are known to reflect cardiac pathology and systemic blood circulation, respectively, in traditional Chinese medicine (TCM), it remains unclear whether fusing these two modalities can meaningfully improve CHD detection. We enrolled 200 CHD patients and 318 healthy controls and acquired standardized paired tongue and facial images. A lightweight U-Net (LightUNet) was built for tongue segmentation, and MediaPipe FaceMesh was used to standardize facial photographs. Two deep learning backbones, ResNet50 and EfficientNet-B1, were arranged in a dual-branch architecture with late concatenation to fuse the extracted features. Ablation experiments and 5-fold stratified cross-validation were conducted to quantify the gain from multimodality, and Grad-CAM was employed to visualize the decision evidence. On the independent test set, the ResNet50 multimodal model achieved an AUC of 0.9992 and an F1 score of 0.9756, which clearly exceeded the tongue-only (F1 = 0.9524) and face-only (F1 = 0.9250) baselines. 5-fold cross-validation confirmed the stability of these gains, yielding AUCs of 0.9980 ± 0.0027 for ResNet50 and 0.9987 ± 0.0014 for EfficientNet-B1, both with very small inter-fold variance. In the Grad-CAM heatmaps, the tongue branch consistently activated on the central tongue body and coating, while the facial branch focused on the perioral and zygomatic-buccal regions. This spatial pattern aligns well with the TCM tenets that “the tongue is the sprout of the heart” and “the heart, its efflorescence is in the face”. The dual-modal fusion of tongue and facial images significantly outperforms unimodal analysis and offers an efficient, interpretable computational approach for non-invasive CHD screening.
The authors' abstract, as published at the source. Scientific Reports, 2026 · DOI ↗
The rest is in the Pofolia app
Takeaways and questions to the paper; new summaries every day for your field. Free.
Sign in on the web to openField: Complementary and alternative medicine
Complementary and alternative medicineMedicine