Scientific Reports· 2026Q1
Geleneksel Çin Tıbbı ve yüz görüntülerini birleştiren açıklanabilir çok modlu derin öğrenme füzyonu ile invaziv olmayan koroner kalp hastalığı taraması
Explainable multimodal deep learning fusion of tongue and facial images for noninvasive coronary heart disease screening
- 0atıf
- Q1SCImago
- 2026yıl
Kısa özet
Dil ve yüz görüntülerini birleştiren çok modlu bir derin öğrenme modeli, invaziv olmayan koroner kalp hastalığı (KKH) taraması için 0.9992 AUC ve 0.9756 F1 skoru elde ederek, tek modlu dil (F1=0.9524) veya yüz (F1=0.9250) modellerini önemli ölçüde geride bıraktı.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- Dil ve yüz görüntülerini birleştiren çok modlu bir derin öğrenme modeli, KKH taraması için 0.9992 AUC ve 0.9756 F1 skoru elde etti.
- Çok modlu model, tek modlu dil (F1=0.9524) ve yüz (F1=0.9250) modellerinden önemli ölçüde daha iyi performans gösterdi.
- Özellik füzyonu için geç birleştirme ile derin öğrenme omurgaları (ResNet50, EfficientNet-B1) kullanıldı.
- Grad-CAM görselleştirmesi, geleneksel Çin tıbbı öğretileriyle uyumlu karar verme bölgelerini doğruladı.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
Abstract There is an urgent need for convenient, non‑invasive and low‑cost tools for early screening of coronary heart disease (CHD). While tongue and facial features are known to reflect cardiac pathology and systemic blood circulation, respectively, in traditional Chinese medicine (TCM), it remains unclear whether fusing these two modalities can meaningfully improve CHD detection. We enrolled 200 CHD patients and 318 healthy controls and acquired standardized paired tongue and facial images. A lightweight U-Net (LightUNet) was built for tongue segmentation, and MediaPipe FaceMesh was used to standardize facial photographs. Two deep learning backbones, ResNet50 and EfficientNet-B1, were arranged in a dual-branch architecture with late concatenation to fuse the extracted features. Ablation experiments and 5-fold stratified cross-validation were conducted to quantify the gain from multimodality, and Grad-CAM was employed to visualize the decision evidence. On the independent test set, the ResNet50 multimodal model achieved an AUC of 0.9992 and an F1 score of 0.9756, which clearly exceeded the tongue-only (F1 = 0.9524) and face-only (F1 = 0.9250) baselines. 5-fold cross-validation confirmed the stability of these gains, yielding AUCs of 0.9980 ± 0.0027 for ResNet50 and 0.9987 ± 0.0014 for EfficientNet-B1, both with very small inter-fold variance. In the Grad-CAM heatmaps, the tongue branch consistently activated on the central tongue body and coating, while the facial branch focused on the perioral and zygomatic-buccal regions. This spatial pattern aligns well with the TCM tenets that “the tongue is the sprout of the heart” and “the heart, its efflorescence is in the face”. The dual-modal fusion of tongue and facial images significantly outperforms unimodal analysis and offers an efficient, interpretable computational approach for non-invasive CHD screening.
Yazarların özeti; kaynağından alınmıştır. Scientific Reports, 2026 · DOI ↗
Devamı Pofolia uygulamasında
Çıkarımlar ve makaleye soru sorma; ilgi alanına göre her gün yeni özetler. Ücretsiz.
Web'de giriş yaparak açAlan: Tamamlayıcı ve Alternatif Tıp
Complementary and alternative medicineMedicine