ACM Transactions on Asian and Low-Resource Language Information Processing· 2026Q2
Arapça Saldırgan Konuşma Tespiti için Çok Görevli Öğrenme ve Aktif Öğrenme
Multi-task Learning with Active Learning for Arabic Offensive Speech Detection
- 1atıf
- Q2SCImago
- 2026yıl
Kısa özet
Çok görevli öğrenme (MTL) ve aktif öğrenmeyi entegre eden yeni bir çerçeve, daha az etiketli örnekle mevcut yöntemleri geride bırakarak Arapça saldırgan konuşma tespiti için %85,42'lik son teknoloji makro F1-skoru elde ediyor.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Ana noktalar
- Arapça saldırgan konuşma tespiti için çok görevli öğrenme (MTL) ve aktif öğrenmeyi entegre eden yeni bir çerçeve.
- Şiddet ve kaba konuşma görevleri üzerinde ortaklaşa eğitim, saldırgan konuşma tespitini iyileştirmek için paylaşılan temsillerden yararlanır.
- Belirsizlik örneklemesi ile aktif öğrenme, veri kıtlığını gidermek için bilgilendirici örnekleri iteratif olarak seçer.
- Önerilen yöntem, OSACT2022 veri kümesinde %85,42'lik son teknoloji makro F1-skoru elde eder.
- Çerçeve, mevcut yöntemlere kıyasla önemli ölçüde daha az ince ayar örneği kullanır.
Yapay zekâ ile başlık ve abstract'tan üretildi; tam metin okunmaz.
Özet (abstract)
The rapid growth of social media has amplified the spread of offensive, violent, and vulgar speech, which poses serious societal and cybersecurity concerns. Detecting such content in Arabic text is particularly complex due to limited labeled data, dialectal variations, and the language’s inherent complexity. This paper proposes a novel framework that integrates multi-task learning (MTL) with active learning to enhance offensive speech detection in Arabic social media text. By jointly training on two auxiliary tasks, violent and vulgar speech, the model leverages shared representations to improve the detection accuracy of offensive speech. Our approach dynamically adjusts task weights during training to balance the contribution of each task and optimize performance. To address the scarcity of labeled data, we employ an active learning strategy through several uncertainty sampling techniques to iteratively select the most informative samples for model training. We also introduce weighted emoji handling to better capture semantic cues. Experimental results using the OSACT2022 dataset show that the proposed framework achieves a state-of-the-art macro F1-score of 85.42%, outperforming existing methods while using significantly fewer fine-tuning samples. The findings of this study highlight the potential of integrating MTL with active learning to efficiently and accurately detect offensive language in resource-constrained settings.
Yazarların özeti; kaynağından alınmıştır. ACM Transactions on Asian and Low-Resource Language Information Processing, 2026 · DOI ↗
Devamı Pofolia uygulamasında
Çıkarımlar ve makaleye soru sorma; ilgi alanına göre her gün yeni özetler. Ücretsiz.
Web'de giriş yaparak açAlan: Yapay Zeka
Artificial IntelligenceComputer Science