PLoS ONE· 2026Q1
English-Persian cross-language plagiarism detection: Multilingual deep learning approach
- 0citations
- Q1SCImago
- 2026year
Short summary
A fine-tuned Logistic Regression model achieved 96% accuracy in detecting English-Persian cross-lingual plagiarism, outperforming complex deep learning models.
AI-generated from the title and abstract; the full text is not read.
Key points
- A fine-tuned Logistic Regression model achieved 96% accuracy in English-Persian cross-lingual plagiarism detection.
- The approach uses BAAI/bge-m3 embeddings and cosine similarity to measure linguistic distances.
- The optimized LR model improved accuracy from 89% to 96% (p < 0.05) compared to other classifiers.
- High-quality embeddings enhance detection in low-resource contexts without complex deep learning models.
AI-generated from the title and abstract; the full text is not read.
Abstract
Cross-lingual plagiarism detection (CLPD) remains a critical challenge in Natural Language Processing, particularly for low-resource language pairs where translation and paraphrasing obscure source materials. This study introduces a robust framework for Persian-English CLPD by integrating state-of-the-art sentence representations with optimized machine learning classifiers. We utilized the BAAI/bge-m3 model to generate high-dimensional embeddings, employing cosine similarity to measure contrastive linguistic distances. To evaluate classification performance, we compared various architectures, including XGBoost, LSTM, and Logistic Regression (LR), across three categories: Exact, Semi-Exact, and Different. While deep learning and ensemble methods were initially considered, experimental results on an expanded dataset demonstrated that a fine-tuned LR model provided superior stability and generalization. The optimized LR approach significantly increased classification accuracy from 89% to 96% (p < 0.05). These findings suggest that high-quality embeddings can enhance detection performance in low-resource contexts without the computational overhead of complex architectures. To ensure transparency and reproducibility, our dataset and source code are available on GitHub.
The authors' abstract, as published at the source. PLoS ONE, 2026 · DOI ↗
Continue with a free account
Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.
Continue free on the webSign in with Google or Apple; no card needed. You come back to this paper.
On your phone:
Field: Safety Research
Safety ResearchSocial Sciences