PofoliaShared via Pofolia

Journal of Air Transportation· 2026Q2

Enhancing Aviation Safety: Natural Language Processing for Automated Classification of Aviation Occurrences

Charilaos C. Kakoulidis, Michail K. Psaropoulos, Anestis I. Kalfas

Short summary

A new NLP framework using a domain-specific BERT model automates aviation safety occurrence classification, achieving 94% F1 score and reducing analysis time from months to seconds.

AI-generated from the title and abstract; the full text is not read.

Key points

  • A domain-specific BERT model was developed for automated classification of aviation safety occurrences.
  • A sentence-based splitting and aggregation method overcomes transformer token limitations, improving F1 score by 6 percentage points.
  • The framework achieves a 94% F1 score, surpassing human analyst consistency.
  • Training on a larger dataset (41,000 reports) yields a 10 percentage point higher F1 score than a smaller balanced subset (6400 reports).

AI-generated from the title and abstract; the full text is not read.

Abstract

This study presents a natural language processing framework for automated classification of aviation safety occurrences. The proposed approach develops a domain-specific Bidirectional Encoder Representations from Transformers model trained on a large corpus of aviation safety reports to classify occurrences according to the European Union Aviation Safety Agency Key Risk Areas. The work addresses limitations of current manual safety assessment processes, which require weeks or months to analyze a single occurrence, delaying the identification of emerging risks. To overcome the 512-token limitation of transformer architectures, a sentence-based splitting and aggregation methodology is introduced. Long reports are divided into sentences, processed independently, and subsequently aggregated, preserving narrative context while avoiding information loss due to truncation. This approach improves classification performance by 6 percentage points in F1 score. Experimental results demonstrate that dataset size has a greater impact on performance than dataset balance. Training on the full dataset of 41,000 reports outperforms a balanced subset of 6400 reports by 10 percentage points in F1 score. The developed framework achieves an F1 score of 94%, exceeding human analyst consistency levels while reducing assessment time from months to seconds. The framework enables standardized, near-real-time aviation safety analysis, supporting proactive safety risk management.

The authors' abstract, as published at the source. Journal of Air Transportation, 2026 · DOI ↗

TakeawaysPremium
Ask the paperFree account

Continue with a free account

Ask the paper: 3 free questions a day about this paper; save it, get its citation, new summaries every day for your field. Takeaways are Premium.

Continue free on the web

Sign in with Google or Apple; no card needed. You come back to this paper.

On your phone:

Field: Radiological and Ultrasound Technology

Radiological and Ultrasound TechnologyHealth Professions