Skip to content

CAST-Eval: A Domain-Specific Benchmark for Large Language Models in Civil Aviation Safety

Semantic Scholar · Article (2024 IEEE 2nd International Conference on Electrical, Automation and Computer Engineering (ICEACE)) · 1757cd9e7ab0786ac0bf0a53d39c44488bb716ac · Published 2024-12-29 · 2024 IEEE 2nd International Conference on Electrical, Automation and Computer Engineering (ICEACE) · 6 authors

Abstract and citation only, verbatim from Semantic Scholar; full text lives there. All credit to the authors and 2024 IEEE 2nd International Conference on Electrical, Automation and Computer Engineering (ICEACE).

Abstract

In this paper, we present CAST-Eval, a novel, comprehensive and domain-specific benchmark designed to assess the knowledge and reasoning capabilities of large language models (LLMs) in the civil aviation domain. CAST-Eval covers critical professional knowledge areas, including airlines, airports, air traffic control, and other key sectors, with tasks ranging from basic university-level content to advanced qualification exams issued by the CAAC. To validate its effectiveness, we evaluated major Chinese LLMs, both closed-source and open-source, and publicly released the results. Our experiments show that in the civil aviation domain, closed-source commercial models generally outperformed smaller-scale open-source models, with most achieving scores above 60. We also found that while general-purpose models performed well in areas such as meteorology and navigation, specialized fields like flight operations and dispatch remain particularly challenging for general AI models. Furthermore, newly released 70B-level open-source models exhibited performance comparable to or even surpassing some closed-source models. CAST-Eval can serves as a versatile evaluation tool for fine-tuned language models across other industry-specific applications, supporting improvements in AI model training and optimization.

Authors

  • Dongxiao Qiao
  • Xianhui Tian
  • Yubin Xu
  • Jie Yang
  • Le Yang
  • Xu-Hui Wang

Keywords

  • Engineering
  • Computer Science

Citation

Dongxiao Qiao, Xianhui Tian, Yubin Xu , et al. (2024). CAST-Eval: A Domain-Specific Benchmark for Large Language Models in Civil Aviation Safety. 2024 IEEE 2nd International Conference on Electrical, Automation and Computer Engineering (ICEACE). Semantic Scholar ID 1757cd9e7ab0786ac0bf0a53d39c44488bb716ac. https://doi.org/10.1109/ICEACE63551.2024.10898372 ↗