Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge
Researchers have evaluated three lightweight language models (LLMs) - Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite - for their ability to perform in-depth free-text diagnostics in the context of 5G domain knowledge and fault analysis. The LLMs were tested on three benchmarks using a 'LLM-as-Judge' methodology, where multiple judges scored the models' outputs. While all models achieved high accuracy (at least 90%) in fault diagnosis, they struggled with recalling
Researchers have evaluated three lightweight language models (LLMs) - Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite - for their ability to perform in-depth free-text diagnostics in the context of 5G domain knowledge and fault analysis. The LLMs were tested on three benchmarks using a 'LLM-as-Judge' methodology, where multiple judges scored the models' outputs. While all models achieved high accuracy (at least 90%) in fault diagnosis, they struggled with recalling specific technical standards (3GPP and O-RAN specifications), scoring below 60% in this task. The study highlights the potential of lightweight LLMs for automating diagnostic reasoning in telecom networks, but also emphasizes the need for further improvement in their ability to recall detailed technical information.
---
Why it matters: This research matters because it investigates the feasibility of using lightweight language models for automating fault analysis and diagnostics in 5G and emerging 6G networks. The results have implications for the development of efficient and accurate diagnostic tools that can be deployed at the edge of telecom networks.
Source: https://arxiv.org/abs/2608.21021
This article was originally published at: https://arxiv.org/abs/2608.21021