AI

Evaluating Large Language Model Performance on International Maritime Dangerous Goods Code Compliance

Researchers have created a benchmark to evaluate the performance of large language models (LLMs) in interpreting the International Maritime Dangerous Goods Code (IMDG). The code is a complex regulatory framework that requires accurate interpretation to prevent safety risks. The benchmark, called DGEval, consists of 1,678 questions across various tasks and was used to test 13 LLMs from six providers. While the best-performing model exceeded human practitioner performance on mu
Researchers have created a benchmark to evaluate the performance of large language models (LLMs) in interpreting the International Maritime Dangerous Goods Code (IMDG). The code is a complex regulatory framework that requires accurate interpretation to prevent safety risks. The benchmark, called DGEval, consists of 1,678 questions across various tasks and was used to test 13 LLMs from six providers. While the best-performing model exceeded human practitioner performance on multiple-choice questions, all models struggled with safety-critical areas such as stowage, segregation, and regulatory recall. This suggests that while LLMs can support compliance tasks, human oversight is still necessary in critical contexts. --- Why it matters: This research matters to engineers working on AI applications in safety-critical domains because it highlights the limitations of current large language models in interpreting complex regulations. Understanding these limitations is crucial for developing reliable and safe AI systems that can augment human decision-making in high-stakes environments. Source: https://arxiv.org/abs/2608.21036

This article was originally published at: https://arxiv.org/abs/2608.21036