AI

Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL

A new approach to evaluating Large Language Model (LLM) performance in text-to-SQL tasks is proposed. The authors argue that current evaluation methods are flawed due to heterogeneity across benchmarks, backbones, and inference protocols. They introduce a taxonomy based on the level of autonomy, ranging from constrained to agentic generation. A case study on the Spider benchmark shows that autonomy can improve robustness but at a cost. The study also highlights the effectiven
A new approach to evaluating Large Language Model (LLM) performance in text-to-SQL tasks is proposed. The authors argue that current evaluation methods are flawed due to heterogeneity across benchmarks, backbones, and inference protocols. They introduce a taxonomy based on the level of autonomy, ranging from constrained to agentic generation. A case study on the Spider benchmark shows that autonomy can improve robustness but at a cost. The study also highlights the effectiveness of chain-of-thought supervision in certain scenarios. --- Why it matters: This research matters because it provides a framework for comparing LLMs across different tasks and protocols, which is essential for advancing the field. Understanding how autonomy affects performance will help researchers design more effective models. Source: https://arxiv.org/abs/2608.15389

This article was originally published at: https://arxiv.org/abs/2608.15389