Agentic-SQL Revisited: Autonomy-Based Taxonomy and Empirical Benchmark Analysis for LLM Text-to-SQL
A new approach to evaluating Large Language Model (LLM) performance in text-to-SQL tasks is proposed. The authors argue that current evaluation methods are flawed due to heterogeneity across benchmarks, backbones, and inference protocols. They introduce a taxonomy based on the level of autonomy, ranging from constrained to agentic generation. A case study on the Spider benchmark shows that autonomy can improve robustness but at a cost. The study also highlights the effectiven
A new approach to evaluating Large Language Model (LLM) performance in text-to-SQL tasks is proposed. The authors argue that current evaluation methods are flawed due to heterogeneity across benchmarks, backbones, and inference protocols. They introduce a taxonomy based on the level of autonomy, ranging from constrained to agentic generation. A case study on the Spider benchmark shows that autonomy can improve robustness but at a cost. The study also highlights the effectiveness of chain-of-thought supervision in certain scenarios.
---
Why it matters: This research matters because it provides a framework for comparing LLMs across different tasks and protocols, which is essential for advancing the field. Understanding how autonomy affects performance will help researchers design more effective models.
Source: https://arxiv.org/abs/2608.15389
This article was originally published at: https://arxiv.org/abs/2608.15389