FormalTCS: Benchmarking End-to-End Frontier Formal Theoretical Computer Science Research of Large Language Models
Researchers have created a benchmark called FormalTCS to evaluate the performance of large language models (LLMs) in conducting end-to-end theoretical computer science research. The benchmark consists of 175 instances drawn from recent papers in top conferences and includes expert-verified Lean formalizations and proofs. Evaluations show that current LLMs struggle with tasks such as autoformalization, where they achieve only a 11.5% success rate in translating natural-languag
Researchers have created a benchmark called FormalTCS to evaluate the performance of large language models (LLMs) in conducting end-to-end theoretical computer science research. The benchmark consists of 175 instances drawn from recent papers in top conferences and includes expert-verified Lean formalizations and proofs. Evaluations show that current LLMs struggle with tasks such as autoformalization, where they achieve only a 11.5% success rate in translating natural-language claims into formal theorem statements.
---
Why it matters: This matters to AI researchers because it highlights the limitations of current large language models in conducting theoretical computer science research and identifies areas for improvement, such as autoformalization and research taste.
Source: https://arxiv.org/abs/2608.20153
This article was originally published at: https://arxiv.org/abs/2608.20153