OmegaUse-OfficeVal: Benchmarking LLM Agents on Long-Horizon Office-Suite Tasks with Economic Grounding
Researchers have created a benchmark called OmegaUse-OfficeVal to evaluate large language model (LLM) agents' ability to complete office-suite tasks e...