ComponentBench: Diagnosing Component-Level Failures in Computer-Use Agents
Researchers have developed a new benchmark called ComponentBench to evaluate the performance of comp...
Researchers have developed a new benchmark called ComponentBench to evaluate the performance of comp...
Researchers have proposed a method to supervise AI models using governance records generated by mach...
A new benchmark called THPT-Ladder has been introduced to evaluate language models on Vietnam's Nati...
Researchers have evaluated the robustness of AI code agents to superficial changes in code. They fou...
Researchers have developed a method to detect when wearable devices for stress c...
Researchers have developed a new framework called SDDL that improves the accurac...
Researchers have created a benchmark called FM-Bench to evaluate the performance...
Researchers have proposed a new framework called UMER for universal multimodal r...