Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking Scheme
A new benchmark called THPT-Ladder has been introduced to evaluate language models on Vietnam's National High School Graduation Examination. The bench...
A new benchmark called THPT-Ladder has been introduced to evaluate language models on Vietnam's National High School Graduation Examination. The bench...
Researchers have evaluated the robustness of AI code agents to superficial changes in code. They found that even top-performing models can be affected...
Researchers have developed a method to detect when wearable devices for stress classification are not working correctly. They found that even if the d...
Researchers have developed a new framework called SDDL that improves the accuracy of language models in solving combinatorial optimization problems. T...
Researchers have created a benchmark called FM-Bench to evaluate the performance of language model agents in long-term decision-making scenarios. The ...
Researchers have proposed a new framework called UMER for universal multimodal retrieval. This framework uses a technique called Pair-Aware Discrimina...
Researchers have developed a new method called HN-CLIP to improve dense-caption retrieval in AI systems. The current methods for this task often suffe...
Researchers have found that using pairwise ranking outperforms single-action reinforcement learning for selecting explanations in offline recommendati...
Researchers have developed a benchmark called FinRCA-Bench to evaluate the performance of artificial intelligence systems in financial evidence retrie...
Researchers have developed a system that combines search and CRM (customer relationship management) functions using AI-powered agents. The goal is to ...
Researchers have developed a framework called FACET to improve the creation of complex tasks for training AI agents. The framework addresses two main ...
Researchers have proposed a new model called BudgetDoc that can estimate the performance of large language models (LLMs) in document-related tasks. Th...