The Metanym Game: An LLM Benchmark Without Ground Truth That Rises With the Models It Measures
Researchers have developed a new benchmark for evaluating large language models (LLMs), called the Metanym Game, which measures their ability to gener...
Researchers have developed a new benchmark for evaluating large language models (LLMs), called the Metanym Game, which measures their ability to gener...
A new benchmark called Know2Guess has been developed to evaluate the reliability of large language models. The benchmark separates supported answering...
Researchers have developed an open-source pipeline called MatMMExtract that extracts image-text pairs from materials science literature. The pipeline ...
Researchers have found that Large Language Model (LLM) personalities can be understood in two different ways. Aggregated features, such as personality...
A new paper explores how world models in AI learn task-relevant information through various routes, including observation reconstruction, recurrent st...
Researchers have developed MM-ToolSandBox, a framework for evaluating visual tool-calling agents. This benchmark and evaluation tool includes over 500...
A new software ontology called Skillware has been introduced to define persistent behavioral artifacts in AI agent systems. These artifacts, known as ...
Researchers have created a new benchmark called MobileForge to evaluate the ability of AI models to generate complete mobile apps from visual designs....
Researchers have developed a system to automatically frame surgical videos during laparoscopic surgery. The system uses machine learning to track the ...
Researchers have proposed a new method called SMOPD for improving model performance in multi-reward reinforcement learning. The issue with existing me...
Researchers have developed a new framework called SparkleDock for flexible macromolecular docking on GPU-accelerated supercomputers. The framework ach...
Researchers from the University of Edinburgh have proposed an ontological framework to define decentralization in computer science. They argue that cu...