When Contextual Inference Fails: Cancelability in Interactive Instruction Following
Researchers have created an interactive benchmark called Build What I Mean (BWIM) to test how well language models can follow instructions in a collab...
Researchers have created an interactive benchmark called Build What I Mean (BWIM) to test how well language models can follow instructions in a collab...
Researchers have developed a new AI framework called Doc-V* that can answer questions about multi-page documents without needing to recognize text fir...
A new method called Token-to-Mask (T2M) has been proposed to improve the performance of diffusion language models. T2M identifies low-confidence posit...
Researchers have developed a new framework called EvoSelect to improve the efficiency of adapting large language models (LLMs) to specific tasks. The ...
Researchers have proposed a new approach to deploying large language model (LLM) judge panels, which are used to evaluate and improve LLMs. The approa...
Researchers have developed a new approach called Self-Harness that enables Large Language Model (LLM)-based agents to improve their own operating harn...
Researchers propose a new approach to medical vision-language models that focuses on calibrated triage rather than autonomy. They evaluate nine confid...
Researchers have proposed a new method called RepSelect for robustly removing unwanted knowledge and tendencies from large language models (LLMs). Thi...
Researchers propose SPyCE (Skill-Policy Co-evolution), a framework that helps multimodal agents learn skills and policies simultaneously. This approac...
Researchers have developed a machine learning model that can interpret lunar geology by analyzing topographic, spectral, and geological maps. The mode...
Researchers have developed LODESTAR, a method to improve the performance of question-answering systems that use retrieval-augmented generation. The ap...
Researchers have found that generative AI models can be used to create fake evidence that degrades the performance of systems designed to detect out-o...