Mixture-of-Expert Blocks Contain Strong Hallucination Detection Signals
Researchers have developed a new method called InnerExpert that can detect hallucinations in large language models (LLMs) at the level of individual t...
Researchers have developed a new method called InnerExpert that can detect hallucinations in large language models (LLMs) at the level of individual t...
Researchers have studied the performance of a type of AI model called prediction cascades when faced with degraded input data. Prediction cascades rou...
Researchers have proposed a new method for monitoring the behavior of long-horizon agents, which are AI systems that operate across multiple steps and...
Researchers have proposed a new way to evaluate the diversity of AI-generated content called 'diversity profiles.' These profiles are curve-valued sum...
Researchers have developed a new method for combining symbolic and neural network-based learning. They call it Baobab, which translates OWL 2 DL ontol...
Researchers have developed a new approach to tackle the complexity of multi-agent decision making under uncertainty. They propose counting policies in...
Researchers have proposed a new protocol called D$^2$ACCI to improve the diagnosis of failures in large language model (LLM) agents' memory systems. T...
Researchers have introduced StartupBench, a benchmark for general-purpose agents that evaluates their ability to complete real-world tasks. Unlike exi...
Researchers have developed ARASH, a method to improve the efficiency of Tabular Foundation Models (TFMs) in tabular prediction tasks. These models are...
Researchers have developed AutoResearch, a two-stage system that aims to ensure the reliability of research processes. The system combines emerging re...
Researchers have developed a new approach to creating 'adaptive policy portfolios' for Markov decision processes. These portfolios are sets of pre-com...
Researchers have developed EvoTS-Agent, a self-evolving language model agent that can detect changes in financial time series data without human inter...