Fragility of Value under Imperfect Alignment
Researchers have proposed a model to address concerns about AI systems being aligned with human values. They argue that optimizing too heavily for imp...
Researchers have proposed a model to address concerns about AI systems being aligned with human values. They argue that optimizing too heavily for imp...
Researchers have proposed a new approach to creating autonomous agents that can tackle complex decision-making tasks. The method combines the strength...
Researchers have developed a new benchmark called BrainBench to evaluate the ability of large language models (LLMs) to understand electroencephalogra...
Researchers have developed Mechanist, an AI system designed to use machine learning models as scientific instruments for discovering the underlying me...
Researchers have created a benchmark to test whether language models can recover the original research idea behind a published paper based on its pre-...
Researchers have developed a new approach called GRIP, which aims to improve the performance of retrieval-augmented generation (RAG) models. These mod...
Researchers have developed a method to optimize machine learning algorithms for low energy consumption while maintaining their performance. The approa...
Researchers have developed a new framework for testing the vulnerability of large language models (LLMs) to audio-based attacks. The framework, called...
Researchers propose an iterative process to improve generative modeling for image generation. They use flow matching, a technique that can lead to hal...
Researchers have revisited the 'Sleeping Beauty problem', a thought experiment in decision-making under uncertainty. The problem involves Sleeping Bea...
Researchers have developed a method called NINJA to 'jailbreak' large language models by appending harmless content to user goals. This allows attacke...
Researchers have created CausalProfiler, a tool that generates synthetic benchmarks for evaluating causal machine learning methods. These benchmarks a...