OenoBench: A Wine-Domain Benchmark for Knowledge-Grounded Evaluation of Large Language Models
Researchers have developed OenoBench, a benchmark for evaluating large language models' knowledge in the wine domain. The benchmark consists of 3,266 ...
Researchers have developed OenoBench, a benchmark for evaluating large language models' knowledge in the wine domain. The benchmark consists of 3,266 ...
Large language models (LLMs) often encounter conflicting evidence from text and numbers when making decisions. Researchers have created a synthetic be...
Researchers have created a benchmark called FormalTCS to evaluate the performance of large language models (LLMs) in conducting end-to-end theoretical...
Researchers have developed a benchmark called ConceptGuard to evaluate the ability of large language models to remove harmful or sensitive knowledge w...
Researchers propose a new framework called HARP for prioritizing vulnerabilities based on specific operational preferences. Unlike existing methods th...
Researchers propose DeltaMomentum, a new method for updating momentum in deep learning optimizers. Unlike traditional exponential moving average (EMA)...
Researchers from Japan investigated whether adding listening behaviors to an AI clone can improve its perceived authenticity. They integrated verbal b...
Researchers have developed StreamSoccer, a system for live soccer commentary that uses event-driven memory to process and generate commentary in real-...
Researchers have proposed a framework for improving the accuracy of medical image captioning. Medical image captioning is a technique that helps docto...
Researchers have developed a method to improve the efficiency of communication topologies in multi-agent systems. The approach, called Reward-Guided A...
Researchers have proposed a hybrid framework for autonomous driving that combines reinforcement learning and PID control with the common-sense reasoni...
Researchers have developed a new framework called Explain-MDRC for recognizing depression in clinical interviews. It combines text, audio, and facial ...