Synthesizing Feature Extractors: An Agentic Approach for Algorithm Selection
Researchers have developed an automated approach to designing feature extractors for constraint satisfaction problems. This method uses large language...
Researchers have developed an automated approach to designing feature extractors for constraint satisfaction problems. This method uses large language...
Researchers evaluated five widely used benchmark suites on 26 open-source small language models to determine their effectiveness in assessing safety. ...
Researchers have developed a new defense mechanism called 'Fool's Gold' to protect open-weight language models from safety-removal attacks. These atta...
Researchers have developed an audit protocol to test the performance of personalized agents in decision-making tasks. They created a synthetic develop...
Researchers proposed a new method to evaluate scientific hypotheses generated by large language models (LLMs). Instead of relying on LLMs as judges or...
Researchers have introduced ASI-Bench, a new benchmark designed to evaluate the capabilities of artificial intelligence systems in exploring the unkno...
A new AI framework called DeAR (Decentralized Agentic Reasoning) has been proposed to improve the accuracy of complex reasoning tasks. Unlike traditio...
Researchers have proposed a new method for training large language models (LLMs) called PlanPO. This approach aims to improve the performance of LLMs ...
Researchers have introduced LiveHouse-TS, an open-world living benchmark for Time Series Foundation Models (TSFMs). Unlike traditional benchmarks that...
Researchers have fine-tuned a large language model, called Qwen2.5-3B-Base, to improve its mathematical reasoning capabilities in signal processing pr...
Researchers from the AIMAE Team have introduced Wuying-Browser-Agent, a unified framework for building browser agents that can perform well in real-wo...
Researchers evaluated three large language models for medical consultation and found that they often provide self-care advice before the patient's con...