SkillEffect: Checked Lowering for Memory-Bounded Agent Tools
Researchers have developed a system called SkillEffect to help language models manage memory usage when interacting with tools. The system uses a chec...
Researchers have developed a system called SkillEffect to help language models manage memory usage when interacting with tools. The system uses a chec...
Researchers propose a new framework to understand how agents allocate information between their own memory and communication with peers. They introduc...
Researchers have proposed a new defense mechanism called DiSCO to prevent text-to-image generative models from producing harmful content. The system o...
Researchers have developed KernelArc, a framework for optimizing GPU kernels across different workloads. The system uses multiple agents that run in p...
Researchers have proposed a method called CASE to improve the accuracy of large language models by selecting the best answer based on hidden-state sig...
Researchers propose a framework called 'cooperative observation' to improve the performance of personal AI systems. They argue that these systems need...
Researchers have developed a new framework called KNOWSIM for evaluating the performance of Large Language Models (LLMs) in collaborating with users o...
Researchers have developed an automated approach to designing feature extractors for constraint satisfaction problems. This method uses large language...
Researchers evaluated five widely used benchmark suites on 26 open-source small language models to determine their effectiveness in assessing safety. ...
Researchers have developed a new defense mechanism called 'Fool's Gold' to protect open-weight language models from safety-removal attacks. These atta...
Researchers have developed an audit protocol to test the performance of personalized agents in decision-making tasks. They created a synthetic develop...
Researchers proposed a new method to evaluate scientific hypotheses generated by large language models (LLMs). Instead of relying on LLMs as judges or...