When to Plan, When to Polish: Noise Level as a Granularity Axis for Diffusion Language Models
Researchers propose a new approach to training language models called Noise Dependent Granularity Control (NDGC). This method uses the level of noise ...
Researchers propose a new approach to training language models called Noise Dependent Granularity Control (NDGC). This method uses the level of noise ...
Researchers investigated whether large language models (LLMs) can accurately identify when their previous responses were manipulated by an 'adversaria...
Researchers have introduced DominoTree, a new method for speculative decoding in large language models. It builds on the Domino drafter by adding a GR...
Researchers have developed a new language model called VEXMLM to improve the performance of multilingual AI systems on low-resource Ge'ez-script langu...
Researchers have developed a benchmark called LexKairos to evaluate the ability of large language models (LLMs) to understand and work with time-relat...
Researchers have developed SAEVerbalizer, a framework that generates explanations for features extracted by sparse autoencoders (SAEs) from large lang...
Researchers have introduced RecurrentGPT, a new type of transformer model that balances expressivity and memory efficiency. Unlike traditional transfo...
Researchers have found that some AI models can produce coherent text even when given no useful input. They studied this phenomenon in speech recogniti...
Researchers have developed a way to measure how polarized online narratives are in relation to real-world conflicts. They analyzed 212 YouTube videos ...
Researchers have developed a new method for modifying language model behavior without updating parameters. The method, called FishBack, corrects a fla...
Douyin Multimodal Embedding (DME) is a new AI model that combines the strengths of contrastive models and CoT-based models for multimodal representati...
Researchers have proposed a new framework called AVA-Encoder for learning video representations that are directly usable by artificial agents. The fra...