AI

presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

Researchers have developed a new framework called presto for placing objects in images. Unlike previous methods that require training or hand-crafted rules, presto uses a multimodal language model to reason about object placement. This allows it to work in open-world scenarios where the objects and scenes are novel. The framework uses an imaginary action space to refine object position and scale, resulting in state-of-the-art performance on multiple benchmarks. Human studies
Researchers have developed a new framework called presto for placing objects in images. Unlike previous methods that require training or hand-crafted rules, presto uses a multimodal language model to reason about object placement. This allows it to work in open-world scenarios where the objects and scenes are novel. The framework uses an imaginary action space to refine object position and scale, resulting in state-of-the-art performance on multiple benchmarks. Human studies suggest that presto produces more perceptually coherent placements than other methods. --- Why it matters: This matters because it enables AI systems to place objects in images without needing extensive training data or hand-crafted rules, making them more flexible and adaptable to new scenarios. Source: https://arxiv.org/abs/2608.21543

This article was originally published at: https://arxiv.org/abs/2608.21543