VibeWorlding: Can Multimodal Agents Construct 3D Open Worlds End-to-End?
Researchers propose VibeWorlding, a framework for training and benchmarking multimodal agents that can create 3D open worlds from user queries. They introduce a large-scale dataset of 2,616 high-quality 3D assets and 6,828 user queries, as well as a joint reinforcement learning post-training framework called VibeWorlding-Gym. Experiments show that current language models struggle to solve the task, but RL training can improve performance and even surpass closed-source frontie
Researchers propose VibeWorlding, a framework for training and benchmarking multimodal agents that can create 3D open worlds from user queries. They introduce a large-scale dataset of 2,616 high-quality 3D assets and 6,828 user queries, as well as a joint reinforcement learning post-training framework called VibeWorlding-Gym. Experiments show that current language models struggle to solve the task, but RL training can improve performance and even surpass closed-source frontiers.
---
Why it matters: This matters because it addresses the challenge of creating interactive 3D worlds from user queries, which is essential for applications like virtual reality and game development. The proposed framework and dataset can help advance multimodal AI research and its practical applications.
Source: https://arxiv.org/abs/2608.15265
This article was originally published at: https://arxiv.org/abs/2608.15265