Insider Brief
- World Labs introduced Atlas, a new world model designed to generate, reconstruct and simulate 3D environments, including Real-to-Sim workflows for robot navigation and manipulation.
- The company said Atlas can reconstruct physical spaces from images or video and generate the RGB and depth data a robot would encounter while moving through those environments.
- For manipulation, World Labs said Atlas can recreate interactions with rigid, articulated and deformable objects and vary objects, positions, robot motions, lighting and backgrounds to create training and testing scenarios.
World Labs has introduced Atlas, a new “spatial intelligence” world model designed to generate, reconstruct and simulate 3D environments, with robotics among the applications targeted by its ability to turn real-world recordings into training and testing environments.
According to World Labs, Atlas supports so-called Real-to-Sim workflows, which recreate real spaces in simulation for robot navigation and manipulation. Atlas can reconstruct physical spaces from a small number of images or video frames and then simulate what a robot would see as it moves through those environments.
For navigation, World Labs said Atlas can build a 3D representation of an environment from cellphone video and generate the RGB images and depth information that a robot’s cameras would encounter along different paths. In demonstrations, the company reconstructed large environments using 24 video frames and simulated different robots navigating through them.
Atlas also supports robotic manipulation and World Labs noted developers can use a small number of real-world recordings to recreate interactions involving rigid, articulated and deformable objects. Once a task is represented in simulation, developers can vary the objects, their positions, robot movements, lighting and backgrounds to generate different training and testing scenarios.
The model combines text, images, video and 3D information within a shared spatial context. World Labs describes Atlas as a multimodal autoregressive diffusion transformer that grounds images and depth maps to specific camera positions, allowing it to reason about the spatial relationships within a scene.
The company also said Atlas can produce explicit 3D outputs, including point clouds and 3D Gaussian splats, from images or video. Those representations can be used in workflows that require a 3D environment rather than only generated images or video.
Beyond robotics, Atlas supports camera-controlled video generation, 3D reconstruction and image generation. World Labs said the model is entering early access with select partners and will also power future versions of its Marble world-generation product.