A
Agentic text-to-3D
Semantically coherent and physically strong, but often requires hours of iterative tool use.
High fidelity · SlowEfficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution
From a single reference image to a physically grounded 3D scene and a diverse family of valid layout variants in minutes.
Base scene↔Layout variant
Base scene↔Layout variant
Each row: one reference, two physically grounded arrangements. All are generated by SceneMosaic.
Overview
The core idea
SceneMosaic combines fast image-based layout priors with local agentic evolution, turning a single reference into a diverse family of simulation-ready layouts.
Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet scalable generation remains challenging. Agentic text-to-3D pipelines can generate high-fidelity scenes but require costly iterative placement and refinement. Parametric image-to-3D models are fast, yet often produce imprecise and physically invalid layouts.
SceneMosaic obtains an initial candidate from a learned image-based prior, then evolves it through VLM agents. By decomposing a scene into independent local units and composing their variants through Cartesian products, the framework delivers efficiency, physical validity, and structured diversity together.
A
Semantically coherent and physically strong, but often requires hours of iterative tool use.
High fidelity · SlowB
Fast feed-forward initialization, but prone to floating, collision, clipping, and boundary errors.
Fast · Physically fragileC
Uses image priors for speed, agents for refinement, and local evolution for combinatorial diversity.
Fast · Valid · DiverseMethod
Hybrid agentic layout evolution
Reconstruct and structure the scene, ground it with physics, evolve local units with Critic–Actor agents, then select representative compositions with a novelty-aware objective.
Register objects, refine instance masks, reconstruct per-object meshes, and organize the scene into anchor-centered local units with explicit support and containment relations.
Correct containment, apply anchor-aware gravity, and ground each local unit before agentic refinement begins.
Agents inspect orthographic layouts, repair semantic and physical errors, and evolve local variants in parallel while memory prevents cyclic edits.
Assemble validated local variants, reject invalid scenes, and select a compact set with novelty-aware dynamic max-min search.
Combinatorial scene diversity
Local units are independently evolved, validated, and recombined. A novelty-aware distance then preserves the most distinct global layouts without sacrificing attach and contain relations.
Results
Fast scene generation
SceneMosaic turns a single reference into a complete, physically grounded 3D scene, then evolves new valid arrangements with only a small additional cost.
Average base scene
~8 min Reconstruction, grounding, and agentic refinementEach additional layout variant
~2 min A new coherent arrangement from evolved local unitsGenerated scenes
Base ↔ Variant
Each paired demo follows the same camera orbit, making the structural changes easy to inspect. Videos load only as they approach the viewport.
Base scene↔Layout variant
Base scene↔Layout variant
Base scene↔Layout variant
Base scene↔Layout variant
Base scene↔Layout variant
Base scene↔Layout variant
Base scene↔Layout variant
Base scene↔Layout variant
Base scene↔Layout variant