SceneMosaic

Efficient and Diverse Simulation-Ready Scene Generation via Hybrid Agentic Layout Evolution

Xingjian Ran1,* Xiaoye Mo2,* Sihao Liu1 Jianyu Zhang1 Li Luo1 Bo Dai1,†

1The University of Hong Kong 2University of Electronic Science and Technology of China

* Equal contribution. † Corresponding author.

From a single reference image to a physically grounded 3D scene and a diverse family of valid layout variants in minutes.

Base sceneLayout variant

Base scene
Layout variant

Base sceneLayout variant

Base scene
Layout variant

Each row: one reference, two physically grounded arrangements. All are generated by SceneMosaic.

Explore the research

The core idea

Fast priors. Deliberate agents.
Many valid worlds.

SceneMosaic combines fast image-based layout priors with local agentic evolution, turning a single reference into a diverse family of simulation-ready layouts.

Diverse and simulation-ready indoor scenes are essential for interactive entertainment and embodied AI, yet scalable generation remains challenging. Agentic text-to-3D pipelines can generate high-fidelity scenes but require costly iterative placement and refinement. Parametric image-to-3D models are fast, yet often produce imprecise and physically invalid layouts.

SceneMosaic obtains an initial candidate from a learned image-based prior, then evolves it through VLM agents. By decomposing a scene into independent local units and composing their variants through Cartesian products, the framework delivers efficiency, physical validity, and structured diversity together.

A

Agentic text-to-3D

Semantically coherent and physically strong, but often requires hours of iterative tool use.

High fidelity · Slow

B

Image-to-3D priors

Fast feed-forward initialization, but prone to floating, collision, clipping, and boundary errors.

Fast · Physically fragile

Hybrid agentic layout evolution

A scene pipeline built around locality.

Reconstruct and structure the scene, ground it with physics, evolve local units with Critic–Actor agents, then select representative compositions with a novelty-aware objective.

The full SceneMosaic pipeline. Open the figure to inspect object relations, local repair, and novelty-aware selection.
  1. 01

    Reconstruct & structure

    Register objects, refine instance masks, reconstruct per-object meshes, and organize the scene into anchor-centered local units with explicit support and containment relations.

    • Perception agent + SAM3
    • Per-object SAM3D reconstruction
    • Hierarchical scene tree
  2. 02

    Physically stabilize

    Correct containment, apply anchor-aware gravity, and ground each local unit before agentic refinement begins.

    • Containment correction
    • Anchor-dependent gravity
    • Collision-hull validation
  3. 03

    Critic–Actor evolution

    Agents inspect orthographic layouts, repair semantic and physical errors, and evolve local variants in parallel while memory prevents cyclic edits.

    • Orthographic 2D abstraction
    • Symbolic pose updates
    • Cross-unit constraints
  4. 04

    Compose for diversity

    Assemble validated local variants, reject invalid scenes, and select a compact set with novelty-aware dynamic max-min search.

    • Cartesian variant assembly
    • Position + distance + rotation
    • Dynamic representative selection

Combinatorial scene diversity

Structured alternatives.

Local units are independently evolved, validated, and recombined. A novelty-aware distance then preserves the most distinct global layouts without sacrificing attach and contain relations.

Local evolutionIndependent anchor-centered edits Physical validationGrounded layouts before composition

Fast scene generation

Realistic scenes, generated in minutes.

SceneMosaic turns a single reference into a complete, physically grounded 3D scene, then evolves new valid arrangements with only a small additional cost.

Average base scene

~8 min Reconstruction, grounding, and agentic refinement

Each additional layout variant

~2 min A new coherent arrangement from evolved local units

Drag or scroll to inspect the full-resolution figure.