GENA3D
  • ECCV 2026
  • arXiv 2511.21945

Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence

Dartmouth College

Texture Geometry
drag the line
Input · 4 views
Bottom left: the occluded input views. The rest is generated by GENA3D. Left of the line is the textured render, right of it the surface normals of the same frame.
67.0 Perceptual coherence (PCS), 4 views VLM-judged, on GSO. Amodal3R: 46.2
30.73 FID, 4 views Lowest of all methods. Amodal3R: 35.15
78.2% Human preference Over Amodal3R at 4 views, 30 objects × 10 raters
1→20 Input views FID keeps improving, 33.91 → 29.43

OverviewFigure 1

Imagine in 2D, settle it in 3D

Objects in real scenes are partly hidden. Recovering the full object means inventing what was never seen, while keeping everything that was seen exactly right.

(a) Reconstruct in 3D

Faithful, but full of holes

Multi-view stereo keeps the visible geometry consistent, yet has nothing to say about the parts no camera saw.

3D coherentnot generative
(b) Complete in 2D, then lift

Complete, but contradictory

2D amodal completion imagines the hidden parts convincingly, but each view imagines something different.

generativeinconsistent across views
(c) GENA3D

2D priors propose, 3D geometry decides

Completed views drive a conditional 3D generator, while partial stereo geometry anchors it to what was actually observed.

generative3D coherent
Fig. 1From sparse, partially occluded inputs: (a) MVS reconstruction (e.g., VGGT), (b) 2D amodal completion lifted to 3D (TRELLIS), (c) GENA3D.

Generating complete 3D objects under partial occlusions (i.e., amodal scenarios) is a practically important yet challenging problem, as large portions of object geometry are unobserved in real-world scenarios. Existing approaches either operate directly in 3D, which ensures geometric consistency but often lacks generative expressiveness, or rely on 2D amodal completion, which provides strong appearance priors but does not guarantee reliable 3D structure. This raises a key question: how can we achieve both generative plausibility and geometric coherence in amodal 3D modeling?

To answer this question, we introduce GENA3D (GENerative Amodal 3D), a framework that integrates learned 2D generative priors with explicit 3D geometric reasoning within a conditional 3D generation paradigm. The 2D priors enable the model to plausibly infer diverse occluded content, while the 3D representation enforces multi-view consistency and spatial validity. Our design incorporates a novel View-Wise Cross-Attention for multi-view alignment and a Stereo-Conditioned Cross-Attention to anchor generative predictions in 3D relationships.

By combining generative imagination with structural constraints, GENA3D generates complete and coherent 3D objects from limited observations without sacrificing geometric fidelity. Experiments demonstrate that our method outperforms existing approaches in both synthetic and real-world amodal scenarios, highlighting the effectiveness of bridging 2D priors and 3D coherence in generating plausible and geometrically consistent 3D structures in complex environments.

§3 MethodFigures 2–4

From 1–4 unposed views to one complete object

GENA3D builds on a two-stage 3D generator. Only the first stage, which decides the object's sparse structure, is retrained, and it gains two conditioning modules.

Fig. 2Given sparse images S, visibility masks Mvis and occlusion masks Mocc of a target object, GENA3D generates a sparse structure by aggregating multi-view information and inferring the complete structure from the partial stereo point cloud. A pretrained amodal SLAT transformer then generates the structured latent, which is decoded into an occlusion-free 3D object.
  1. Observe

    K sparse, unposed views (typically 1–4). Visibility masks and a text prompt come from foundation models such as SAM and Florence-2.

    input · Mvis
  2. Complete in 2D

    Each view is amodally completed on its own. The results look plausible but disagree with each other.

    OAAC · I, Mocc
  3. Fuse the views

    View-Wise Cross Attention reads every completed view in parallel and fuses them by visibility.

    new · §3.1
  4. Anchor in 3D

    Stereo-Conditioned Cross Attention gates attention with the partial point cloud from an MVS model.

    new · §3.2
  5. Generate detail

    A pretrained amodal SLAT transformer adds appearance and decodes the complete object.

    stage 2 · frozen
§3.1 · multi-view alignment

View-Wise Cross Attention

Conditioning on one view per sampling step, or concatenating all views, lets later views overwrite earlier structure and lets inconsistent completions pile up as geometric drift. GENA3D instead attends to every view in parallel at every step, and weights each view by how much of the object it actually saw.

z'ₙ = CA(z, DINO(iⁿ)) for every view n z' = (1/K) Σₙ wₙ z'ₙ , wₙ = τₙ / Σⱼ τⱼ τₙ = mvⁿ / (moⁿ + mvⁿ) visibility ratio

Why it works: the base generator always produces objects in a canonical pose, whatever the camera. Views can therefore be fused without knowing their poses. In 100 GSO objects × 10 viewpoints, 97.6% of generations land in the canonical axis-aligned frame.

Without it (2 views): FID 32.12 → 38.64
§3.2 · geometric anchoring

Stereo-Conditioned Cross Attention

Geometry added as extra tokens is easy to ignore, especially when it is partial and noisy. Here the partial point cloud of the target, lifted from an MVS model (π³) and cut out by the visibility masks, is voxelized at 64³, encoded, and turned into a gate inside the attention logits.

P_O = ∪ₙ { p ∈ MVS(sⁿ) | mvⁿ(p) = 1 } c_geo = Linear(Patchify(Enc(Voxelize(P_O, 64³)))) g = σ(MLP(c_geo)) α = softmax( QKᵀ/√D + log(g + ε) )

Robust to bad stereo: when the reconstructed points are misaligned or overlap across views, the generative prior takes over and corrects them (Fig. 7).

Without it (2 views): FID 32.12 → 42.92, MMD 5.50 → 5.91
Fig. 4View-Wise Cross Attention and Stereo-Conditioned Cross Attention in detail.
Fig. 3The original model, conditioned on multiple views, is biased and accumulates artifacts such as a doubled headboard.
13,957 objects3D-FUTURE (9,472) and ABO (4,485), made watertight.
20–60% hidden3D-consistent occlusion masks grown face by face over each mesh.
1–4 viewsSampled per batch; stereo point clouds are computed on the fly.
~16 hours12 epochs on 8 RTX 6000 GPUs with a flow-matching objective.

§4 ExperimentsTables 1–2 · Figure 5

The most coherent objects at every view count

All 1,030 objects of Google Scanned Objects, with 3D-consistent occlusions and views whose visible part covers 30–70% of the render. Baselines get the same 2D amodal completion (OAAC) where they need it.

Table 1Amodal 3D object generation on GSO with 1, 2 and 4 views. Bold: best, underlined: second best within each setting. OAAC: per-view 2D amodal completion. ‡ OAAC on one view, then multi-view diffusion (OAAC + MV). SAM3D is evaluated single-view only, following its official setup.
ViewsMethod2D amodalFID ↓KID (%) ↓CLIP (%) ↑MMD (‰) ↓COV (%) ↑PCS (%) ↑
1 viewTRELLISOAAC49.683.2176.796.0833.3721.1
FreeSplatter ‡OAAC + MV72.884.6878.147.2435.7825.5
Amodal3RN/A39.050.4480.065.5238.6446.3
SAM3DN/A34.680.4682.165.5039.0346.5
GENA3D (ours)OAAC33.910.4682.235.6839.1961.9
2 viewsTRELLISOAAC46.233.2877.016.2734.1225.0
FreeSplatterOAAC91.326.7774.2410.4427.7412.4
Amodal3RN/A35.520.4481.845.5238.9247.1
GENA3D (ours)OAAC32.120.4582.335.5039.0866.3
4 viewsTRELLISOAAC45.412.7978.175.9434.4223.2
TRELLIS ‡OAAC + MV44.851.9476.246.7134.2428.6
FreeSplatterOAAC94.104.6475.5810.4832.318.4
Amodal3RN/A35.150.4382.045.5138.8346.2
GENA3D (ours)OAAC30.730.4382.535.4839.4867.0

FID and KID measure render quality, CLIP the similarity to the input, MMD and COV the geometry, and PCS a VLM-rated perceptual coherence (Qwen3-VL 32B, 0–100). A human study agrees with PCS: GENA3D is preferred over Amodal3R in 65.4%, 74.1% and 78.2% of comparisons at 1, 2 and 4 views.

Fig. 5Amodal 3D generation on GSO with 1, 2 and 4 views. For FreeSplatter, Hunyuan-MV generates the extra views from the single input.
Table 2Faithfulness to what was observed: SSIM, PSNR and LPIPS computed only inside the visibility masks, at the input views.
ViewsMethodSSIM ↑PSNR ↑LPIPS ↓
1 viewTRELLIS0.72114.150.325
SAM3D0.79715.540.226
Amodal3R0.81416.410.241
GENA3D (ours)0.83216.520.231
2 viewsTRELLIS0.68613.980.346
Amodal3R0.81916.520.238
GENA3D (ours)0.83716.990.212
4 viewsTRELLIS0.67413.660.357
Amodal3R0.83216.640.236
GENA3D (ours)0.83816.980.208

§4.4 · Appendix BFigures 6, 9, 10, B.1–B.5

In the wild, in the scene

Real captures from COCO (single view) and Mip-NeRF 360 (sparse views), and indoor scenes from Hypersim where many objects hide each other at once.

Fig. 9In-the-wild amodal 3D generation. Single-view examples come from COCO, sparse-view examples from Mip-NeRF 360; Amodal3R is shown for comparison.

§4.3 AblationTables 3–5 · Figures 7, 8

Each piece earns its place

Removing either module hurts both image quality and geometry. Metrics improve or hold steady up to 20 input views, and the method works with different 2D amodal completion front-ends.

Table 3Removing each module, 2 views.
ModelFID ↓MMD (‰) ↓COV (%) ↑
Full model32.125.5039.08
− Gating MLP32.875.5837.68
− View-Wise CA38.645.7638.34
− Stereo-Cond. CA42.925.9135.26
− all proposed46.235.9434.42

Every removal hurts all three metrics. Stereo conditioning matters most, above all for geometry (MMD); view-wise fusion mostly helps render quality (FID).

Table 4Number of input views, up to 20. Changes are relative to the row above.
ViewsFID ↓CLIP (%) ↑MMD (‰) ↓
133.9182.235.52
232.12−1.7982.33+0.105.50−0.02
430.73−1.3982.53+0.205.48−0.02
1029.56−1.1782.98+0.455.45−0.03
2029.43−0.1382.95−0.035.45±0.00

Quality improves or holds steady as views are added, so the model also handles semi-dense captures.

Table 5Different 2D amodal completion front-ends, 2 views.
2D completionFID ↓CLIP (%) ↑COV (%) ↑
pix2gestalt38.6182.0737.26
Flux inpainting33.0784.2639.49
OAAC (default)32.1282.3339.08

Results stay close across front-ends, and stronger completion helps.

Fig. 8Visual ablation. Stereo conditioning keeps the structure faithful even when the 2D completions disagree or contain artifacts.
Fig. 7GENA3D tolerates misaligned or overlapping MVS point clouds.

Citation

BibTeX

@inproceedings{zhou2026gena3d,
  title={GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence},
  author={Zhou, Junwei and Tai, Yu-Wing},
  booktitle={European Conference on Computer Vision},
  pages={318--337},
  year={2026},
  organization={Springer}
}