- ECCV 2026
- arXiv 2511.21945
Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence
Dartmouth College
OverviewFigure 1
Imagine in 2D, settle it in 3D
Objects in real scenes are partly hidden. Recovering the full object means inventing what was never seen, while keeping everything that was seen exactly right.
Faithful, but full of holes
Multi-view stereo keeps the visible geometry consistent, yet has nothing to say about the parts no camera saw.
Complete, but contradictory
2D amodal completion imagines the hidden parts convincingly, but each view imagines something different.
2D priors propose, 3D geometry decides
Completed views drive a conditional 3D generator, while partial stereo geometry anchors it to what was actually observed.
Generating complete 3D objects under partial occlusions (i.e., amodal scenarios) is a practically important yet challenging problem, as large portions of object geometry are unobserved in real-world scenarios. Existing approaches either operate directly in 3D, which ensures geometric consistency but often lacks generative expressiveness, or rely on 2D amodal completion, which provides strong appearance priors but does not guarantee reliable 3D structure. This raises a key question: how can we achieve both generative plausibility and geometric coherence in amodal 3D modeling?
To answer this question, we introduce GENA3D (GENerative Amodal 3D), a framework that integrates learned 2D generative priors with explicit 3D geometric reasoning within a conditional 3D generation paradigm. The 2D priors enable the model to plausibly infer diverse occluded content, while the 3D representation enforces multi-view consistency and spatial validity. Our design incorporates a novel View-Wise Cross-Attention for multi-view alignment and a Stereo-Conditioned Cross-Attention to anchor generative predictions in 3D relationships.
By combining generative imagination with structural constraints, GENA3D generates complete and coherent 3D objects from limited observations without sacrificing geometric fidelity. Experiments demonstrate that our method outperforms existing approaches in both synthetic and real-world amodal scenarios, highlighting the effectiveness of bridging 2D priors and 3D coherence in generating plausible and geometrically consistent 3D structures in complex environments.
§3 MethodFigures 2–4
From 1–4 unposed views to one complete object
GENA3D builds on a two-stage 3D generator. Only the first stage, which decides the object's sparse structure, is retrained, and it gains two conditioning modules.
-
Observe
K sparse, unposed views (typically 1–4). Visibility masks and a text prompt come from foundation models such as SAM and Florence-2.
input · Mvis -
Complete in 2D
Each view is amodally completed on its own. The results look plausible but disagree with each other.
OAAC · I, Mocc -
Fuse the views
View-Wise Cross Attention reads every completed view in parallel and fuses them by visibility.
new · §3.1 -
Anchor in 3D
Stereo-Conditioned Cross Attention gates attention with the partial point cloud from an MVS model.
new · §3.2 -
Generate detail
A pretrained amodal SLAT transformer adds appearance and decodes the complete object.
stage 2 · frozen
View-Wise Cross Attention
Conditioning on one view per sampling step, or concatenating all views, lets later views overwrite earlier structure and lets inconsistent completions pile up as geometric drift. GENA3D instead attends to every view in parallel at every step, and weights each view by how much of the object it actually saw.
z'ₙ = CA(z, DINO(iⁿ)) for every view n
z' = (1/K) Σₙ wₙ z'ₙ , wₙ = τₙ / Σⱼ τⱼ
τₙ = mvⁿ / (moⁿ + mvⁿ) visibility ratio
Why it works: the base generator always produces objects in a canonical pose, whatever the camera. Views can therefore be fused without knowing their poses. In 100 GSO objects × 10 viewpoints, 97.6% of generations land in the canonical axis-aligned frame.
Without it (2 views): FID 32.12 → 38.64Stereo-Conditioned Cross Attention
Geometry added as extra tokens is easy to ignore, especially when it is partial and noisy. Here the partial point cloud of the target, lifted from an MVS model (π³) and cut out by the visibility masks, is voxelized at 64³, encoded, and turned into a gate inside the attention logits.
P_O = ∪ₙ { p ∈ MVS(sⁿ) | mvⁿ(p) = 1 }
c_geo = Linear(Patchify(Enc(Voxelize(P_O, 64³))))
g = σ(MLP(c_geo))
α = softmax( QKᵀ/√D + log(g + ε) )
Robust to bad stereo: when the reconstructed points are misaligned or overlap across views, the generative prior takes over and corrects them (Fig. 7).
Without it (2 views): FID 32.12 → 42.92, MMD 5.50 → 5.91§4 ExperimentsTables 1–2 · Figure 5
The most coherent objects at every view count
All 1,030 objects of Google Scanned Objects, with 3D-consistent occlusions and views whose visible part covers 30–70% of the render. Baselines get the same 2D amodal completion (OAAC) where they need it.
| Views | Method | 2D amodal | FID ↓ | KID (%) ↓ | CLIP (%) ↑ | MMD (‰) ↓ | COV (%) ↑ | PCS (%) ↑ |
|---|---|---|---|---|---|---|---|---|
| 1 view | TRELLIS | OAAC | 49.68 | 3.21 | 76.79 | 6.08 | 33.37 | 21.1 |
| FreeSplatter ‡ | OAAC + MV | 72.88 | 4.68 | 78.14 | 7.24 | 35.78 | 25.5 | |
| Amodal3R | N/A | 39.05 | 0.44 | 80.06 | 5.52 | 38.64 | 46.3 | |
| SAM3D | N/A | 34.68 | 0.46 | 82.16 | 5.50 | 39.03 | 46.5 | |
| GENA3D (ours) | OAAC | 33.91 | 0.46 | 82.23 | 5.68 | 39.19 | 61.9 | |
| 2 views | TRELLIS | OAAC | 46.23 | 3.28 | 77.01 | 6.27 | 34.12 | 25.0 |
| FreeSplatter | OAAC | 91.32 | 6.77 | 74.24 | 10.44 | 27.74 | 12.4 | |
| Amodal3R | N/A | 35.52 | 0.44 | 81.84 | 5.52 | 38.92 | 47.1 | |
| GENA3D (ours) | OAAC | 32.12 | 0.45 | 82.33 | 5.50 | 39.08 | 66.3 | |
| 4 views | TRELLIS | OAAC | 45.41 | 2.79 | 78.17 | 5.94 | 34.42 | 23.2 |
| TRELLIS ‡ | OAAC + MV | 44.85 | 1.94 | 76.24 | 6.71 | 34.24 | 28.6 | |
| FreeSplatter | OAAC | 94.10 | 4.64 | 75.58 | 10.48 | 32.31 | 8.4 | |
| Amodal3R | N/A | 35.15 | 0.43 | 82.04 | 5.51 | 38.83 | 46.2 | |
| GENA3D (ours) | OAAC | 30.73 | 0.43 | 82.53 | 5.48 | 39.48 | 67.0 |
FID and KID measure render quality, CLIP the similarity to the input, MMD and COV the geometry, and PCS a VLM-rated perceptual coherence (Qwen3-VL 32B, 0–100). A human study agrees with PCS: GENA3D is preferred over Amodal3R in 65.4%, 74.1% and 78.2% of comparisons at 1, 2 and 4 views.
| Views | Method | SSIM ↑ | PSNR ↑ | LPIPS ↓ |
|---|---|---|---|---|
| 1 view | TRELLIS | 0.721 | 14.15 | 0.325 |
| SAM3D | 0.797 | 15.54 | 0.226 | |
| Amodal3R | 0.814 | 16.41 | 0.241 | |
| GENA3D (ours) | 0.832 | 16.52 | 0.231 | |
| 2 views | TRELLIS | 0.686 | 13.98 | 0.346 |
| Amodal3R | 0.819 | 16.52 | 0.238 | |
| GENA3D (ours) | 0.837 | 16.99 | 0.212 | |
| 4 views | TRELLIS | 0.674 | 13.66 | 0.357 |
| Amodal3R | 0.832 | 16.64 | 0.236 | |
| GENA3D (ours) | 0.838 | 16.98 | 0.208 |
Supplementary360° renders
Every object, all the way around
Turntable renders from the supplementary material, each shown with the occluded input views it was generated from. Switch to Geometry to see the surface normals, hover a card to flip just that one, or click the inputs to enlarge them.
Videos load as you scroll and pause when out of view. Tap a render on touch screens to flip it.
§4.4 · Appendix BFigures 6, 9, 10, B.1–B.5
In the wild, in the scene
Real captures from COCO (single view) and Mip-NeRF 360 (sparse views), and indoor scenes from Hypersim where many objects hide each other at once.
§4.3 AblationTables 3–5 · Figures 7, 8
Each piece earns its place
Removing either module hurts both image quality and geometry. Metrics improve or hold steady up to 20 input views, and the method works with different 2D amodal completion front-ends.
| Model | FID ↓ | MMD (‰) ↓ | COV (%) ↑ |
|---|---|---|---|
| Full model | 32.12 | 5.50 | 39.08 |
| − Gating MLP | 32.87 | 5.58 | 37.68 |
| − View-Wise CA | 38.64 | 5.76 | 38.34 |
| − Stereo-Cond. CA | 42.92 | 5.91 | 35.26 |
| − all proposed | 46.23 | 5.94 | 34.42 |
Every removal hurts all three metrics. Stereo conditioning matters most, above all for geometry (MMD); view-wise fusion mostly helps render quality (FID).
| Views | FID ↓ | CLIP (%) ↑ | MMD (‰) ↓ |
|---|---|---|---|
| 1 | 33.91 | 82.23 | 5.52 |
| 2 | 32.12−1.79 | 82.33+0.10 | 5.50−0.02 |
| 4 | 30.73−1.39 | 82.53+0.20 | 5.48−0.02 |
| 10 | 29.56−1.17 | 82.98+0.45 | 5.45−0.03 |
| 20 | 29.43−0.13 | 82.95−0.03 | 5.45±0.00 |
Quality improves or holds steady as views are added, so the model also handles semi-dense captures.
| 2D completion | FID ↓ | CLIP (%) ↑ | COV (%) ↑ |
|---|---|---|---|
| pix2gestalt | 38.61 | 82.07 | 37.26 |
| Flux inpainting | 33.07 | 84.26 | 39.49 |
| OAAC (default) | 32.12 | 82.33 | 39.08 |
Results stay close across front-ends, and stronger completion helps.
Citation
BibTeX
@inproceedings{zhou2026gena3d, title={GENA3D: Generative Amodal 3D Modeling by Bridging 2D Priors and 3D Coherence}, author={Zhou, Junwei and Tai, Yu-Wing}, booktitle={European Conference on Computer Vision}, pages={318--337}, year={2026}, organization={Springer} }