Skip to content
RobotWorld
Back to Papers

PAPER DEEP DIVE

3DGS场景编辑相机放置

Look Before You Edit: Attention-Guided Camera Placement and Multi-View Alignment for 3D Gaussian Splatting Editing

Text-driven 3D scene editing with 3D Gaussian Splatting (3DGS) typically applies a 2D diffusion editor to views rendered from fixed training cameras, limiting both the spatial coverage of edits and the user's freedom to target specific objects in complex scenes. We present LB-Edit, a framework that addresses two coupled problems: where to place editing cameras for localized edits, and how to make per-view edits agree with one another so that the 3D scene remains consistent after fine-tuning. First, Attention-Guided Editing Camera Placement (ACP) probes the diffusion model's self- and cross-attention at multiple candidate camera distances to find where attention is well-contained in the region of interest, then places a compact, geometrically diverse editing camera set at that attention-optimal distance. Second, Multi-view Attention Alignment (MAA) steers the editor toward the same edit across views along two axes: it aligns appearance by sharing self-attention features via token-level correspondence, and aligns spatial location by lifting cross-attention maps onto the 3D Gaussians as a shared 3D attention field, suppressing both appearance and spatial drift. Experiments on multi-object and single-object scenes show that our method achieves the highest user preference in instruction fidelity, multi-view consistency, and editing locality, using as few as 5 editing views and reducing latency by up to 7x over existing methods.

Jaeyeon Park, Taeho Kang, Youngki LeeJuly 22, 202611 min read
中文

LB-Edit: Look Before You Edit — Attention-Guided Camera Placement and Multi-View Alignment for 3DGS Editing

Authors: see paper | arXiv:2607.19777 | Journal: TOG | Domain: 3D Gaussian Splatting / text-driven editing | Single NVIDIA RTX A6000

One-Sentence Summary

LB-Edit tackles two coupled problems in text-driven 3DGS editing—where to place editing cameras and how to make per-view edits agree—with Attention-Guided Camera Placement (ACP), which uses the diffusion model's own self/cross-attention statistics to find the attention-optimal camera distance, and Multi-View Attention Alignment (MAA), which synchronizes both self-attention (appearance) and cross-attention (spatial) within a single U-Net forward pass; it achieves the highest user preference with as few as 5 editing views and up to 7x lower latency.

Background and Motivation

3D Gaussian Splatting (3DGS) offers real-time rendering via explicit discrete Gaussian primitives and has become a powerful alternative to NeRF. Leveraging 2D diffusion priors, recent methods demonstrate text-driven 3DGS editing: GaussianEditor introduces language-guided ROI localization, GaussCtrl and VcEdit improve multi-view consistency via attention sharing, and DGE achieves direct editing via extended self-attention inspired by TokenFlow. These methods all apply the 2D editor to views rendered from the fixed COLMAP training cameras—cameras optimized for reconstruction quality, not for editing a specific region.

This design works when the edit target is a single dominant object occupying a large fraction of every training view, but breaks down in practical AR/VR settings where users want to edit a specific small object in a complex multi-object scene: the target may be cropped or observed from too distant a view, preventing reliable attention localization. Moreover, independently applying a 2D editor per view produces two drifts: appearance drift (inconsistent style/shape across views) and spatial drift (mis-aligned edit locations across views).

LB-Edit addresses both coupled problems simultaneously. The core insight is that editing-camera placement should be determined not by reconstruction heuristics but by the diffusion editor's own attention behavior at that camera position. If cross-attention has heavy background activation outside the ROI, or self-attention leaks ROI features to the background, that camera distance is unsuitable for editing. ACP uses attention statistics to select the optimal distance; MAA synchronizes both self/cross-attention in a single forward pass to eliminate both drifts.


Figure 1: LB-Edit overview. COLMAP cameras are optimized for reconstruction, not editing: target objects may be cropped or observed from too-distant views. Given a text-specified ROI, the method selects a compact set of editing-aware views at attention-effective distances, enabling more localized 3DGS edits with substantially lower latency.

Method

Problem Formulation

Given a 3DGS scene $\mathcal{G}$ and a text edit instruction, place $K$ editing cameras $\mathcal{C}=\{c_k\}_{k=1}^{K}$ such that applying a 2D diffusion editor to their rendered views produces ROI-focused edits with high 3D consistency. Let $\mathcal{M}_{\text{ROI}} \subseteq \mathcal{G}$ be the ROI Gaussians obtained from language-guided 2D masks lifted into Gaussian space, with center $\mathbf{c}$ and intrinsic scale $r_{\text{obj}} = \sqrt{\lambda_{\max}}$ from the largest eigenvalue of the opacity-weighted covariance. The canonical front direction is the predominant viewing direction of training cameras:

$$\mathbf{v}_{\text{front}} = \text{normalize}(\mathbf{c}_{\text{scene}} - \mathbf{c})$$

where $\mathbf{c}_{\text{scene}}$ is the mean COLMAP camera position.

Attention-Guided Editing Camera Placement (ACP)


Figure 2: Framework overview—given source 3DGS, ROI mask, and text instruction, ACP selects a compact editing camera set at the attention-optimal distance, MAA synchronizes self/cross-attention in a single forward pass, then 3DGS is fine-tuned.

ACP has two stages: choosing the attention-optimal camera distance, then constructing a diverse camera set at that distance. Distance probing: for candidate distances $d_i = m_i \, r_{\text{obj}}$ (multipliers $m_i$ sampled from a predefined range), a probe camera is placed at $\mathbf{c} + d_i \mathbf{v}_{\text{front}}$ looking toward $\mathbf{c}$, a frontal view is rendered, and InstructPix2Pix is run for only 5 denoising steps (attention spatial structure stabilizes within the first few steps) to extract: (i) the projected 2D ROI mask $M_i$, (ii) the cross-attention map $A_i$ for edit tokens, (iii) the self-attention matrix $\bar{S}_i \in \mathbb{R}^{N \times N}$.

Cross-attention concentration measures whether edit tokens localize on the ROI without background spill. The F1 score $F_i$ between thresholded cross-attention $\hat{A}_i$ and projected ROI mask $M_i$ measures ROI alignment; the 90th-percentile cross-attention value $\ell_i$ in the background measures residual activation:

$$S_i^{\text{ca}} = F_i \cdot (1 - \ell_i)$$

Self-attention concentration measures whether ROI features leak to the background—background tokens drawing from ROI tokens causes edit bleed. Leakage is the average attention from background queries to ROI keys, normalized by ROI occupancy $o_i = \|\mathbf{m}\|_1 / N$ (without normalization, larger ROIs inflate the metric simply because background queries have more ROI keys):

$$\mathcal{L}_i^{\text{sa}} = \frac{1}{o_i} \cdot \frac{\bar{\mathbf{m}}^\top (\bar{S}_i \, \mathbf{m})}{\|\bar{\mathbf{m}}\|_1}$$

where $\mathbf{m} \in \{0,1\}^N$ is the downsampled ROI mask and $\bar{\mathbf{m}} = \mathbf{1} - \mathbf{m}$. The concentration score is $S_i^{\text{sa}} = 1 - \mathcal{L}_i^{\text{sa}}$.

Attention-optimal distance selection jointly maximizes both scores:

$$d^* = \arg\max_{d_i} \; S_i^{\text{ca}} + S_i^{\text{sa}}$$

Both scores share comparable empirical ranges, allowing direct summation. Distances with strong self-attention leakage or weak cross-attention localization are naturally penalized.

Energy-based camera placement: with $d^*$ fixed, candidate directions are sampled on a sphere via Fibonacci lattice, cameras placed at $\mathbf{p}_{\text{cam}} = \mathbf{c} + d^* \mathbf{d}$. Each candidate is scored by:

$$E_k = w_{\text{vis}} \, S_k^{\text{vis}} + w_{\text{can}} \, S_k^{\text{can}}$$

$S_k^{\text{vis}}$ is the projected ROI visibility ratio; $S_k^{\text{can}}$ favors canonical poses aligned with intrinsic ROI axes. From top-energy candidates, $K$ cameras are selected via farthest-point sampling in angular space for diversity.

Multi-View Attention Alignment (MAA)

Even with carefully placed cameras, editing each view independently causes appearance and spatial drift. In the U-Net, self-attention controls appearance coherence and cross-attention controls edit location. Prior work addresses each separately: TokenFlow/DGE synchronizes self-attention via token correspondence, VcEdit synchronizes spatial via 3D-lifted cross-attention. MAA integrates both within a single U-Net forward pass, leveraging ACP's compact camera set for efficiency.

Reference-target framework avoids all-pairs attention exchange. At each diffusion timestep, the $K$ rendered views are partitioned into references $\mathcal{V}_{\text{ref}}$ and targets $\mathcal{V}_{\text{tgt}}$. References perform a full U-Net forward pass establishing self/cross-attention features; targets inherit aligned features through two replacements.


Figure 4: Ablation studies. (a) ACP ablation; (b) MAA consistency editing ablation.

Self-attention alignment follows the TokenFlow/DGE token-correspondence paradigm. References jointly perform Extended Cross-View Self-Attention (keys/values concatenated across references), caching output as $\mathbf{o}_{\text{ref}}^{\text{self}}$. For each target token $p$, the nearest reference token is found via cosine similarity on L2-normalized layer-normalized hidden states:

$$q^\star(p) = \arg\max_q \; \tilde{\mathbf{h}}_i(p)^\top \tilde{\mathbf{h}}_{\text{ref}}(q)$$

The target self-attention output is replaced with the matched reference output:

$$\mathbf{o}_i^{\text{self}}(p) \leftarrow \mathbf{o}_{\text{ref}}^{\text{self}}(q^\star(p))$$

Each target token inherits the appearance representation already harmonized across references, suppressing appearance drift.

Cross-attention alignment lifts cross-attention into a shared 3D field via inverse splatting. Cross-attention probability maps $A_v^{(t)} \in \mathbb{R}^{HW \times |T|}$ from the reference pass are lifted onto ROI Gaussians as a per-Gaussian field $M_{\text{3D}}^{(t)}$, then re-rendered into each target view via differentiable Gaussian rasterization to yield geometrically consistent $\hat{A}_i^{(t)}$. The standard cross-attention output is replaced by:

$$\mathbf{o}_i^{\text{cross}} = \hat{A}_i^{(t)} \, \mathbf{V}_T$$

where $\mathbf{V}_T = W_V E_{\text{text}}[T] \in \mathbb{R}^{|T| \times d}$ are the value projections of edit-relevant tokens. Because all target views query the same 3D attention field, the text-to-spatial mapping is view-consistent by construction, suppressing spatial drift.

Joint synchronization within a single forward pass: both replacements are applied within the same U-Net forward—SA replacement at the attn1 stage, CA replacement at attn2 of each transformer block. While each mechanism has appeared individually in prior work, this is the first 3D editing pipeline that synchronizes both within a single 2D editing pass. Combined with ACP's compact camera set, high-quality multi-view-consistent editing is achieved with as few as 5 editing cameras.

Experimental Results

Evaluation uses five scenes from 3D-OVS (room, covered_desk, blue_sofa—multi-object forward-facing) and IN2N (face, bear), with 20 edit prompts per scene. Baselines are GaussianEditor, VcEdit, DGE—all relying on COLMAP cameras and aligning at most one of self/cross-attention (Table 1). Fine-tuning follows IN2N's loss schedule combining $L_1$ and LPIPS, updating only ROI Gaussians.

SceneGSEditorVcEditDGEOurs
blue_sofa9.707.5411.7913.10
covered_desk4.7216.2219.8418.48
room7.2112.7919.4619.78
face9.5010.1019.3023.00
bear12.0023.8019.2016.20

Table 2: CLIP directional similarity (higher is better). Best on four of five scenes. On bear, VcEdit scores higher, but analysis shows CLIPsim rewards edits that bleed color across the entire image—a pathology ACP explicitly avoids.

A user study (21 participants, 14 tasks) shows LB-Edit achieves the highest preference across instruction fidelity, multi-view consistency, and editing locality, with overall preference of 37.4% far exceeding GSEditor's 24.2%:

CriterionOursGSEditorVcEditDGE
Instruction fidelity41.7%19.8%18.7%19.8%
Multi-view consistency37.7%22.2%22.6%17.5%
Editing locality32.9%30.6%21.4%15.1%
Overall37.4%24.2%20.9%17.5%

Table 3: User study results—percentage of trials on which a method was preferred under each criterion.


Figure 5: Performance-efficiency trade-off on room. With only 5 ACP-selected cameras (89s), the method matches DGE at 60 views (231s) and exceeds VcEdit/GSEditor at 20 views (>14 min), yielding up to 7x latency reduction.

For efficiency, 5 ACP cameras (89s) match DGE's 60-view CLIPsim (231s) and exceed VcEdit/GSEditor at 20 views (>14 min), reducing latency by up to 7x. The speedup can be formalized: to reach equivalent CLIPsim, view count drops from 20-60 to 5, and ACP needs only 5 denoising steps for probing, cutting total latency from ~14 min to 89s.


Figure 3: Qualitative comparison. Baselines fail or apply weak edits when ROI is small in multi-object scenes (yellow boxes) and show appearance/spatial drift across views (red/blue boxes); LB-Edit is more stable due to ACP's attention-optimal distance and MAA's dual synchronization.

System Architecture Diagram

flowchart TB
  subgraph Input["Input"]
    GS["Source 3DGS scene G"]
    ROI["ROI mask M_ROI"]
    TXT["Text edit instruction"]
  end
  subgraph ACP["ACP Attention-Guided Camera Placement"]
    PROBE["Distance probing
d_i = m_i * r_obj
5-step denoise extract attention"] CAC["Cross-attn concentration S_ca
F1*(1-bg activation)"] SAC["Self-attn concentration S_sa
1-leakage"] OPT["Optimal distance d*
argmax(S_ca+S_sa)"] PLACE["Energy camera placement
Fibonacci+farthest-point K cameras"] end subgraph MAA["MAA Multi-View Attention Alignment"] REF["Reference full forward
extended cross-view self-attn"] SA["Self-attn alignment
token cosine match replace o_self"] CA["Cross-attn alignment
3D attention field re-render o_cross"] SYNC["Single-pass joint sync
attn1=SA, attn2=CA"] end FT["3DGS fine-tune
L1+LPIPS update ROI Gaussians only"] GS --> PROBE ROI --> PROBE TXT --> PROBE PROBE --> CAC --> OPT PROBE --> SAC --> OPT OPT --> PLACE --> REF REF --> SA --> SYNC REF --> CA --> SYNC SYNC --> FT

Limitations

CLIPsim-perceived quality mismatch (analyzed by authors). On bear, VcEdit achieves higher CLIPsim, but analysis reveals CLIPsim rewards edits that bleed color across the entire image—a pathology ACP explicitly avoids. This shows automated metrics correlate imperfectly with perceptual quality, especially when edits affect global appearance rather than the target object. The user study mitigates this, but its sample (21 participants, 14 tasks) is limited.

Dependence on 2D editor attention internals. Both ACP and MAA require access to InstructPix2Pix U-Net self/cross-attention internal features, coupling the method to a specific editor architecture. Switching to an editor that does not expose attention or has a different architecture (e.g., Transformer-based diffusion) would require re-adaptation.

Primarily forward-facing scenes evaluated. The five scenes are mostly forward-facing; only bear is forward-facing 360. For fully 360 or large-range orbital scenes, whether ACP's canonical front direction and Fibonacci sphere sampling remain effective is not thoroughly validated.

Editing view count still manually set. While 5-20 views is far fewer than baselines' 20-60, the exact number is still chosen manually by scene complexity, lacking an adaptive mechanism.

Summary and Outlook

LB-Edit elevates two overlooked coupled problems in diffusion-guided 3DGS editing—editing camera placement and multi-view consistency—to first-class design objects. ACP is the first to make editing-camera placement an attention-driven decision, using the editor's own attention statistics to determine the optimal distance rather than reusing COLMAP cameras optimized for reconstruction. MAA is the first to synchronize both self-attention (appearance) and cross-attention (spatial) within a single U-Net forward pass, rather than addressing each in isolation as prior work does. Together they enable the highest user preference and up to 7x latency reduction with as few as 5 editing views.

From a broader perspective, LB-Edit embodies the idea of "using the generative model's own signals to guide its application"—diffusion attention is not just an internal computation mechanism but an observable, measurable, optimizable signal source. This idea generalizes to other 3D tasks guided by generative priors: using diffusion attention concentration to select optimal observation conditions, and 3D-lifting attention fields to ensure cross-view consistency. As 3D editing moves from research toward AR/VR practicality, seemingly engineering questions like "where to place editing cameras" actually determine whether editing can reliably land in complex scenes.

Golden Lines

"Editing-camera placement should be determined not by reconstruction heuristics but by the diffusion editor's own attention behavior at that position."
"Self-attention governs appearance, cross-attention governs location—synchronizing both, rather than each in isolation, is what eliminates both drifts in a single forward pass."

Related Papers

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

WilLaGS: Latent-Conditional 3D Appearance Fields for Robust Gaussian Splatting In-the-Wild

3D Gaussian Splatting (3DGS) delivers real-time and high-fidelity rendering but remains challenged by unconstrained in-the-wild scenes, where drastic appearance variations and transient objects violate multi-view consistency. Existing methods are fundamentally limited by independent and discrete embeddings that struggle to capture continuous environmental changes or model spatially-varying local illumination. To address these limitations, we propose \textbf{WilLaGS}, a unified framework for robust 3D scene reconstruction and generative appearance synthesis under unconstrained settings. Specifically, we introduce a generative appearance model where a $β$-VAE learns a structured and continuous manifold of global appearance. Conditioned on the latent code, we construct a 3D neural appearance field that generates dynamic Tri-Plane features to encode spatially-varying local illumination effects. Furthermore, to suppress transient artifacts, we present a self-supervised perceptual masking mechanism that leverages a Teacher-Student (EMA) architecture to derive a stable scene consensus, robustly identifying inconsistent regions via perceptual discrepancies. Extensive experiments on multiple datasets demonstrate that \textbf{WilLaGS} achieves state-of-the-art performance in reconstruction quality and novel view appearance synthesis, while maintaining real-time rendering efficiency.

3DGSGaussian Splatting新视角合成Aug 28, 2026
DL-SLAM: Enabling High-Fidelity Gaussian Splatting SLAM in Dynamic Environments based on Dual-Level Probability

DL-SLAM: Enabling High-Fidelity Gaussian Splatting SLAM in Dynamic Environments based on Dual-Level Probability

Recent advances in 3D Gaussian Splatting (3DGS) have enabled significant progress in dense dynamic Simultaneous Localization And Mapping (SLAM). Prevailing methods typically discard predefined dynamic objects, ignoring that transiently static objects offer valuable geometric constraints for pose estimation. A recent work attempts to leverage this potential by employing per-pixel uncertainty maps to quantify the magnitude of motion. While this approach enables transiently static objects to enhance pose estimation, it erroneously integrates these objects into the static map, resulting in persistent artifacts. Moreover, its reliance on purely geometric information leads to ambiguous object boundaries in the uncertainty maps. To overcome these limitations, we present DL-SLAM, a monocular Gaussian Splatting SLAM system built upon a novel dual-level probabilistic framework. Our method computes dynamic probability maps by combining semantic and geometric information. These pixel-level probabilities are lifted to 3D and aggregated to derive an object-level dynamic probability for each instance. Object-level probability enables the categorical pruning of dynamic Gaussians, resulting in an artifact-free static map. The static map, in turn, provides a geometrically consistent guidance to refine the pixel-wise probabilities, enhancing their reliability. Experimental results demonstrate that DL-SLAM outperforms existing approaches, improving tracking accuracy by up to 13\% while generating high-fidelity semantic maps.

动态环境Dynamic EnvironmentsSLAMJul 2, 2026
VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

VLK: Learning Humanoid Loco-Manipulation from Synthetic Interactions in Reconstructed Scenes

VLK synthesizes paired vision-language-kinematics supervision inside 3DGS-reconstructed real scenes: it generates navigation and object-interaction trajectories with privileged scene info, renders egocentric views after the fact, and produces 48,000 paired trajectories to train a policy predicting Unitree G1 whole-body motion, enabling sim-to-real perception-based humanoid loco-manipulation.

humanoid人形机器人loco-manipulationJun 29, 2026
TRACE: Ergodic Trajectory Optimization for Active Scene Reconstruction

TRACE: Ergodic Trajectory Optimization for Active Scene Reconstruction

Existing active reconstruction systems with Gaussian-splatting maps select observations greedily, optimizing a single next-best-view (NBV) at each step and connecting the chosen views by short-horizon path planning. This greedy decoupling disregards the global structure of scene information, producing inefficient trajectories that waste sensing capacity in transit between selected views. In this work, we study active reconstruction as an ergodic coverage problem: the time-averaged spatial statistics of the sensor trajectory should match a target information distribution induced by the current map. Our approach derives this target distribution online from uncertainty and visibility, and calculates ergodic trajectories via a kernel-ergodic horizon planner with gradient flow and footprint depletion, closing the loop between mapping and trajectory optimization. We thoroughly evaluate TRACE on the Replica dataset against the Next-Best-View (NBV) baselines, improving PSNR by 1.5 dB. Code: https://github.com/spikelab-jhu/trace-active-reconstruction.

主动感知轨迹优化遍历覆盖Aug 3, 2026