- Published on
CVPR 2026 — 3D Vision & Depth
3D Vision & Depth
719 papers
1. DirectFisheye-GS: Enabling Native Fisheye Input in Gaussian Splatting with Cross-View Joint Optimization
- Link: Open Access
- arXiv: 2604.00648
2. PanoEnv: Exploring 3D Spatial Intelligence in Panoramic Environments with Reinforcement Learning
- Link: Open Access
- arXiv: 2602.21992
3. REArtGS++: Generalizable Articulation Reconstruction with Temporal Geometry Constraint via Planar Gaussian Splatting
- Link: Open Access
- arXiv: 2511.17059
4. HeroGS: Hierarchical Guidance for Robust 3D Gaussian Splatting under Sparse Views
- Link: Open Access
- arXiv: 2603.01099
5. SRGCD: Stability-Driven Region Growth Framework for 3D Change Detection
- Link: Open Access
6. SAM 3D Body: Robust Full-Body Human Mesh Recovery
- Link: Open Access
- arXiv: 2602.15989
7. VLM-Guided Group Preference Alignment for Diffusion-based Human Mesh Recovery
- Link: Open Access
- arXiv: 2602.19180
8. PhysIR-Splat: Physically Consistent Thermal Infrared Radiative Transfer in 3D Gaussian Splatting
- Link: Open Access
9. Generalized-CVO: Fast and Correspondence-Free Local Point Cloud Registration with Second Order Riemannian Optimization
- Link: Open Access
- arXiv: 2606.10019
10. RT-Splatting: Joint Reflection-Transmission Modeling with Gaussian Splatting
- Link: Open Access
- arXiv: 2605.18263
11. FreeArtGS: Articulated Gaussian Splatting Under Free-moving Scenario
- Link: Open Access
- arXiv: 2603.22102
12. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
- Link: Open Access
- arXiv: 2604.01646
13. CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstruction
- Link: Open Access
14. GauMVC: Generative Decoupled Gaussian Representation for Human-centric Multi-view Video Compression
- Link: Open Access
15. ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
- Link: Open Access
- arXiv: 2602.06226
16. Learning Generalizable 3D Medical Image Representations from Mask-Guided Self-Supervision
- Link: Open Access
- arXiv: 2603.13660
17. BiMotion: B-spline Motion for Text-guided Dynamic 3D Character Generation
- Link: Open Access
- arXiv: 2602.18873
18. EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Grounding
- Link: Open Access
19. C-GenReg: Training-Free 3D Point Cloud Registration by Multi-View-Consistent Geometry-to-Image Generation with Probabilistic Modalities Fusion
- Link: Open Access
- arXiv: 2604.16680
20. MeshFlow: Efficient Artistic Mesh Generation via MeshVAE and Flow-based Diffusion Transformer
- Link: Open Access
21. Refining Few-Step Text-to-Multiview Diffusion via Reinforcement Learning
- Link: Open Access
- arXiv: 2505.20107
22. Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction
- Link: Open Access
- arXiv: 2604.22865
23. RetimeGS: Continuous-Time Reconstruction of 4D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.13783
24. Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy
- Link: Open Access
- arXiv: 2508.04728
25. Prune Wisely, Reconstruct Sharply: Compact 3D Gaussian Splatting via Adaptive Pruning and Difference-of-Gaussian Primitives
- Link: Open Access
- arXiv: 2602.24136
26. Paparazzo: Active Mapping of Moving 3D Objects
- Link: Open Access
- arXiv: 2604.19556
27. QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence
- Link: Open Access
28. DENALI: A Dataset Enabling Non-Line-of-Sight Spatial Reasoning with Low-Cost LiDARs
- Link: Open Access
- arXiv: 2604.16201
29. OnlineHMR: Video-based Online World-Grounded Human Mesh Recovery
- Link: Open Access
- arXiv: 2603.17355
30. MSCD-GS: Motion-Separated Cooperative Deblurring Dynamic Reconstruction via Gaussian Splatting
- Link: Open Access
31. 3D-Aware Multi-Task Learning with Cross-View Correlations for Dense Scene Understanding
- Link: Open Access
- arXiv: 2511.20646
32. Repurposing 3D Generative Model for Autoregressive Layout Generation
- Link: Open Access
- arXiv: 2604.16299
33. RnG: A Unified Transformer for Complete 3D Modeling from Partial Observations
- Link: Open Access
- arXiv: 2603.01194
34. WonderZoom: Multi-Scale 3D World Generation
- Link: Open Access
- arXiv: 2512.09164
35. GeoFree-CoSeg: Unsupervised Point Cloud-Image Cross-Modal Co-Segmentation Without Geometric Alignment
- Link: Open Access
36. Any Resolution Any Geometry: From Multi-View To Multi-Patch
- Link: Open Access
- arXiv: 2603.03026
37. LoG3D: Ultra-High-Resolution 3D Shape Modeling via Local-to-Global Partitioning
- Link: Open Access
- arXiv: 2511.10040
38. PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & Editing
- Link: Open Access
39. LiteSense: Lifting Lightweight ToF with RGB for High-Resolution Metric Depth Estimation
- Link: Open Access
40. Write Where It Matters: Policy-Guided Watermarks for 3D Gaussian Splatting
- Link: Open Access
41. TerraSeg: Self-Supervised Ground Segmentation for Any LiDAR
- Link: Open Access
42. Speeding Up the Learning of 3D Gaussians with Much Shorter Gaussian Lists
- Link: Open Access
- arXiv: 2603.09277
43. Scaling Self-Supervised and Cross-Modal Pretraining for Volumetric CT Transformers
- Link: Open Access
44. OpenVO: Open-World Visual Odometry with Temporal Dynamics Awareness
- Link: Open Access
- arXiv: 2602.19035
45. S2D: Sparse to Dense Lifting for 3D Reconstruction with Minimal Inputs
- Link: Open Access
- arXiv: 2603.10893
46. STAvatar: Soft Binding and Temporal Density Control for Monocular 3D Head Avatars Reconstruction
- Link: Open Access
- arXiv: 2511.19854
47. PiLoT: Neural Pixel-to-3D Registration for UAV-based Ego and Target Geo-localization
- Link: Open Access
- arXiv: 2603.20778
48. Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic Image
- Link: Open Access
- arXiv: 2603.05908
49. MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs
- Link: Open Access
50. A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation
- Link: Open Access
- arXiv: 2511.19004
51. STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
- Link: Open Access
52. Pantheon360: Taming Digital Twin Generation via 3D-Aware 360deg Video Diffusion
- Link: Open Access
53. Correspondence-Attention Alignment for Multi-View Diffusion Models
- Link: Open Access
- arXiv: 2512.03045
54. Geometry-Guided 3D Visual Token Pruning for Video-Language Models
- Link: Open Access
- arXiv: 2604.18260
55. Residual Primitive Fitting of 3D Shapes with SuperFrusta
- Link: Open Access
- arXiv: 2512.09201
56. Anti-Degradation Lifelong Multi-View Clustering
- Link: Open Access
57. ReScene4D: Temporally Consistent Semantic Instance Segmentation of Evolving Indoor 3D Scenes
- Link: Open Access
58. GaussianZoom: Progressive Zoom-in Generative 3D Gaussian Splatting with Geometric and Semantic Guidance
- Link: Open Access
- arXiv: 2605.18252
59. An Instance-Centric Panoptic Occupancy Prediction Benchmark for Autonomous Driving
- Link: Open Access
- arXiv: 2603.27238
60. FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head Avatar from Few Casual Captures
- Link: Open Access
- arXiv: 2603.23345
61. UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditions
- Link: Open Access
62. U4D: Uncertainty-Aware 4D World Modeling from LiDAR Sequences
- Link: Open Access
- arXiv: 2512.02982
63. ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation
- Link: Open Access
- arXiv: 2601.08325
64. Probabilistic Discrepancy Learning for Roadside LiDAR Scene Completion
- Link: Open Access
65. FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation
- Link: Open Access
- arXiv: 2603.01515
66. P2GS: Physical Prior-guided Gaussian Splatting for Photometrically Consistent Urban Reconstruction
- Link: Open Access
- arXiv: 2605.16925
67. MeshSplatting: Differentiable Rendering with Opaque Meshes
- Link: Open Access
- arXiv: 2512.06818
68. CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- Link: Open Access
- arXiv: 2510.20776
69. Bidirectional Cross-Modal Prompting for Event-Frame Asymmetric Stereo
- Link: Open Access
- arXiv: 2604.15312
70. Disco-GS: Gaussian Splatting in Dynamic Color Lighting
- Link: Open Access
71. Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
- Link: Open Access
- arXiv: 2603.21511
72. MoCapAnything: Unified 3D Motion Capture for Arbitrary Skeletons from Monocular Videos
- Link: Open Access
- arXiv: 2512.10881
73. Bootstrapping Multi-view Learning for Test-time Noisy Correspondence
- Link: Open Access
74. Aesthetic Camera Viewpoint Suggestion with 3D Aesthetic Field
- Link: Open Access
- arXiv: 2602.20363
75. MeshRipple: Structured Autoregressive Generation of Artist-Meshes
- Link: Open Access
- arXiv: 2512.07514
76. Text-Image Conditioned 3D Generation
- Link: Open Access
- arXiv: 2603.21295
77. SV-GS: Sparse View 4D Reconstruction with Skeleton-Driven Gaussian Splatting
- Link: Open Access
- arXiv: 2601.00285
78. Think 360deg: Beyond Depth: Evaluating the Width-centric Reasoning Capability of MLLMs
- Link: Open Access
79. ORBIT: Benchmarking SfM in the Wild with 360deg Video
- Link: Open Access
80. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
- Link: Open Access
81. SharpTimeGS: Sharp and Stable Dynamic Gaussian Splatting via Lifespan Modulation
- Link: Open Access
- arXiv: 2602.02989
82. Gau-Occ: Geometry-Completed Gaussians for Multi-Modal 3D Occupancy Prediction
- Link: Open Access
- arXiv: 2603.22852
83. GeoSAM2: Unleashing the Power of SAM2 for 3D Part Segmentation
- Link: Open Access
- arXiv: 2508.14036
84. MV2UV: Generating High-quality UV Texture Maps with Multiview Prompts
- Link: Open Access
- arXiv: 2603.15436
85. Pip-Stereo: Progressive Iterations Pruner for Iterative Optimization based Stereo Matching
- Link: Open Access
- arXiv: 2602.20496
86. Learning Hierarchical Hyperbolic Mixture Model for Part-aware 3D Generation
- Link: Open Access
87. Unblur-SLAM: Dense Neural SLAM for Blurry Inputs
- Link: Open Access
- arXiv: 2603.26810
88. Direction-aware 3D Large Multimodal Models
- Link: Open Access
- arXiv: 2602.19063
89. Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
- Link: Open Access
- arXiv: 2604.19318
90. Lifting Unlabeled Internet-level Data for 3D Scene Understanding
- Link: Open Access
- arXiv: 2604.01907
91. UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
- Link: Open Access
- arXiv: 2605.17742
92. Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
- Link: Open Access
- arXiv: 2601.01695
93. Bootstrap Dynamic-Aware 3D Visual Representation for Scalable Robot Learning
- Link: Open Access
- arXiv: 2512.00074
94. CLIPoint3D: Language-Grounded Few-Shot Unsupervised 3D Point Cloud Domain Adaptation
- Link: Open Access
- arXiv: 2602.20409
95. Sparsity-Aware Voxel Attention and Foreground Modulation for 3D Semantic Scene Completion
- Link: Open Access
- arXiv: 2604.05780
96. Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors
- Link: Open Access
- arXiv: 2603.23324
97. ObjectMorpher: 3D-Aware Image Editing via Deformable 3DGS
- Link: Open Access
- arXiv: 2603.28152
98. NimbusGS: Unified 3D Scene Reconstruction under Hybrid Weather
- Link: Open Access
- arXiv: 2603.27228
99. C-LaV: Conditional Latent Velocity Field Denoising for Weather-Robust LiDAR Place Recognition
- Link: Open Access
100. CoLC: Communication-Efficient Collaborative Perception with LiDAR Completion
- Link: Open Access
- arXiv: 2603.00682
101. FoleyDesigner: Immersive Stereo Foley Generation with Precise Spatio-Temporal Alignment for Film Clips
- Link: Open Access
- arXiv: 2604.05731
102. Adapting Point Cloud Analysis via Multimodal Bayesian Distribution Learning
- Link: Open Access
- arXiv: 2603.22070
103. Physically Inspired Gaussian Splatting for HDR Novel View Synthesis
- Link: Open Access
- arXiv: 2603.28020
104. ManifoldNeuS: Manifold-aware View Optimizability for Pose-Free Neural Surface Reconstruction
- Link: Open Access
105. MatMart: Material Reconstruction of 3D Objects via Diffusion
- Link: Open Access
- arXiv: 2511.18900
106. SparseSplat: Towards Applicable Feed-Forward 3D Gaussian Splatting with Pixel-Unaligned Prediction
- Link: Open Access
- arXiv: 2604.03069
107. E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction
- Link: Open Access
- arXiv: 2603.14684
108. FreeScale: Scaling 3D Scenes via Certainty-Aware Free-View Generation
- Link: Open Access
- arXiv: 2604.10512
109. PoseGaussian: 6D Pose Estimation for Unseen Objects via Sparse-View Object-Level 3D Gaussian Splatting
- Link: Open Access
110. ResiHMR: Residual-Limb Aware Single-Image 3D Human Mesh Recovery for Individuals with Limb Loss
- Link: Open Access
- arXiv: 2604.28025
111. Neural Distribution Prior for LiDAR Out-of-Distribution Detection
- Link: Open Access
- arXiv: 2604.09232
112. HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language Models
- Link: Open Access
- arXiv: 2603.25411
113. EfficientMonoHair: Fast Strand-Level Reconstruction from Monocular Video via Multi-View Direction Fusion
- Link: Open Access
- arXiv: 2604.05794
114. Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud Analysis
- Link: Open Access
- arXiv: 2603.17538
115. Energy-GS: Image Energy-guided Pose Alignment Gaussian Splatting with redesigned pose gradient flow
- Link: Open Access
116. SGSoft: Learning Fused Semantic-Geometric Features for 3D Shape Correspondence via Template-Guided Soft Signals
- Link: Open Access
- arXiv: 2605.18039
117. Neural Gabor Splatting: Enhanced Gaussian Splatting with Neural Gabor for High-frequency Surface Reconstruction
- Link: Open Access
- arXiv: 2604.15941
118. E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimation
- Link: Open Access
- arXiv: 2604.08543
119. Proxy-GS: Unified Occlusion Priors for Training and Inference in Structured 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2509.24421
120. AnyPcc: Compressing Any Point Cloud with a Single Universal Model
- Link: Open Access
- arXiv: 2510.20331
121. ORV: 4D Occupancy-centric Robot Video Generation
- Link: Open Access
- arXiv: 2506.03079
122. SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings
- Link: Open Access
- arXiv: 2601.09665
123. OneOcc: Semantic Occupancy Prediction for Legged Robots with a Single Panoramic Camera
- Link: Open Access
- arXiv: 2511.03571
124. XPaintNet: An eXtreme Lightweight Framework for Stereoscopic Conversion without Inpainting Network
- Link: Open Access
125. PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding
- Link: Open Access
- arXiv: 2601.02457
126. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection
- Link: Open Access
127. VesMamba: 3D Pulmonary Vessel Segmentation from CT images via Mamba with Structural Perception and Scale-aware Filtering
- Link: Open Access
128. ConsisVLA-4D: Advancing Spatiotemporal Consistency in Efficient 3D-Perception and 4D-Reasoning for Robotic Manipulation
- Link: Open Access
- arXiv: 2605.05126
129. Plug-and-Play Incomplete Multi-View Clustering via Janus-Faced Affinity Learning with Topology Harmonization
- Link: Open Access
130. SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
- Link: Open Access
- arXiv: 2603.28760
131. Cross-Instance Gaussian Splatting Registration via Geometry-Aware Feature-Guided Alignment
- Link: Open Access
- arXiv: 2603.21936
132. MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
- Link: Open Access
- arXiv: 2603.27542
133. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Grasping
- Link: Open Access
134. Easy3E: Feed-Forward 3D Asset Editing via Rectified Voxel Flow
- Link: Open Access
- arXiv: 2602.21499
135. Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
- Link: Open Access
- arXiv: 2605.05328
136. PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency
- Link: Open Access
- arXiv: 2604.01791
137. Unlocking 3D Affordance Segmentation with 2D Semantic Knowledge
- Link: Open Access
- arXiv: 2510.08316
138. GaussianDWM: 3D Gaussian Driving World Model for Unified Scene Understanding and Multi-Modal Generation
- Link: Open Access
- arXiv: 2512.23180
139. R3-PCQA: Ray-Reprojection-Reinforcement for No-Reference 3D Point Cloud Quality Assessment
- Link: Open Access
140. Learning Anchor in Dual Orthogonal Space for Fast Multi-view Clustering
- Link: Open Access
141. Are We Ready for RL in Text-to-3D Generation? A Progressive Investigation
- Link: Open Access
- arXiv: 2512.10949
142. MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation
- Link: Open Access
143. 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single Image
- Link: Open Access
- arXiv: 2604.04406
144. CrowdGaussian: Reconstructing High-Fidelity 3D Gaussians for Human Crowd from a Single Image
- Link: Open Access
- arXiv: 2603.17779
145. EmoTaG: Emotion-Aware Talking Head Synthesis on Gaussian Splatting with Few-Shot Personalization
- Link: Open Access
- arXiv: 2603.21332
146. Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator
- Link: Open Access
- arXiv: 2603.14726
147. Material Magic Wand: Material-Aware Grouping of 3D Parts in Untextured Meshes
- Link: Open Access
- arXiv: 2603.17370
148. LEADER: Learning Reliable Local-to-Global Correspondences for LiDAR Relocalization
- Link: Open Access
- arXiv: 2604.11355
149. Unsupervised 3d Motion Estimation Using Event Camera
- Link: Open Access
150. Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature Alignment
- Link: Open Access
- arXiv: 2512.08930
151. Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2512.15508
152. Learning Compact 3D Representations from Feed-Forward Novel View Synthesis
- Link: Open Access
153. SceneTok: A Compressed, Diffusable Token Space for 3D Scenes
- Link: Open Access
- arXiv: 2602.18882
154. MAPo: Motion-Aware Partitioning of Deformable 3D Gaussian Splatting for High-Fidelity Dynamic Scene Reconstruction
- Link: Open Access
- arXiv: 2508.19786
155. EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- Link: Open Access
- arXiv: 2511.18173
156. RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection
- Link: Open Access
- arXiv: 2507.19856
157. High-Precision Dichotomous Image Segmentation via Depth Integrity-Prior and Fine-Grained Patch Strategy
- Link: Open Access
- arXiv: 2503.06100
158. Ani3DHuman: Photorealistic 3D Human Animation with Self-guided Stochastic Sampling
- Link: Open Access
- arXiv: 2602.19089
159. MambaSIC: Mamba-based Stereo Image Compression with Bi-directional Multi-reference Entropy Model
- Link: Open Access
160. ProOOD: Prototype-Guided Out-of-Distribution 3D Occupancy Prediction
- Link: Open Access
- arXiv: 2604.01081
161. FUSER: Feed-Forward Multiview 3D Registration Transformer and SE(3) Diffusion Refinement
- Link: Open Access
- arXiv: 2512.09373
162. PointCSP: Cross-Sample Semantic Propagation and Stability Preservation in Self-Supervised Point Cloud Learning
- Link: Open Access
- arXiv: 2605.01759
163. SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping
- Link: Open Access
164. Affostruction: 3D Affordance Grounding with Generative Reconstruction
- Link: Open Access
- arXiv: 2601.09211
165. Gaussian-Mixture Latent Flow for Stochastic 3D Human Motion Prediction
- Link: Open Access
166. MD2E: Modeling Depth-to-Edge Cues for Monocular Metric Depth Estimation
- Link: Open Access
167. Omni-3DEdit: Generalized Versatile 3D Editing in One-Pass
- Link: Open Access
- arXiv: 2603.17841
168. AffordMatcher: Affordance Learning in 3D Scenes from Visual Signifiers
- Link: Open Access
- arXiv: 2603.27970
169. Exact-GS: Mathematically Rigorous and Accurate 3D Gaussian Splatting for 3D X-ray Reconstruction
- Link: Open Access
170. SDGS: Spatial Difference Guided Gaussian Splatting for Simultaneous Localization and 3D Reconstruction
- Link: Open Access
171. MotionScale: Reconstructing Appearance, Geometry, and Motion of Dynamic Scenes with Scalable 4D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.29296
172. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction
- Link: Open Access
173. Towards Foundation Models for 3D Scene Understanding: Instance-Aware Self-Supervised Learning for Point Clouds
- Link: Open Access
- arXiv: 2603.25165
174. CoRoGS: Contextual Gaussian Splatting for Robust Large-Deviation View Synthesis
- Link: Open Access
175. iSplat: Iterative Learning for Fine-Grained Gaussian Splatting
- Link: Open Access
176. 3D-Aware Implicit Motion Control for View-Adaptive Human Video Generation
- Link: Open Access
- arXiv: 2602.03796
177. KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization
- Link: Open Access
178. Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
- Link: Open Access
- arXiv: 2507.16861
179. PRIMU: Uncertainty Estimation for Novel Views in Gaussian Splatting from Primitive-Based Representations of Error and Coverage
- Link: Open Access
- arXiv: 2508.02443
180. CraftMesh: High-Fidelity Generative Mesh Manipulation via Poisson Seamless Fusion
- Link: Open Access
181. PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
- Link: Open Access
- arXiv: 2603.01650
182. Bezier Degradation Modeling for LiDAR-based Human Motion Capture
- Link: Open Access
183. Adaptive Anisotropic Gaussian Splatting for Multi-contrast MRI Arbitrary-Scale Super-Resolution with Anatomy Guidance
- Link: Open Access
184. DualSplat: Robust 3D Gaussian Splatting via Pseudo-Mask Bootstrapping from Reconstruction Failures
- Link: Open Access
- arXiv: 2604.21631
185. PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image
- Link: Open Access
- arXiv: 2511.13648
186. S-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- Link: Open Access
- arXiv: 2512.01223
187. MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
- Link: Open Access
- arXiv: 2511.18370
188. 4C4D: 4 Camera 4D Gaussian Splatting
- Link: Open Access
- arXiv: 2604.04063
189. CROWn: A Unified Framework for Anti-Aliased Downsampling and Phase-Calibrated Fusion in 3D Medical Segmentation
- Link: Open Access
190. GVLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
- Link: Open Access
- arXiv: 2511.21688
191. UAVLight: A Benchmark for Illumination-Robust 3D Reconstruction in Unmanned Aerial Vehicle (UAV) Scenes
- Link: Open Access
- arXiv: 2511.21565
192. Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesis
- Link: Open Access
- arXiv: 2603.12903
193. E-RayZer: Self-supervised 3D Reconstruction as Spatial Visual Pre-training
- Link: Open Access
- arXiv: 2512.10950
194. RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D Detection
- Link: Open Access
195. SparseWorld-TC: Trajectory-Conditioned Sparse Occupancy World Model
- Link: Open Access
- arXiv: 2511.22039
196. Universal 3D Shape Matching via Coarse-to-Fine Language Guidance
- Link: Open Access
- arXiv: 2602.19112
197. SplatSuRe: Selective Super-Resolution for Multi-view Consistent 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2512.02172
198. SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimation
- Link: Open Access
199. Fast SceneScript: Fast and Accurate Language-Based 3D Scene Understanding via Multi-Token Prediction
- Link: Open Access
- arXiv: 2512.05597
200. Uncertainty-driven 3D Gaussian Splatting Active Mapping via Anisotropic Visibility Field
- Link: Open Access
201. Radiance Meshes for Volumetric Reconstruction
- Link: Open Access
- arXiv: 2512.04076
202. LaS-Comp: Zero-shot 3D Completion with Latent-Spatial Consistency
- Link: Open Access
- arXiv: 2602.18735
203. AIMDepth: Asymmetric Image-Event Mamba for Monocular Depth Estimation
- Link: Open Access
204. Geometry-as-context: Modulating Explicit 3D in Scene-consistent Video Generation to Geometry Context
- Link: Open Access
- arXiv: 2602.21929
205. VoDaSuRe: A Large-Scale Dataset Revealing Domain Shift in Volumetric Super-Resolution
- Link: Open Access
206. TokenHand: Discrete Token Representation for Efficient Hand Mesh Reconstruction
- Link: Open Access
207. Distilling Unsigned Distance Function for Surface Reconstruction from 3D Gaussian Splatting
- Link: Open Access
208. TouchDream: 3D Object Completion through Imagined Touch
- Link: Open Access
209. eRetinexGS: Retinex Modeling for Low-Light Scene Enhancement via Event Streams and 3D Gaussian Splatting
- Link: Open Access
210. The Midas Touch for Metric Depth
- Link: Open Access
- arXiv: 2605.11578
211. Vinedresser3D: Towards Agentic Text-guided 3D Editing
- Link: Open Access
212. Long-SCOPE: Fully Sparse Long-Range Cooperative 3D Perception
- Link: Open Access
- arXiv: 2604.09206
213. TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures
- Link: Open Access
- arXiv: 2602.19679
214. Extend3D: Town-Scale 3D Generation
- Link: Open Access
- arXiv: 2603.29387
215. GeoPredict: Leveraging Predictive Kinematics and 3D Gaussian Geometry for Precise VLA Manipulation
- Link: Open Access
- arXiv: 2512.16811
216. PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
- Link: Open Access
- arXiv: 2512.10840
217. VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- Link: Open Access
- arXiv: 2603.00912
218. UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pair
- Link: Open Access
- arXiv: 2603.19616
219. How Much 3D Do Video Foundation Models Encode?
- Link: Open Access
- arXiv: 2512.19949
220. 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects
- Link: Open Access
- arXiv: 2605.10204
221. VAD-GS: Visibility-Aware Densification for 3D Gaussian Splatting in Dynamic Urban Scenes
- Link: Open Access
222. Z-Order Transformer for Feed-Forward Gaussian Splatting
- Link: Open Access
- arXiv: 2605.13465
223. DepthFocus: Controllable Depth Estimation for See-Through Scenes
- Link: Open Access
- arXiv: 2511.16993
224. Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- Link: Open Access
- arXiv: 2512.11130
225. Representing 3D Faces with Learnable B-Spline Volumes
- Link: Open Access
- arXiv: 2604.12894
226. TGSFormer: Scalable Temporal Gaussian Splatting for Embodied Semantic Scene Completion
- Link: Open Access
- arXiv: 2512.00300
227. Geometric-Aware Hypergraph Reasoning for Novel Class Discovery in Point Cloud Segmentation
- Link: Open Access
228. RARE: Learn to RAnk and REtrieve for Monocular 3D Object Detection
- Link: Open Access
229. CaliTex: Geometry-Calibrated Attention for View-Coherent 3D Texture Generation
- Link: Open Access
- arXiv: 2511.21309
230. DynamicTree: Interactive Real Tree Animation via Sparse Voxel Spectrum
- Link: Open Access
- arXiv: 2510.22213
231. Yo'City: Personalized and Boundless 3D Realistic City Scene Generation via Self-Critic Expansion
- Link: Open Access
232. Talking Together: Synthesizing Co-Located 3D Conversations from Audio
- Link: Open Access
- arXiv: 2603.08674
233. Learning Multi-View Spatial Reasoning from Cross-View Relations
- Link: Open Access
- arXiv: 2603.27967
234. HandDreamer: Zero-Shot Text to 3D Hand Model Generation using Corrective Hand Shape Guidance
- Link: Open Access
- arXiv: 2604.04425
235. Prospective Dynamic 3D MRI Reconstruction via Latent-Space Motion Tracking from Single Measurement
- Link: Open Access
236. PixARMesh: Autoregressive Mesh-Native Single-View Scene Reconstruction
- Link: Open Access
- arXiv: 2603.05888
237. STAC: Plug-and-Play Spatio-Temporal Aware Cache Compression for Streaming 3D Reconstruction
- Link: Open Access
- arXiv: 2603.20284
238. Context-Nav: Context-Driven Exploration and Viewpoint-Aware 3D Spatial Reasoning for Instance Navigation
- Link: Open Access
- arXiv: 2603.09506
239. Seeing through boxes: Non-Line-of-Sight 3D Reconstruction from Radar Signals
- Link: Open Access
240. SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
- Link: Open Access
241. Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking
- Link: Open Access
242. EcoSplat: Efficiency-controllable Feed-forward 3D Gaussian Splatting from Multi-view Images
- Link: Open Access
- arXiv: 2512.18692
243. Dr.Occ: Depth- and Region-Guided 3D Occupancy from Surround-View Cameras for Autonomous Driving
- Link: Open Access
244. Rethinking Pose Refinement in 3D Gaussian Splatting under Pose Prior and Geometric Uncertainty
- Link: Open Access
- arXiv: 2603.16538
245. Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
- Link: Open Access
- arXiv: 2603.24721
246. Grounded 3D-Aware Spatial Vision-Language Modeling
- Link: Open Access
247. LATTICE: Democratize High-Fidelity 3D Generation at Scale
- Link: Open Access
- arXiv: 2512.03052
248. Dual Graph Regularized Deep Unfolding Network for Guided Depth Map Super-resolution
- Link: Open Access
249. Illustrator's Depth: Monocular Layer Index Prediction for Image Decomposition
- Link: Open Access
250. OpenVoxel: Training-Free Grouping and Captioning Voxels for Open-Vocabulary 3D Scene Understanding
- Link: Open Access
- arXiv: 2601.09575
251. ZipMap: Linear-Time Stateful 3D Reconstruction via Test-Time Training
- Link: Open Access
- arXiv: 2603.04385
252. Dehallu3D: Hallucination-Mitigated 3D Generation from a Single Image via Cyclic View Consistency Refinement
- Link: Open Access
253. VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- Link: Open Access
- arXiv: 2512.22351
254. QueryOcc: Query-based Self-Supervision for 3D Semantic Occupancy
- Link: Open Access
- arXiv: 2511.17221
255. MV-TAP: Tracking Any Point in Multi-View Videos
- Link: Open Access
- arXiv: 2512.02006
256. MHopReg: Efficient Hierarchical Multi-Hop Graph Search for Point Cloud Registration
- Link: Open Access
257. EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy
- Link: Open Access
258. ST4R-Splat: Spatio-Temporal Referring Segmentation in 4D Gaussian Splatting
- Link: Open Access
259. Aligning Text, Images and 3D Structure Token-by-Token
- Link: Open Access
- arXiv: 2506.08002
260. LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
- Link: Open Access
- arXiv: 2603.03765
261. CineScene: Implicit 3D as Effective Scene Representation for Cinematic Video Generation
- Link: Open Access
- arXiv: 2602.06959
262. Stochastic Ray Tracing for the Reconstruction of 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.23637
263. FUN REC * Reconstructing Functional 3D Scenes from Egocentric Interaction Videos
- Link: Open Access
264. Task-Driven Implicit Representations for Automated Design of LiDAR Systems
- Link: Open Access
- arXiv: 2505.22344
265. Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and Alignment
- Link: Open Access
- arXiv: 2605.08064
266. MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation
- Link: Open Access
- arXiv: 2604.08916
267. Chorus: Multi-Teacher Pretraining for Holistic 3D Gaussian Scene Encoding
- Link: Open Access
- arXiv: 2512.17817
268. Faster-GS: Analyzing and Improving Gaussian Splatting Optimization
- Link: Open Access
- arXiv: 2602.09999
269. MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior Dynamics
- Link: Open Access
270. Geometry-Aligned and Anomaly-Aware Reconstruction for 3D Anomaly Detection
- Link: Open Access
271. TraceGen: World Modeling in 3D Trace Space Enables Learning from Cross-Embodiment Videos
- Link: Open Access
- arXiv: 2511.21690
272. FluidGaussian: Propagating Simulation-Based Uncertainty Toward Functionally-Intelligent 3D Reconstruction
- Link: Open Access
- arXiv: 2603.21356
273. DreamStereo: Towards Real-Time Stereo Inpainting for HD Videos
- Link: Open Access
- arXiv: 2604.12270
274. AeroDGS: Physically Consistent Dynamic Gaussian Splatting for Single-Sequence Aerial 4D Reconstruction
- Link: Open Access
- arXiv: 2602.22376
275. SMVRT: Implicit Human 3D Modeling Using Sparse Multi-View Volumetric Reconstruction with Transformer Fusion
- Link: Open Access
276. AnchorFlow: Training-Free 3D Editing via Latent Anchor-Aligned Flows
- Link: Open Access
- arXiv: 2511.22357
277. Multimodal Semantic Bias Mitigation for Diverse Text-To-3D Generation
- Link: Open Access
278. Iris: Bringing Real-World Priors into Diffusion Model for Monocular Depth Estimation
- Link: Open Access
- arXiv: 2603.16340
279. Captain Safari: A World Engine with Pose-Aligned 3D Memory
- Link: Open Access
- arXiv: 2511.22815
280. TopoMA: Topology-Guided Multi-Agent Dense RGB 3D Reconstruction via Distributed Inference
- Link: Open Access
281. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
- Link: Open Access
- arXiv: 2603.23276
282. I-Scene: 3D Instance Models are Implicit Generalizable Spatial Learners
- Link: Open Access
- arXiv: 2512.13683
283. Zero-Shot Depth Completion with Vision-Language Model
- Link: Open Access
284. FreqSIC: Frequency-aware Stereo Image Compression with Bi-directional Checkerboard Context Model
- Link: Open Access
285. Stereo World Model: Camera-Guided Stereo Video Generation
- Link: Open Access
- arXiv: 2603.17375
286. MS^2Gait: A Multi-Scale Spatio-Temporal Fusion Network for LiDAR-based Gait Recognition
- Link: Open Access
287. Gen3R: 3D Scene Generation Meets Feed-Forward Reconstruction
- Link: Open Access
- arXiv: 2601.04090
288. REALM: An MLLM-Agent Framework for Open World 3D Reasoning Segmentation and Editing on Gaussian Splatting
- Link: Open Access
- arXiv: 2510.16410
289. Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation
- Link: Open Access
- arXiv: 2604.23604
290. L3DR: 3D-aware LiDAR Diffusion and Rectification
- Link: Open Access
- arXiv: 2602.19064
291. VarSplat: Uncertainty-aware 3D Gaussian Splatting for Robust RGB-D SLAM
- Link: Open Access
- arXiv: 2603.09673
292. Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface
- Link: Open Access
- arXiv: 2511.04029
293. GaussFusion: Improving 3D Reconstruction in the Wild with A Geometry-Informed Video Generator
- Link: Open Access
- arXiv: 2603.25053
294. ShelfOcc: Native 3D Supervision beyond LiDAR for Vision-Based Occupancy Estimation
- Link: Open Access
- arXiv: 2511.15396
295. SpeeDe3DGS: Speedy Deformable 3D Gaussian Splatting with Temporal Pruning and Motion Grouping
- Link: Open Access
- arXiv: 2506.07917
296. FastGS: Training 3D Gaussian Splatting in 100 Seconds
- Link: Open Access
- arXiv: 2511.04283
297. Edit-As-Act: Goal-Regressive Planning for Open-Vocabulary 3D Indoor Scene Editing
- Link: Open Access
- arXiv: 2603.17583
298. OccuFly: A 3D Vision Benchmark for Semantic Scene Completion from the Aerial Perspective
- Link: Open Access
- arXiv: 2512.20770
299. Spatial-SAM: Spatially Consistent 3D Electron Microscopy Segmentation with SDF Memory and Semi-Supervised Learning
- Link: Open Access
300. Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes
- Link: Open Access
- arXiv: 2602.22667
301. NERFIFY: A Multi-Agent Framework for Turning NeRF Papers into Code
- Link: Open Access
- arXiv: 2603.00805
302. Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2505.04594
303. DiffuView: Multi-View Diffusion Pretraining for 3D Aware Robotic Manipulation
- Link: Open Access
304. Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
- Link: Open Access
- arXiv: 2512.19402
305. PromptDepth: Efficient and Promptable Geometric 3D Vision Model for Embodied Intelligence
- Link: Open Access
306. PoInit-of-View: Poisoning Initialization of Views Transfers Across Multiple 3D Reconstruction Systems
- Link: Open Access
- arXiv: 2604.16540
307. RecEdit-Drive: 3D Reconstruction-Guided Spatiotemporal Video Editing for Autonomous Driving Scenes
- Link: Open Access
308. Particulate: Feed-Forward 3D Object Articulation
- Link: Open Access
- arXiv: 2512.11798
309. Ego-1K - A Large-Scale Multiview Video Dataset for Egocentric Vision
- Link: Open Access
- arXiv: 2603.13741
310. Seeing Depth Through Frequency and Motion: A Progressive Training Paradigm for Monocular Depth Estimation
- Link: Open Access
311. MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global Assembly
- Link: Open Access
- arXiv: 2509.19995
312. Seele: A Unified Acceleration Framework for Real-Time Gaussian Splatting on Mobile Devices
- Link: Open Access
313. Gaussian Splatting-based Low-Rank Tensor Representation for Multi-Dimensional Image Recovery
- Link: Open Access
- arXiv: 2511.14270
314. Efficient Encoder-Free Fourier-based 3D Large Multimodal Model
- Link: Open Access
- arXiv: 2602.23153
315. Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction
- Link: Open Access
- arXiv: 2604.08542
316. Eulerian Gaussian Splatting using Hashed Probability Pyramids
- Link: Open Access
317. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
- Link: Open Access
- arXiv: 2604.18476
318. Where, What, Why: Toward Explainable 3D-GS Watermarking
- Link: Open Access
- arXiv: 2603.08809
319. SPE-MVS: Spatial Position Encoding Enhanced Multi-View Stereo with Monocular Depth Priors
- Link: Open Access
320. No Calibration, No Depth, No Problem: Cross-Sensor View Synthesis with 3D Consistency
- Link: Open Access
- arXiv: 2602.23559
321. Hybrid Robust Collaborative Perception with LiDAR-4D Radar Fusion under Adverse Weather Conditions
- Link: Open Access
322. CARD: A Multi-Modal Automotive Dataset for Dense 3D Reconstruction in Challenging Road Topography
- Link: Open Access
323. PQDT: Pseudo-Query Dual Transformer for Robust Point Cloud Restoration
- Link: Open Access
324. Human Interaction-Aware 3D Reconstruction from a Single Image
- Link: Open Access
- arXiv: 2604.05436
325. PolarGuide-GSDR: 3D Gaussian Splatting Driven by Polarization Priors and Deferred Reflection for Real-World Reflective Scenes
- Link: Open Access
- arXiv: 2512.02664
326. Parallelised Differentiable Straightest Geodesics for 3D Meshes
- Link: Open Access
- arXiv: 2603.15780
327. Fusion of Depth and Semantics for Probabilistic Floorplan Localization
- Link: Open Access
328. GaussianVision: Vision-Language Alignment from Compressed Image Representations using 2D Gaussian Splatting
- Link: Open Access
- arXiv: 2509.22615
329. From 3D Pose to Prose: Biomechanics-Grounded Vision-Language Coaching
- Link: Open Access
- arXiv: 2603.26938
330. 3D-IDE: 3D Implicit Depth Emergent
- Link: Open Access
- arXiv: 2604.03296
331. 3D Gaussian Splatting at Arbitrary Resolutions with Compact Proxy Anchors
- Link: Open Access
332. Depth Any Panoramas: A Foundation Model for Panoramic Depth Estimation
- Link: Open Access
- arXiv: 2512.16913
333. Expanding mmWave Datasets for Human Pose Estimation with Unlabeled Data and LiDAR Datasets
- Link: Open Access
- arXiv: 2603.14507
334. MVGGT: Multimodal Visual Geometry Grounded Transformer for Multiview 3D Referring Expression Segmentation
- Link: Open Access
- arXiv: 2601.06874
335. PointWorld: Scaling 3D World Models for In-The-Wild Robotic Manipulation
- Link: Open Access
- arXiv: 2601.03782
336. GeoGuide: Hierarchical Geometric Guidance for Open-Vocabulary 3D Semantic Segmentation
- Link: Open Access
- arXiv: 2603.26260
337. Volumetric Functional Maps
- Link: Open Access
- arXiv: 2506.13212
338. QD-PCQA: Quality-Aware Domain Adaptation for Point Cloud Quality Assessment
- Link: Open Access
- arXiv: 2603.03726
339. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
- Link: Open Access
- arXiv: 2604.03305
340. ShapeR: Robust Conditional 3D Shape Generation from Casual Captures
- Link: Open Access
- arXiv: 2601.11514
341. MatSpray: Fusing 2D Material World Knowledge on 3D Geometry
- Link: Open Access
- arXiv: 2512.18314
342. UniTEX: Universal High Fidelity Generative Texturing for 3D Shapes
- Link: Open Access
- arXiv: 2505.23253
343. Cluster-aware Anchor Learning for Multi-View Clustering
- Link: Open Access
344. Fast3Dcache: Training-free 3D Geometry Synthesis Acceleration
- Link: Open Access
- arXiv: 2511.22533
345. Fast Markov Random Field Optimisation for Topologically Noisy 3D Shape Matching
- Link: Open Access
346. 3DrawAgent: Teaching LLM to Draw in 3D with Early Contrastive Experience
- Link: Open Access
- arXiv: 2604.08042
347. GaussianGrow: Geometry-aware Gaussian Growing from 3D Point Clouds with Text Guidance
- Link: Open Access
- arXiv: 2604.05721
348. Layered 4D-Rotor Gaussian Splatting: A Compressed Representation for Long Dynamic Scenes
- Link: Open Access
349. More Natural, More Real: Object-aware Gaussian Splatting for 3D Visual Decoding from Human Brain
- Link: Open Access
350. Real2Sim2Real: RetinalDepth-64K for Depth Estimation in Posterior Segment Ophthalmic Surgery
- Link: Open Access
351. Learning 3D Shape Fidelity Metric from Real-world Distortions
- Link: Open Access
352. SAGE: Scalable Agentic 3D Scene Generation for Embodied AI
- Link: Open Access
- arXiv: 2602.10116
353. PartGS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2506.17212
354. PE3R: Perception-Efficient 3D Reconstruction
- Link: Open Access
- arXiv: 2503.07507
355. TWINGS: Thin Plate Splines Warp-aligned Initialization for Sparse-View Gaussian Splatting
- Link: Open Access
- arXiv: 2605.22069
356. Intrinsic Image Fusion for Multi-View 3D Material Reconstruction
- Link: Open Access
357. PhysGS: Bayesian-Inferred Gaussian Splatting for Physical Property Estimation
- Link: Open Access
- arXiv: 2511.18570
358. GeoBridge: A Semantic-Anchored Multi-View Foundation Model Bridging Images and Text for Geo-Localization
- Link: Open Access
- arXiv: 2512.02697
359. Underground Plant Exploration: Non-Destructive 3D Root Assessment with GPR Based on Point Graph Neural Network
- Link: Open Access
360. AnchorSplat: Feed-Forward 3D Gaussian Splatting With 3D Geometric Priors
- Link: Open Access
- arXiv: 2604.07053
361. CoSMo3D: Open-World Promptable 3D Semantic Segmentation through LLM-Guided Canonical Spatial Modeling
- Link: Open Access
362. Node-RF: Learning Generalized Continuous Space-Time Scene Dynamics with Neural ODE-based NeRFs
- Link: Open Access
- arXiv: 2603.12078
363. LaRP: Efficient Multi-View Inpainting with Latent Reprojection Priors
- Link: Open Access
364. Geometry-Aware Cross-Modal Graph Alignment for Referring Segmentation in 3D Gaussian Splatting
- Link: Open Access
365. Guardians of the Hair: Rescuing Soft Boundaries in Depth, Stereo, and Novel Views
- Link: Open Access
366. SceMoS: Scene-Aware 3D Human Motion Synthesis by Planning with Geometry-Grounded Tokens
- Link: Open Access
- arXiv: 2602.20476
367. Muses: Designing, Composing, Generating Nonexistent Fantasy 3D Creatures without Training
- Link: Open Access
- arXiv: 2601.03256
368. SuP: Sub-cloud Driven Point Cloud Registration
- Link: Open Access
369. VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth Estimation
- Link: Open Access
- arXiv: 2603.18943
370. HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
- Link: Open Access
- arXiv: 2603.02329
371. Localizing, Structuring, and Rendering: Bridging 3D and 2D Vision-Language-Action Models for Robotic Manipulation
- Link: Open Access
372. Masked-Diffusion Autoencoders for 3D Medical Vision Representation Learning
- Link: Open Access
373. Reliev3R: Relieving Feed-forward 3D Reconstruction from Multi-View Geometric Annotations
- Link: Open Access
374. Point Cloud as a Foreign Language for Multi-modal Large Language Model
- Link: Open Access
- arXiv: 2603.09173
375. Generative Video Motion Editing with 3D Point Tracks
- Link: Open Access
- arXiv: 2512.02015
376. Edges Compete for Trust: Group Relative Edge Optimization for Building Reconstruction from Point Clouds
- Link: Open Access
377. Wave-Former: Through-Occlusion 3D Reconstruction via Wireless Shape Completion
- Link: Open Access
- arXiv: 2511.14152
378. ParkGaussian: Surround-view 3D Gaussian Splatting for Autonomous Parking
- Link: Open Access
- arXiv: 2601.01386
379. MorphAny3D: Unleashing the Power of Structured Latent in 3D Morphing
- Link: Open Access
- arXiv: 2601.00204
380. PackUV: Packed Gaussian UV Maps for 4D Volumetric Video
- Link: Open Access
- arXiv: 2602.23040
381. UniDAC: Universal Metric Depth Estimation for Any Camera
- Link: Open Access
- arXiv: 2603.27105
382. Haptic Neural Fields: Bringing Tactile Interactions to 3D Rendered Scenes
- Link: Open Access
383. SIMSPINE: A Biomechanics-Aware Simulation Framework for 3D Spine Motion Annotation and Benchmarking
- Link: Open Access
- arXiv: 2602.20792
384. Ghosts in the Point Clouds: De-glaring LiDAR in the Transient Domain
- Link: Open Access
385. Geometric-Photometric Event-based 3D Gaussian Ray Tracing
- Link: Open Access
- arXiv: 2512.18640
386. BA-GS: Bayesian Adaptive Gaussian Splatting for SFM-Free 3D Reconstruction
- Link: Open Access
387. BuildAnyPoint: 3D Building Structured Abstraction from Diverse Point Clouds
- Link: Open Access
- arXiv: 2602.23645
388. R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
- Link: Open Access
- arXiv: 2603.11566
389. When CLIP Sees More, It Fights Back Harder: Multi-View Guided Adaptive Counterattacks for Test-Time Adversarial Robustness
- Link: Open Access
390. Radar-Guided Polynomial Fitting for Metric Depth Estimation
- Link: Open Access
- arXiv: 2503.17182
391. 240FPS Stereo Vision from Monocular Mixed Spikes
- Link: Open Access
392. LoST: Level of Semantics Tokenization for 3D Shapes
- Link: Open Access
- arXiv: 2603.17995
393. Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routing
- Link: Open Access
394. HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face Avatars
- Link: Open Access
- arXiv: 2507.02803
395. ExMesh: EXplicit Mesh Reconstruction with Topology Adaptation
- Link: Open Access
396. Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
- Link: Open Access
- arXiv: 2601.05251
397. EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstruction
- Link: Open Access
398. FlexAvatar: Learning Complete 3D Head Avatars with Partial Supervision
- Link: Open Access
399. SAM3D: Scale-controllable Part Segmentation of 3D Point Clouds
- Link: Open Access
400. SCE-Depth: A Spherical Compound Eye Framework for Wide FOV Depth Estimation
- Link: Open Access
401. Learning Spatial-Temporal Consistency for 3D Semantic Scene Completion
- Link: Open Access
402. Deformable Gaussian Occupancy: Decoupling Rigid and Nonrigid Motion with Factorized Distillation
- Link: Open Access
403. Text-Driven 3D Hand Motion Generation from Sign Language Data
- Link: Open Access
- arXiv: 2508.15902
404. UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching
- Link: Open Access
405. : Low-Light Dynamic Gaussian Splatting
- Link: Open Access
406. Native and Compact Structured Latents for 3D Generation
- Link: Open Access
- arXiv: 2512.14692
407. Zoo3D: Zero-Shot 3D Object Detection at Scene Level
- Link: Open Access
- arXiv: 2511.20253
408. FilterGS: Traversal-Free Parallel Filtering and Adaptive Shrinking for Large-Scale LoD 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.23891
409. Unsupervised Multi-Scale Segmentation of 3D Subcellular World with Stable Diffusion Foundation Model
- Link: Open Access
410. LiDAR-to-4DRadar Diffusion Bridge via Cross-Modal Alignment and Translation in Latent Space
- Link: Open Access
411. Test-Time Training for LiDAR Semantic Segmentation under Corruption via Geometric Inlier Discrimination
- Link: Open Access
412. Orthogonal Spatial-Aware Multi-View Anchor Graph Clustering for Incomplete Remote Sensing Data
- Link: Open Access
413. Global-Graph Guided and Local-Graph Weighted Contrastive Learning for Unified Clustering on Incomplete and Noise Multi-View Data
- Link: Open Access
- arXiv: 2512.21516
414. Gallant: Voxel Grid-based Humanoid Locomotion and Local-navigation across 3-D Constrained Terrains
- Link: Open Access
415. Multi-View Hierarchical Alignment Learning for Spatial Transcriptomics
- Link: Open Access
416. Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
- Link: Open Access
- arXiv: 2507.12336
417. 3D-LATTE: Latent Space 3D Editing from Textual Instructions
- Link: Open Access
- arXiv: 2509.00269
418. Generative Diffusion Priors for 3D Mapping of the Dark Universe
- Link: Open Access
419. PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure Matching
- Link: Open Access
- arXiv: 2603.20818
420. Any2Any 3D Diffusion Models with Knowledge Transfer: A Radiotherapy Planning Study
- Link: Open Access
- arXiv: 2605.09622
421. BrepGaussian: CAD reconstruction from Multi-View Images with Gaussian Splatting
- Link: Open Access
- arXiv: 2602.21105
422. Bringing Your Portrait to 3D Presence
- Link: Open Access
- arXiv: 2511.22553
423. Dynamic Visual SLAM using a General 3D Prior
- Link: Open Access
- arXiv: 2512.06868
424. Turbo-GS: Accelerating 3D Gaussian Fitting for High-Resolution Radiance Fields
- Link: Open Access
425. Nestwork: Conditional 3D Furnished House Layout Generation through Latent Heterogeneous Graph Diffusion
- Link: Open Access
426. Lafite: A Generative Latent Field for 3D Native Texturing
- Link: Open Access
- arXiv: 2512.04786
427. PartDiffuser: Part-wise 3D Mesh Generation via Discrete Diffusion
- Link: Open Access
- arXiv: 2511.18801
428. Seeing through Light and Darkness: Sensor-Physics Grounded Deblurring HDR NeRF from Single-Exposure Images and Events
- Link: Open Access
- arXiv: 2601.15475
429. Perceptual 3D Simulation With Physical World Modeling
- Link: Open Access
430. Towards Realistic and Consistent Orbital Video Generation via 3D Foundation Priors
- Link: Open Access
- arXiv: 2604.12309
431. GEM: Generating LiDAR World Model via Deformable Mamba
- Link: Open Access
- arXiv: 2605.07326
432. Featurising Pixels from Dynamic 3D Scenes with Linear In-Context Learners
- Link: Open Access
- arXiv: 2604.26488
433. AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with Backend
- Link: Open Access
- arXiv: 2511.20343
434. PatchScene: Patch-based Voxel Diffusion Model for Large-Scale Scene Completion
- Link: Open Access
435. SwiftTailor: Efficient 3D Garment Generation with Geometry Image Representation
- Link: Open Access
- arXiv: 2603.19053
436. Lite Any Stereo: Efficient Zero-Shot Stereo Matching
- Link: Open Access
- arXiv: 2511.16555
437. GIFSplat: Generative Prior-Guided Iterative Feed-Forward 3D Gaussian Splatting from Sparse Views
- Link: Open Access
- arXiv: 2602.22571
438. Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
- Link: Open Access
- arXiv: 2604.10573
439. REL-SF4PASS: Panoramic Semantic Segmentation with REL Depth Representation and Spherical Fusion
- Link: Open Access
- arXiv: 2601.16788
440. Scene Reconstruction as Mapping Priors for 3D Detection
- Link: Open Access
441. AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
- Link: Open Access
442. InfiniDepth: Arbitrary-Resolution and Fine-Grained Depth Estimation with Neural Implicit Fields
- Link: Open Access
- arXiv: 2601.03252
443. ReLaGS: Relational Language Gaussian Splatting
- Link: Open Access
- arXiv: 2603.17605
444. EmoDiffTalk: Emotion-aware Diffusion for Editable 3D Gaussian Talking Head
- Link: Open Access
445. 3D-Object Perception Transformer (3PT)
- Link: Open Access
446. FunFact: Building Probabilistic Functional 3D Scene Graphs via Factor-Graph Reasoning
- Link: Open Access
447. MuM: Multi-View Masked Image Modeling for 3D Vision
- Link: Open Access
448. LightSplat: Fast and Memory-Efficient Open-Vocabulary 3D Scene Understanding in Five Seconds
- Link: Open Access
- arXiv: 2603.24146
449. Realiz3D: 3D Generation Made Photorealistic via Domain-Aware Learning
- Link: Open Access
- arXiv: 2605.13852
450. EMGauss: Continuous Slice-to-3D Reconstruction via Dynamic Gaussian Modeling in Volume Electron Microscopy
- Link: Open Access
- arXiv: 2512.06684
451. Foundry: Distilling 3D Foundation Models for the Edge
- Link: Open Access
- arXiv: 2511.20721
452. RelightAnyone: A Generalized Relightable 3D Gaussian Head Model
- Link: Open Access
453. TopoMesh: High-Fidelity Mesh Autoencoding via Topological Unification
- Link: Open Access
- arXiv: 2603.24278
454. TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online Reconstruction
- Link: Open Access
- arXiv: 2512.02341
455. VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network
- Link: Open Access
- arXiv: 2605.07552
456. SGS-Intrinsic: Semantic-Invariant Gaussian Splatting for Sparse-View Indoor Inverse Rendering
- Link: Open Access
- arXiv: 2603.27516
457. Sparse-View Localization via Online Neural 3D Regression
- Link: Open Access
458. PhysGen: Physically Grounded 3D Shape Generation for Industrial Design
- Link: Open Access
- arXiv: 2512.00422
459. FMPose3D: monocular 3D pose estimation via flow matching
- Link: Open Access
- arXiv: 2602.05755
460. Bridging the 2D-3D Gap: A Hierarchical Semantic-Geometric Map for Vision Language Navigation
- Link: Open Access
461. Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
- Link: Open Access
- arXiv: 2501.05264
462. Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
- Link: Open Access
- arXiv: 2604.07997
463. TROPHIES: Temporal Reconstruction of Places, Humans, and Cameras from Multi-view Videos
- Link: Open Access
464. SGAD-SLAM: Splatting Gaussians at Adjusted Depth for Better Radiance Fields in RGBD SLAM
- Link: Open Access
- arXiv: 2603.21055
465. GaussianPile: A Unified Sparse Gaussian Splatting Framework for Slice-based Volumetric Reconstruction
- Link: Open Access
- arXiv: 2603.20611
466. Skullptor: High Fidelity 3D Head Reconstruction in Seconds with Multi-View Normal Prediction
- Link: Open Access
- arXiv: 2602.21100
467. ODGS-SLAM: Omnidirectional Gaussian Splatting SLAM
- Link: Open Access
468. GS-ASM: 2DGS-Supervised Active Stereo Matching
- Link: Open Access
469. EV-CGNet: Co-visible Focused 3D-guided 2D Event Keypoint Detection Network
- Link: Open Access
470. Depth Peeling for High-Fidelity Gaussian-Enhanced Surfel Rendering
- Link: Open Access
471. Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
- Link: Open Access
- arXiv: 2604.03972
472. Confidence-Guided Multi-Scale Aggregation for Sparse-View High-Resolution 3D Gaussian Splatting
- Link: Open Access
473. EmbodiedSplat: Online Feed-Forward Semantic 3DGS for Open-Vocabulary 3D Scene Understanding
- Link: Open Access
- arXiv: 2603.04254
474. SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors
- Link: Open Access
- arXiv: 2505.22499
475. SAMosaic3D: Modular Scene Assembly for Real-Time 3D Segment Anything
- Link: Open Access
476. VDFE: Difference-Aware 3D Scene Editing with Non-Intrusive Video Diffusion Priors for Multi-View Consistency and Efficiency
- Link: Open Access
477. Endless World: Real-Time 3D-Aware Long Video Generation
- Link: Open Access
- arXiv: 2512.12430
478. Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modeling
- Link: Open Access
479. MoRE: 3D Visual Geometry Reconstruction Meets Mixture-of-Experts
- Link: Open Access
- arXiv: 2510.27234
480. LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World
- Link: Open Access
- arXiv: 2605.05390
481. Points-to-3D: Structure-Aware 3D Generation with Point Cloud Priors
- Link: Open Access
- arXiv: 2603.18782
482. ViLearn: Accelerating Training Convergence of Image-to-3D Generation via Visibility Learning
- Link: Open Access
483. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
- Link: Open Access
- arXiv: 2604.14563
484. Robust3DGSW: Toward Robust Watermarking for Quantization-Aware 3D Gaussian Splatting
- Link: Open Access
485. PointAlign: Feature-Level Alignment Regularization for 3D Vision-Language Models
- Link: Open Access
- arXiv: 2603.00412
486. OccAny: Generalized Unconstrained Urban 3D Occupancy
- Link: Open Access
- arXiv: 2603.23502
487. Beyond 3D VQAs: Injecting 3D Spatial Priors into Vision-Language Models for Enhanced Geometric Reasoning
- Link: Open Access
488. SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
- Link: Open Access
- arXiv: 2512.10957
489. Block-Sparse Global Attention for Efficient Multi-View Geometry Transformers
- Link: Open Access
- arXiv: 2509.07120
490. GHPT: Real-Time Relightable Gaussian Splatting using Hybrid Path Tracing
- Link: Open Access
491. WorldGen: From Text to Traversable and Interactive 3D Worlds
- Link: Open Access
- arXiv: 2511.16825
492. HumanNOVA: Photorealistic, Universal and Rapid 3D Human Avatar Modeling from a Single Image
- Link: Open Access
493. ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2605.04730
494. Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence
- Link: Open Access
- arXiv: 2605.01450
495. MetricHMSR: Metric Human Mesh and Scene Recovery from Monocular Images
- Link: Open Access
496. SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
- Link: Open Access
497. Voxify3D: Pixel Art Meets Volumetric Rendering
- Link: Open Access
- arXiv: 2512.07834
498. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification
- Link: Open Access
499. R2G: A Multi-View Circuit Graph Benchmark Suite from RTL to GDSII
- Link: Open Access
- arXiv: 2604.08810
500. JUMP-Hand: Learning Joint-wise Uncertainty to Gate Mixture of View Experts for Multi-View 3D Hand Reconstruction
- Link: Open Access
501. Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction
- Link: Open Access
502. FlashMesh: Faster and Better Autoregressive Mesh Synthesis via Structured Speculation
- Link: Open Access
- arXiv: 2511.15618
503. SeeThrough3D: Occlusion Aware 3D Control in Text-to-Image Generation
- Link: Open Access
- arXiv: 2602.23359
504. 3D Gaussian Splatting with Self-Constrained Priors for High Fidelity Surface Reconstruction
- Link: Open Access
- arXiv: 2603.19682
505. Generalizing Visual Geometry Priors to Sparse Gaussian Occupancy Prediction
- Link: Open Access
- arXiv: 2602.21552
506. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal
- Link: Open Access
- arXiv: 2603.28224
507. Scalable Multi-View Subspace Clustering with Tensorized Anchor Guidance
- Link: Open Access
508. Image-Guided Geometric Stylization of 3D Meshes
- Link: Open Access
- arXiv: 2604.07795
509. Zero-Shot Reconstruction of Animatable 3D Avatars with Cloth Dynamics from a Single Image
- Link: Open Access
- arXiv: 2603.14772
510. TokenGS: Decoupling 3D Gaussian Prediction from Pixels with Learnable Tokens
- Link: Open Access
- arXiv: 2604.15239
511. Catalyst4D: High-Fidelity 3D-to-4D Scene Editing via Dynamic Propagation
- Link: Open Access
- arXiv: 2603.12766
512. ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding
- Link: Open Access
513. DROID-SLAM in the Wild
- Link: Open Access
- arXiv: 2603.19076
514. ExtrinSplat: Decoupling Geometry and Semantics for Open-Vocabulary Understanding in 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2509.22225
515. SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2511.06702
516. Dark3R: Learning Structure from Motion in the Dark
- Link: Open Access
- arXiv: 2603.05330
517. MajutsuCity: Language-driven Aesthetic-adaptive City Generation with Controllable 3D Assets and Layouts
- Link: Open Access
- arXiv: 2511.20415
518. OpenDance: Multimodal Controllable 3D Dance Generation with Large-scale Internet Data
- Link: Open Access
- arXiv: 2506.07565
519. AniMimic: Imitating 3D Animation from Video Priors
- Link: Open Access
520. OnlinePG: Online Open-Vocabulary Panoptic Mapping with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2603.18510
521. PanDA: Unsupervised Domain Adaptation for Multimodal 3D Panoptic Segmentation in Autonomous Driving
- Link: Open Access
- arXiv: 2604.19379
522. Human Geometry Distribution for 3D Animation Generation
- Link: Open Access
- arXiv: 2512.07459
523. Towards Visual Query Localization in the 3D World
- Link: Open Access
- arXiv: 2605.01498
524. ArtLLM: Generating Articulated Assets via 3D LLM
- Link: Open Access
- arXiv: 2603.01142
525. Intrinsic Geometry-Appearance Consistency Optimization for Sparse-View Gaussian Splatting
- Link: Open Access
- arXiv: 2603.02893
526. SpatialStack: Layered Geometry-Language Fusion for 3D VLM Spatial Reasoning
- Link: Open Access
- arXiv: 2603.27437
527. Abstract 3D Perception for Spatial Intelligence in Vision-Language Models
- Link: Open Access
- arXiv: 2511.10946
528. Learning to Infer Parameterized Representations of Plants from 3D Scans
- Link: Open Access
- arXiv: 2505.22337
529. Structure-to-Intensity Diffusion for Adverse-Weather LiDAR Generation
- Link: Open Access
530. VIAFormer: Voxel-Image Alignment Transformer for High-Fidelity Voxel Refinement
- Link: Open Access
- arXiv: 2601.13664
531. IDESplat: Iterative Depth Probability Estimation for Generalizable 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2601.03824
532. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images
- Link: Open Access
533. From None to All: Self-Supervised 3D Reconstruction via Novel View Synthesis
- Link: Open Access
- arXiv: 2603.27455
534. Motion 3-to-4: 3D Motion Reconstruction for 4D Synthesis
- Link: Open Access
- arXiv: 2601.14253
535. CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
- Link: Open Access
- arXiv: 2603.05042
536. Spatial Matters: Position-Guided 3D Referring Expression Segmentation
- Link: Open Access
537. Benchmarking PhD-Level Coding in 3D Geometric Computer Vision
- Link: Open Access
- arXiv: 2603.30038
538. SpiderCam: Low-Power Snapshot Depth from Differential Defocus
- Link: Open Access
- arXiv: 2603.17910
539. VoxTell: Free-Text Promptable Universal 3D Medical Image Segmentation
- Link: Open Access
- arXiv: 2511.11450
540. Action-Geometry Prediction with 3D Geometric Prior for Bimanual Manipulation
- Link: Open Access
- arXiv: 2602.23814
541. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- Link: Open Access
- arXiv: 2512.12887
542. Imbalanced View Contribution Evaluation and Refinement for Deep Incomplete Multi-View Clustering
- Link: Open Access
543. Mining Attribute Subspaces for Efficient Fine-tuning of 3D Foundation Models
- Link: Open Access
- arXiv: 2604.10095
544. Vista4D: Video Reshooting with 4D Point Clouds
- Link: Open Access
- arXiv: 2604.21915
545. NVGS: Neural Visibility for Occlusion Culling in 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2511.19202
546. MimicTalker: A Multimodal Interactive and Memory-Enhanced Framework for Real-Time Dyadic 3D Head Generation
- Link: Open Access
547. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection
- Link: Open Access
548. VideoWeaver: Multimodal Multi-View Video-to-Video Transfer for Embodied Agents
- Link: Open Access
- arXiv: 2603.25420
549. TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free Reconstruction
- Link: Open Access
- arXiv: 2603.00697
550. Cross-Domain Few-Shot Segmentation via Multi-view Progressive Adaptation
- Link: Open Access
- arXiv: 2602.05217
551. Think-Then-Generate: Structural Chain-of-Thought Reasoning for Consistent 3D Generation
- Link: Open Access
552. AERGS-SLAM: Auto-Exposure-Robust Stereo 3D Gaussian Splatting SLAM
- Link: Open Access
553. Photo3D: Advancing Photorealistic 3D Generation through Structure-Aligned Detail Enhancement
- Link: Open Access
- arXiv: 2512.08535
554. LiREC-Net: A Target-Free and Learning-Based Network for LiDAR, RGB, and Event Calibration
- Link: Open Access
- arXiv: 2602.21754
555. Cross-Axis Feature Fusion with Joint-Wise Motion Difference Prediction for Text-Based 3D Human Motion Editing
- Link: Open Access
556. Efficient Unrolled Networks for Large-Scale 3D Inverse Problems
- Link: Open Access
- arXiv: 2601.02141
557. Mesh-Pro: Asynchronous Advantage-guided Ranking Preference Optimization for Artist-style Quadrilateral Mesh Generation
- Link: Open Access
- arXiv: 2603.00526
558. Urban-GS: A Unified 3D Gaussian Splatting Framework for Compact and High-Fidelity Aerial-to-Street Reconstruction
- Link: Open Access
559. Masking Matters: Unlocking the Spatial Reasoning Capabilities of LLMs for 3D Scene-Language Understanding
- Link: Open Access
- arXiv: 2512.02487
560. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
- Link: Open Access
- arXiv: 2603.25159
561. What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?
- Link: Open Access
- arXiv: 2504.16930
562. GenieDrive: Towards Physics-Aware Driving World Model with 4D Occupancy Guided Video Generation
- Link: Open Access
- arXiv: 2512.12751
563. Let it Snow! Animating 3D Gaussian Scenes with Dynamic Weather Effects via Physics-Guided Score Distillation
- Link: Open Access
564. Mark4D: Temporally-Consistent Watermarking for 4D Gaussian Splatting
- Link: Open Access
565. 3D Gaussian Splatting from Unposed Spike Stream
- Link: Open Access
566. EXOTIC: External Vision-driven Incomplete Multi-view Classification
- Link: Open Access
567. CompetitorFormer: Mitigating Query Conflicts for 3D Instance Segmentation via Competitive Strategy
- Link: Open Access
568. HOG-Layout: Hierarchical 3D Scene Generation, Optimization and Editing via Vision-Language Models
- Link: Open Access
- arXiv: 2604.10772
569. 4D Local Modeling Toward Dynamic Global Perception for Ambiguity-free Rotation-Invariant Point Cloud Analysis
- Link: Open Access
570. Mocap-2-to-3: Multi-view Lifting for Monocular Motion Recovery with 2D Pretraining
- Link: Open Access
- arXiv: 2503.03222
571. Revisiting 3D Reconstruction Kernels as Low-Pass Filters
- Link: Open Access
- arXiv: 2601.17900
572. 3D sans 3D Scans: Scalable Pre-training from Video-Generated Point Clouds
- Link: Open Access
- arXiv: 2512.23042
573. Rascene: High-Fidelity 3D Scene Imaging with mmWave Communication Signals
- Link: Open Access
- arXiv: 2604.02603
574. Depth Hypothesis Guided Iterative Refinement for Event-Image Monocular Depth Estimation
- Link: Open Access
575. M3DLayout: A Multi-Source Dataset of 3D Indoor Layouts and Structured Descriptions for 3D Generation
- Link: Open Access
- arXiv: 2509.23728
576. Depth Any Endoscopy: Towards Self-Supervised Generalizable Depth Estimation in Monocular Endoscopy
- Link: Open Access
577. VLM-Loc: Localization in Point Cloud Maps via Vision-Language Models
- Link: Open Access
- arXiv: 2603.09826
578. TR2M: Transferring Monocular Relative Depth to Metric Depth with Language Descriptions and Dual-Level Scale-Oriented Contrast
- Link: Open Access
- arXiv: 2506.13387
579. ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
- Link: Open Access
- arXiv: 2603.19776
580. Parallax to Align Them All: An OmniParallax Attention Mechanism for Distributed Multi-View Image Compression
- Link: Open Access
- arXiv: 2603.03615
581. FILTR: Extracting Topological Features from Pretrained 3D Models
- Link: Open Access
- arXiv: 2604.22334
582. MSGNav: Unleashing the Power of Multi-modal 3D Scene Graph for Zero-Shot Embodied Navigation
- Link: Open Access
- arXiv: 2511.10376
583. 3D Space as a Scratchpad for Editable Text-to-Image Generation
- Link: Open Access
- arXiv: 2601.14602
584. Learning 3D Reconstruction with Priors in Test Time
- Link: Open Access
- arXiv: 2604.03878
585. Hyper-PCN: Hypergraph-Based Point Cloud Completion via High-Order Correlation Modeling
- Link: Open Access
586. FISHuman: Fine-grained Single-image 3D Human Reconstruction via Multi-view 4D Remeshing
- Link: Open Access
587. StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift Compensation
- Link: Open Access
588. Multi-Hierarchical Contrastive Spectral Fusion for Multi-View Clustering
- Link: Open Access
589. MoRGS: Efficient Per-Gaussian Motion Reasoning for Streamable Dynamic 3D Scenes
- Link: Open Access
- arXiv: 2603.25042
590. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning
- Link: Open Access
- arXiv: 2602.19206
591. VGG-T: Offline Feed-Forward 3D Reconstruction at Scale
- Link: Open Access
- arXiv: 2602.23361
592. Emergent Extreme-View Geometry in 3D Foundation Models
- Link: Open Access
- arXiv: 2511.22686
593. InstructMix2Mix: Consistent Sparse-View Editing Through Multi-View Model Personalization
- Link: Open Access
- arXiv: 2511.14899
594. Decoding 3D Perception via BrainSSD: Synergistic Fusion of EEG Representations from Static and Dynamic Visual Streams
- Link: Open Access
595. Ov3R: Open-Vocabulary Semantic 3D Reconstruction from RGB Videos
- Link: Open Access
- arXiv: 2507.22052
596. Image-to-Point Cloud Feature Back-Projection for Multimodal Training of 3D Semantic Segmentation
- Link: Open Access
597. 2D-LFM: Lifting Foundation Model without 3D Supervision
- Link: Open Access
598. DualPrim: Compact 3D Reconstruction with Positive and Negative Primitives
- Link: Open Access
- arXiv: 2603.16133
599. TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Tracking
- Link: Open Access
- arXiv: 2512.01329
600. LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Link: Open Access
- arXiv: 2511.20648
601. ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body
- Link: Open Access
- arXiv: 2512.14234
602. Unlocking the Power of Critical Factors for 3D Visual Geometry Estimation
- Link: Open Access
- arXiv: 2604.21713
603. LangRef3DGS: Natural Language-Guided 3D Referential Segmentation from Partial Observations via 3D Gaussian Splatting
- Link: Open Access
604. Elastic3D: Controllable Stereo Video Conversion with Guided Latent Decoding
- Link: Open Access
- arXiv: 2512.14236
605. RAM: Recover Any 3D Human Motion in-the-Wild
- Link: Open Access
- arXiv: 2603.19929
606. Photo-Guided Tooth Segmentation on 3D Oral Scan Model
- Link: Open Access
607. HAD: Hallucination-Aware Diffusion Priors for 3D Reconstruction
- Link: Open Access
- arXiv: 2605.16873
608. Taming Video Models for 3D and 4D Generation via Zero-Shot Camera Control
- Link: Open Access
- arXiv: 2509.15130
609. M4Human: A Large-Scale Multimodal mmWave Radar Benchmark for Human Mesh Reconstruction
- Link: Open Access
- arXiv: 2512.12378
610. ELiC: Efficient LiDAR Geometry Compression via Cross-Bit-depth Feature Propagation and Bag-of-Encoders
- Link: Open Access
- arXiv: 2511.14070
611. PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding
- Link: Open Access
612. Learning Differentiable Hierarchies in 3D Gaussian Splatting
- Link: Open Access
613. Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
- Link: Open Access
- arXiv: 2508.03643
614. Circular-DPO: Aligning Multi-Stage 3D Generative Models via Preference Feedback Loop
- Link: Open Access
615. ScenDi: 3D-to-2D Scene Diffusion Cascades for Urban Generation
- Link: Open Access
- arXiv: 2601.15221
616. EventHub: Data Factory for Generalizable Event-Based Stereo Networks without Active Sensors
- Link: Open Access
- arXiv: 2604.02331
617. GS^2: Graph-based Spatial Distribution Optimization for Compact 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2604.01884
618. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
- Link: Open Access
- arXiv: 2506.21076
619. Deformation-based In-Context Learning for Point Cloud Understanding
- Link: Open Access
- arXiv: 2604.02845
620. WorldStereo: Bridging Camera-Guided Video Generation and Scene Reconstruction via 3D Geometric Memories
- Link: Open Access
- arXiv: 2603.02049
621. SPEAR-1: Scaling Beyond Robot Demonstrations via 3D Understanding
- Link: Open Access
- arXiv: 2511.17411
622. Dense Metric Depth Completion from Sparse Direct Time-of-Flight Sensors
- Link: Open Access
623. MonoVLM: Monocular 3D Visual Grounding with Vision Language Models
- Link: Open Access
624. FG-Portrait: 3D Flow Guided Editable Portrait Animation
- Link: Open Access
- arXiv: 2603.23381
625. ARMFlow: AutoRegressive MeanFlow for Online 3D Human Reaction Generation
- Link: Open Access
- arXiv: 2512.16234
626. GaussianMatch: Semi-Supervised Regression with Pseudo-Label Filtering via Multi-View Gaussian Consistency
- Link: Open Access
627. Curvature-Aware Captioning: Leveraging Geodesic Attention for 3D Scene Understanding
- Link: Open Access
628. MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid Cameras
- Link: Open Access
629. SonoWorld: From One Image to a 3D Audio-Visual Scene
- Link: Open Access
- arXiv: 2603.28757
630. AREA3D: Active Reconstruction Agent with Unified Feed-Forward 3D Perception and Vision-Language Guidance
- Link: Open Access
- arXiv: 2512.05131
631. MANSION: Multi-floor lANguage-to-3D Scene generatIOn for loNg-horizon tasks
- Link: Open Access
- arXiv: 2603.11554
632. GOR-IS: 3D Gaussian Object Removal In the Intrinsic Space
- Link: Open Access
- arXiv: 2605.00498
633. Mamba Learns in Context: Structure-Aware Domain Generalization for Multi-Task Point Cloud Understanding
- Link: Open Access
- arXiv: 2603.20739
634. Color-Encoded Illumination for High-Speed Volumetric Scene Reconstruction
- Link: Open Access
- arXiv: 2604.26920
635. SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D Matching
- Link: Open Access
636. MV-Fashion: Towards Enabling Virtual Try-On and Size Estimation with Multi-View Paired Data
- Link: Open Access
637. Generalizable Sparse-View 3D Reconstruction from Unconstrained Images
- Link: Open Access
- arXiv: 2604.28193
638. Topology-aware Feature Propagation for Unsupervised Non-rigid Point Cloud Correspondence
- Link: Open Access
639. FastEventDGS: Deformable Gaussian Splatting for Fast Dynamic Scenes from a Single Event Camera
- Link: Open Access
640. iLRM: An Iterative Large 3D Reconstruction Model
- Link: Open Access
- arXiv: 2507.23277
641. Test-Time 3D Occupancy Prediction
- Link: Open Access
- arXiv: 2503.08485
642. Scenes as Tokens: Multi-Scale Normal Distributions Transform Tokenizer for General 3D Vision-Language Understanding
- Link: Open Access
- arXiv: 2511.21191
643. ESAM++: Efficient Online 3D Perception on the Edge
- Link: Open Access
644. StereoWorld: Geometry-Aware Monocular-to-Stereo Video Generation
- Link: Open Access
- arXiv: 2512.09363
645. Low-Rank Test-Time Training for Pre-Trained Point Cloud Models
- Link: Open Access
646. SR3R: Rethinking Super-Resolution 3D Reconstruction With Feed-Forward Gaussian Splatting
- Link: Open Access
- arXiv: 2602.24020
647. SoPE: Spherical Coordinate-Based Positional Embedding for Enhancing Spatial Perception of 3D LVLMs
- Link: Open Access
648. Forecasting 3D Scanpaths in Egocentric Video
- Link: Open Access
649. MU-GeNeRF: Multi-view Uncertainty-guided Generalizable Neural Radiance Fields for Distractor-aware Scene
- Link: Open Access
- arXiv: 2604.17965
650. H^2A^2: Homogeneity-Aware and Heterogeneity-Aware Feature Perception for Unified Indoor 3D Object Detection
- Link: Open Access
651. tttLRM: Test-Time Training for Long Context and Autoregressive 3D Reconstruction
- Link: Open Access
- arXiv: 2602.20160
652. MVInverse: Feed-forward Multiview Inverse Rendering in Seconds
- Link: Open Access
653. D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- Link: Open Access
- arXiv: 2512.12622
654. GP-4DGS: Probabilistic 4D Gaussian Splatting from Monocular Video via Variational Gaussian Processes
- Link: Open Access
- arXiv: 2604.02915
655. Noise-Aware Few-Shot Learning through Bi-directional Multi-View Prompt Alignment
- Link: Open Access
- arXiv: 2603.11617
656. ActionMesh: Animated 3D Mesh Generation with Temporal 3D Diffusion
- Link: Open Access
- arXiv: 2601.16148
657. Towards Intrinsic-Aware Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2603.27059
658. Plug-and-Play PDE Optimization for 3D Gaussian Splatting: Toward High-Quality Rendering and Reconstruction
- Link: Open Access
- arXiv: 2509.13938
659. AeroGS: Scale-Aware Gaussian Splatting for Pose-Free Dynamic UAV Scene Reconstruction
- Link: Open Access
660. Learning to Control Physically-simulated 3D Characters via Generating and Mimicking 2D Motions
- Link: Open Access
- arXiv: 2512.08500
661. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
- Link: Open Access
- arXiv: 2603.10703
662. Minimal Constraint Relaxation for Multiview Autocalibration
- Link: Open Access
663. PhysHO: Physics-Based Dynamic 3D Gaussian Human and Object from Monocular Video
- Link: Open Access
664. EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision
- Link: Open Access
- arXiv: 2605.13152
665. Evidential Neural Radiance Fields
- Link: Open Access
- arXiv: 2602.23574
666. Coupled Diffusion Sampling for Training-Free Multi-View Image Editing
- Link: Open Access
- arXiv: 2510.14981
667. From 2D Alignment to 3D Plausibility: Unifying Heterogeneous 2D Priors and Penetration-Free Diffusion for Occlusion-Robust Two-Hand Reconstruction
- Link: Open Access
- arXiv: 2503.17788
668. Order Matters: 3D Shape Generation from Sequential VR Sketches
- Link: Open Access
- arXiv: 2512.04761
669. DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classification
- Link: Open Access
670. Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation
- Link: Open Access
- arXiv: 2512.06158
671. Revisiting Monocular SLAM with Spatio-Temporal Scene Modeling
- Link: Open Access
672. MORE-STEM: Long-Short MemOry REcall and Spatio-TEmporal Consistency Model for Query-Driven 3D/4D Point Cloud Segmentation
- Link: Open Access
673. Multi-view Consistent 3D Gaussian Head Avatars 'without' Multi-view Generation
- Link: Open Access
674. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
- Link: Open Access
- arXiv: 2605.14110
675. RoSAMDepth: Robust Self-supervised Depth Estimation Leveraging Segment Anything Model
- Link: Open Access
676. MLLMSplat: A 2D MLLM-Powered Framework for 3D Gaussian Splatting Understanding, Generation, and Editing
- Link: Open Access
677. Flow4DGS-SLAM: Optical Flow-Guided 4D Gaussian Splatting SLAM
- Link: Open Access
- arXiv: 2604.22339
678. NeuroSeg Meets DINOv3: Transferring 2D Self-Supervised Visual Priors to 3D Neuron Segmentation via DINOv3 Initialization
- Link: Open Access
- arXiv: 2603.23104
679. Rethinking 2D-3D Registration: A Novel Network for High-Value Zone Selection and Representation Consistency Alignment
- Link: Open Access
680. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
- Link: Open Access
- arXiv: 2604.19432
681. mmWaveFlow: Unified Enhancement and Generation of mmWave Human Point Clouds
- Link: Open Access
682. Learning Surgical Robotic Manipulation with 3D Spatial Priors
- Link: Open Access
- arXiv: 2603.03798
683. DSERT-RoLL: Robust Multi-Modal Perception for Diverse Driving Conditions with Stereo Event-RGB-Thermal Cameras, 4D Radar, and Dual-LiDAR
- Link: Open Access
- arXiv: 2604.03685
684. ClipGStream: Clip-Stream Gaussian Splatting for Any Length and Any Motion Multi-View Dynamic Scene Reconstruction
- Link: Open Access
- arXiv: 2604.13746
685. Multi-view Pyramid Transformer: Look Coarser to See Broader
- Link: Open Access
- arXiv: 2512.07806
686. PointGS: Semantic-Consistent Unsupervised 3D Point Cloud Segmentation with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2605.11520
687. Iris: Integrating Language into Diffusion-based Monocular Depth Estimation
- Link: Open Access
- arXiv: 2411.16750
688. Merge3D: Efficient 3D Multimodal LLMs via Joint 2D-3D Token Merging
- Link: Open Access
689. Event-Based Motion Deblurring Using Task-Oriented 3D Gaussian Event Representations
- Link: Open Access
690. SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splatting
- Link: Open Access
- arXiv: 2604.19202
691. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training
- Link: Open Access
692. Changes in Real Time: Online Scene Change Detection with Multi-View Fusion
- Link: Open Access
- arXiv: 2511.12370
693. Reliable Clustering Number Estimation for Contrastive Multi-View Clustering
- Link: Open Access
694. NG-GS: NeRF-guided 3D Gaussian Splatting Segmentation
- Link: Open Access
- arXiv: 2604.14706
695. VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D Reconstruction
- Link: Open Access
- arXiv: 2505.20279
696. PrITTI: Primitive-based Generation of Controllable and Editable 3D Semantic Urban Scenes
- Link: Open Access
- arXiv: 2506.19117
697. Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-View
- Link: Open Access
698. 3D-VCD: Hallucination Mitigation in 3D-LLM Embodied Agents through Visual Contrastive Decoding
- Link: Open Access
- arXiv: 2604.08645
699. ORD: Object-Relation Decoupling for Generalized 3D Visual Grounding
- Link: Open Access
700. Structural Action Transformer for 3D Dexterous Manipulation
- Link: Open Access
- arXiv: 2603.03960
701. PointNSP: Autoregressive 3D Point Cloud Generation with Next-Scale Level-of-Detail Prediction
- Link: Open Access
702. ProgressiveAvatars: Progressive Animatable 3D Gaussian Avatars
- Link: Open Access
- arXiv: 2603.16447
703. SunFaded: Illumination-Aware Gaussian Splatting for Dark Scenes with Camera-Mounted Active Lighting
- Link: Open Access
704. SAQN: Semantic-based Adaptive Query Network for 3D Referring Expression Segmentation
- Link: Open Access
705. PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic Imagery
- Link: Open Access
- arXiv: 2603.17571
706. RAP: Fast Feedforward Rendering-Free Attribute-Guided Primitive Importance Score Prediction for Efficient 3D Gaussian Splatting Processing
- Link: Open Access
- arXiv: 2602.19753
707. Sketch2CT: Multimodal Diffusion for Structure-Aware 3D Medical Volume Generation
- Link: Open Access
- arXiv: 2603.22509
708. Adaptive 3D Perception for Small Aerial Targets Under Sparse Sampling via Reinforcement Learning
- Link: Open Access
709. Hierarchical Visual Relocalization with Nearest View Synthesis from Feature Gaussian Splatting
- Link: Open Access
- arXiv: 2603.29185
710. RemedyGS: Defend 3D Gaussian Splatting Against Computation Cost Attacks
- Link: Open Access
- arXiv: 2511.22147
711. ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
- Link: Open Access
- arXiv: 2604.22202
712. Teaching DINOv3 About Partial 3D Geometry: A Self-Supervised Geometry-Aware Approach
- Link: Open Access
713. UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg Latents
- Link: Open Access
- arXiv: 2512.09435
714. Routing on Demand: DSNet for Efficient Progressive Point Cloud Denoising
- Link: Open Access
715. Dropping Anchor and Spherical Harmonics for Sparse-view Gaussian Splatting
- Link: Open Access
- arXiv: 2602.20933
716. Good Can Sometimes be Bad: A Unified Attack against 3D Point Cloud Classifier by a Flexible Isotropic Resampling
- Link: Open Access
717. ARES: Unifying Asymmetric RGB-Event Stereo for Probabilistic Scene Flow Estimation
- Link: Open Access
718. Learning Explicit Continuous Motion Representation for Dynamic Gaussian Splatting from Monocular Videos
- Link: Open Access
- arXiv: 2603.25058
719. PointTPA: Dynamic Network Parameter Adaptation for 3D Scene Understanding
- Link: Open Access
- arXiv: 2604.04933