- Published on
CVPR 2026 — Detection & Recognition
Detection & Recognition
758 papers
1. Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor Distances
- Link: Open Access
2. An Efficient Token Compression Framework for Visual Object Tracking
- Link: Open Access
- arXiv: 2605.08329
3. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning
- Link: Open Access
- arXiv: 2512.15693
4. SRGCD: Stability-Driven Region Growth Framework for 3D Change Detection
- Link: Open Access
5. GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision
- Link: Open Access
- arXiv: 2511.20994
6. MultiAnimate: Pose-Guided Image Animation Made Extensible
- Link: Open Access
- arXiv: 2602.21581
7. Does YOLO Really Need to See Every Training Image in Every Epoch?
- Link: Open Access
- arXiv: 2603.17684
8. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label
- Link: Open Access
- arXiv: 2604.01646
9. CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstruction
- Link: Open Access
10. ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning
- Link: Open Access
- arXiv: 2509.03951
11. Rethinking Occlusion Modeling for UAV Tracking
- Link: Open Access
12. ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos
- Link: Open Access
- arXiv: 2602.06226
13. Make it SING: Analyzing Semantic Invariants in Classifiers
- Link: Open Access
- arXiv: 2603.14610
14. PRISM: Prototype-based Reasoning with Inter-modal Semantic Mining for Interpretable Image Recognition
- Link: Open Access
15. JRM: Joint Reconstruction Model for Multiple Objects without Alignment
- Link: Open Access
- arXiv: 2603.25985
16. EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Grounding
- Link: Open Access
17. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos
- Link: Open Access
- arXiv: 2503.22174
18. Refracting Reality: Generating Images with Realistic Transparent Objects
- Link: Open Access
- arXiv: 2511.17340
19. Tri-Modal Fusion Transformers for UAV-based Object Detection
- Link: Open Access
- arXiv: 2604.16630
20. RetFormer: Multimodal Retrieval for Enhancing Image Recognition
- Link: Open Access
21. Globally Optimal Pose from Orthographic Silhouettes
- Link: Open Access
- arXiv: 2604.09199
22. Mind the Gap: Transferring Labels to Align Object Detection Datasets
- Link: Open Access
23. Decoupled Generative Modeling for Human-Object Interaction Synthesis
- Link: Open Access
- arXiv: 2512.19049
24. Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy
- Link: Open Access
- arXiv: 2508.04728
25. Paparazzo: Active Mapping of Moving 3D Objects
- Link: Open Access
- arXiv: 2604.19556
26. QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence
- Link: Open Access
27. DREAM: Document Recognition with Explicit Adaptive Memory
- Link: Open Access
28. RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos
- Link: Open Access
- arXiv: 2605.17014
29. Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video
- Link: Open Access
- arXiv: 2604.07786
30. Harnessing the Power of Foundation Models for Accurate Material Classification
- Link: Open Access
- arXiv: 2603.17390
31. HamiPose: Hamiltonian Optimization for Unsupervised Domain Adaptive Pose Estimation
- Link: Open Access
32. AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detection
- Link: Open Access
33. Refacade: Editing Object with Given Reference Texture
- Link: Open Access
34. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation
- Link: Open Access
- arXiv: 2603.00493
35. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition
- Link: Open Access
- arXiv: 2603.09921
36. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions
- Link: Open Access
- arXiv: 2603.05970
37. Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU Detection
- Link: Open Access
38. VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving
- Link: Open Access
- arXiv: 2602.20794
39. FBTA: Enabling Single-GPU End-to-End Gigapixel WSI Classification with Feature Bridging and Translation Alignment
- Link: Open Access
40. Agile Deliberation: Concept Deliberation for Subjective Visual Classification
- Link: Open Access
- arXiv: 2512.10821
41. UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVs
- Link: Open Access
42. DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Models
- Link: Open Access
- arXiv: 2512.12633
43. From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding
- Link: Open Access
- arXiv: 2605.15951
44. ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection
- Link: Open Access
45. Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognition
- Link: Open Access
- arXiv: 2603.07911
46. SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection
- Link: Open Access
- arXiv: 2603.26109
47. InterRVOS: Interaction-Aware Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2506.02356
48. ICTPolarReal: A Polarized Reflection and Material Dataset of Real World Objects
- Link: Open Access
- arXiv: 2603.24912
49. Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic Image
- Link: Open Access
- arXiv: 2603.05908
50. MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs
- Link: Open Access
51. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video
- Link: Open Access
- arXiv: 2603.13912
52. T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding
- Link: Open Access
- arXiv: 2603.06973
53. All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermark
- Link: Open Access
- arXiv: 2602.23523
54. Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction
- Link: Open Access
- arXiv: 2602.18996
55. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection
- Link: Open Access
56. FedSDR: Federated Graph Learning with Structural Noise Detection and Reconstruction
- Link: Open Access
57. Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves
- Link: Open Access
- arXiv: 2603.20850
58. DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces
- Link: Open Access
- arXiv: 2602.19565
59. Reinforcing Video Object Segmentation to Think before it Segments
- Link: Open Access
60. Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration
- Link: Open Access
61. FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head Avatar from Few Casual Captures
- Link: Open Access
- arXiv: 2603.23345
62. TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramas
- Link: Open Access
63. Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts
- Link: Open Access
- arXiv: 2503.23348
64. UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditions
- Link: Open Access
65. Your One-Stop Solution for AI-Generated Video Detection
- Link: Open Access
- arXiv: 2601.11035
66. Common Inpainted Objects In-N-Out of Context
- Link: Open Access
- arXiv: 2506.00721
67. One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers
- Link: Open Access
68. Hierarchical Concept Embedding & Pursuit for Interpretable Image Classification
- Link: Open Access
- arXiv: 2602.11448
69. Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection
- Link: Open Access
- arXiv: 2604.02966
70. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection
- Link: Open Access
- arXiv: 2603.08483
71. CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognition
- Link: Open Access
72. OneHOI: Unifying Human-Object Interaction Generation and Editing
- Link: Open Access
- arXiv: 2604.14062
73. FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation
- Link: Open Access
- arXiv: 2603.01515
74. GenMatter: Perceiving Physical Objects with Generative Matter Models
- Link: Open Access
- arXiv: 2604.22160
75. MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images
- Link: Open Access
- arXiv: 2602.06965
76. UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language Modeling
- Link: Open Access
77. Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.19386
78. CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
- Link: Open Access
- arXiv: 2510.20776
79. Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection
- Link: Open Access
- arXiv: 2603.21511
80. GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation
- Link: Open Access
- arXiv: 2512.17495
81. VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition
- Link: Open Access
- arXiv: 2605.02834
82. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding
- Link: Open Access
- arXiv: 2510.09110
83. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition
- Link: Open Access
84. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models
- Link: Open Access
85. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition
- Link: Open Access
86. One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer
- Link: Open Access
- arXiv: 2511.22940
87. Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State Manipulation
- Link: Open Access
- arXiv: 2604.00452
88. Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes
- Link: Open Access
- arXiv: 2604.19318
89. UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation
- Link: Open Access
- arXiv: 2605.17742
90. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection
- Link: Open Access
91. Learnability-Driven Submodular Optimization for Active Roadside 3D Detection
- Link: Open Access
- arXiv: 2601.01695
92. TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition
- Link: Open Access
93. Clothe and Pose
- Link: Open Access
94. RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection
- Link: Open Access
95. TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs
- Link: Open Access
- arXiv: 2512.14698
96. HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition
- Link: Open Access
- arXiv: 2506.04764
97. Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors
- Link: Open Access
- arXiv: 2603.23324
98. ObjectMorpher: 3D-Aware Image Editing via Deformable 3DGS
- Link: Open Access
- arXiv: 2603.28152
99. C-LaV: Conditional Latent Velocity Field Denoising for Weather-Robust LiDAR Place Recognition
- Link: Open Access
100. Enhancing Accuracy of Uncertainty Estimation in Appearance-based Gaze Tracking with Probabilistic Evaluation and Calibration
- Link: Open Access
- arXiv: 2501.14894
101. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation
- Link: Open Access
- arXiv: 2604.10485
102. Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensation
- Link: Open Access
- arXiv: 2604.21772
103. ManifoldNeuS: Manifold-aware View Optimizability for Pose-Free Neural Surface Reconstruction
- Link: Open Access
104. Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectors
- Link: Open Access
105. MatMart: Material Reconstruction of 3D Objects via Diffusion
- Link: Open Access
- arXiv: 2511.18900
106. Phrase-Grounding-Aware Supervised Fine-Tuning for Chart Recognition via Side-Masked Attention
- Link: Open Access
107. E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction
- Link: Open Access
- arXiv: 2603.14684
108. TrackMAE: Video Representation Learning via Track Mask and Predict
- Link: Open Access
- arXiv: 2603.27268
109. Translating Signals to Languages for sEMG-Based Activity Recognition
- Link: Open Access
- arXiv: 2605.22403
110. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics
- Link: Open Access
- arXiv: 2512.00885
111. Composite-Attribute Person Re-Identification via Pose-Guided Disentanglement
- Link: Open Access
112. SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Network
- Link: Open Access
113. PoseGaussian: 6D Pose Estimation for Unseen Objects via Sparse-View Object-Level 3D Gaussian Splatting
- Link: Open Access
114. Neural Distribution Prior for LiDAR Out-of-Distribution Detection
- Link: Open Access
- arXiv: 2604.09232
115. Physical Adversarial Clothing Evades Visible-Thermal Detectors via Non-Overlapping RGB-T Pattern
- Link: Open Access
- arXiv: 2605.04675
116. Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection
- Link: Open Access
117. OSMO: Open-vocabulary Self-eMOtion Tracking
- Link: Open Access
118. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation
- Link: Open Access
- arXiv: 2604.03723
119. Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection
- Link: Open Access
120. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression
- Link: Open Access
121. SpikeTrack: A Spike-driven Framework for Efficient Visual Tracking
- Link: Open Access
- arXiv: 2602.23963
122. Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition
- Link: Open Access
- arXiv: 2602.19385
123. Energy-GS: Image Energy-guided Pose Alignment Gaussian Splatting with redesigned pose gradient flow
- Link: Open Access
124. OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition
- Link: Open Access
- arXiv: 2512.16727
125. Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding
- Link: Open Access
- arXiv: 2511.23151
126. Neural Gabor Splatting: Enhanced Gaussian Splatting with Neural Gabor for High-frequency Surface Reconstruction
- Link: Open Access
- arXiv: 2604.15941
127. Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer
- Link: Open Access
- arXiv: 2603.01000
128. E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimation
- Link: Open Access
- arXiv: 2604.08543
129. SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual Learning
- Link: Open Access
- arXiv: 2512.00539
130. Drift-Resilient Temporal Priors for Visual Tracking
- Link: Open Access
- arXiv: 2604.02654
131. Recovering Physically Plausible Human-Object Interactions from Monocular Videos
- Link: Open Access
132. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos
- Link: Open Access
133. EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer
- Link: Open Access
- arXiv: 2602.20985
134. Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations
- Link: Open Access
- arXiv: 2511.20295
135. CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance
- Link: Open Access
- arXiv: 2603.03281
136. FlowComposer: Composable Flows for Compositional Zero-Shot Learning
- Link: Open Access
- arXiv: 2603.16641
137. NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2603.21069
138. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection
- Link: Open Access
139. Egocentric Visibility-Aware Human Pose Estimation
- Link: Open Access
- arXiv: 2602.23618
140. Visual Grounding for Object Questions
- Link: Open Access
141. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection
- Link: Open Access
- arXiv: 2509.12546
142. D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network
- Link: Open Access
143. Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective
- Link: Open Access
- arXiv: 2603.02629
144. PhaseWin Search Framework Enable Efficient Object-Level Interpretation
- Link: Open Access
- arXiv: 2511.10914
145. Plug-and-Play Incomplete Multi-View Clustering via Janus-Faced Affinity Learning with Topology Harmonization
- Link: Open Access
146. SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild
- Link: Open Access
- arXiv: 2603.28760
147. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction
- Link: Open Access
- arXiv: 2512.11988
148. Real-World Point Tracking with Verifier-Guided Pseudo-Labeling
- Link: Open Access
149. Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classification
- Link: Open Access
150. AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models
- Link: Open Access
151. MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
- Link: Open Access
- arXiv: 2603.27542
152. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Grasping
- Link: Open Access
153. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection
- Link: Open Access
154. Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift
- Link: Open Access
- arXiv: 2605.05328
155. PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency
- Link: Open Access
- arXiv: 2604.01791
156. DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization
- Link: Open Access
- arXiv: 2508.01873
157. Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly Detection
- Link: Open Access
158. DeepProtect: Proactive Face-Swapping Defense using Identity Blending and Attribute Distortion
- Link: Open Access
159. OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation
- Link: Open Access
- arXiv: 2506.02015
160. Protego: User-Centric Pose-Invariant Privacy Protection Against Face Recognition-Induced Digital Footprint Exposure
- Link: Open Access
- arXiv: 2508.02034
161. MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation
- Link: Open Access
162. Adaptive Capacity Autoregressive Visual Tracking
- Link: Open Access
163. EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing
- Link: Open Access
- arXiv: 2603.19224
164. Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection
- Link: Open Access
- arXiv: 2603.24166
165. Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator
- Link: Open Access
- arXiv: 2603.14726
166. DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection
- Link: Open Access
- arXiv: 2603.18757
167. Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detection
- Link: Open Access
168. HierUQ: Hierarchical Uncertainty Quantification with Adaptive Granularity Reconciliation for Degraded Image Classification
- Link: Open Access
169. Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2512.15508
170. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2605.21261
171. EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses
- Link: Open Access
- arXiv: 2511.18173
172. RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection
- Link: Open Access
- arXiv: 2507.19856
173. From Few-way to Many-way: Rethinking Few-shot Fine-grained Image Classification
- Link: Open Access
174. A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Images
- Link: Open Access
175. SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping
- Link: Open Access
176. CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection
- Link: Open Access
- arXiv: 2605.09802
177. Affostruction: 3D Affordance Grounding with Generative Reconstruction
- Link: Open Access
- arXiv: 2601.09211
178. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.00507
179. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction
- Link: Open Access
180. MMGait: Towards Multi-Modal Gait Recognition
- Link: Open Access
- arXiv: 2604.15979
181. PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
- Link: Open Access
- arXiv: 2603.04598
182. Physical Object Understanding with a Physically Controllable World Model
- Link: Open Access
183. GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding
- Link: Open Access
- arXiv: 2604.02093
184. DFD-HR: Generalizable Deepfake Detection via Hierarchical Routing Learning
- Link: Open Access
185. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets
- Link: Open Access
- arXiv: 2304.02296
186. ConsistCompose: Unified Multimodal Layout Control for Image Composition
- Link: Open Access
- arXiv: 2511.18333
187. MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation
- Link: Open Access
- arXiv: 2603.29029
188. KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization
- Link: Open Access
189. Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection
- Link: Open Access
- arXiv: 2507.16861
190. Similarity-Consistent Likelihood Diffusion enables Hidden Person Detection from Wall Reflections
- Link: Open Access
191. Batman: Benign Knowledge Alignment Through Malicious Null Space in Federated Backdoor Attack
- Link: Open Access
192. ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposals
- Link: Open Access
- arXiv: 2602.22666
193. Enabling Supervised Learning of Generative Signatures for Generalized AI-Generated Images Detection
- Link: Open Access
194. TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models
- Link: Open Access
- arXiv: 2512.16523
195. DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning
- Link: Open Access
196. Think Before You Drive: World Model-Inspired Multimodal Grounding
- Link: Open Access
- arXiv: 2512.03454
197. S-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance
- Link: Open Access
- arXiv: 2512.01223
198. UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting
- Link: Open Access
- arXiv: 2604.02905
199. MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer
- Link: Open Access
- arXiv: 2511.18370
200. Advancing Image Classification with Discrete Diffusion Classification Modeling
- Link: Open Access
201. RecoverMark: Robust Watermarking for Localization and Recovery of Manipulated Faces
- Link: Open Access
- arXiv: 2602.20618
202. Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesis
- Link: Open Access
- arXiv: 2603.12903
203. Active Inference for Micro-Gesture Recognition: EFE-Guided Temporal Sampling and Adaptive Learning
- Link: Open Access
- arXiv: 2603.07559
204. DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images
- Link: Open Access
- arXiv: 2512.03004
205. RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D Detection
- Link: Open Access
206. Rounded or Streamlined Head? Bridging Concept Bottleneck Models and Attribute-Described Object Parts
- Link: Open Access
207. MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- Link: Open Access
- arXiv: 2512.02906
208. PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training
- Link: Open Access
- arXiv: 2604.00503
209. UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register
- Link: Open Access
- arXiv: 2605.19622
210. Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking
- Link: Open Access
211. Distilling Unsigned Distance Function for Surface Reconstruction from 3D Gaussian Splatting
- Link: Open Access
212. TouchDream: 3D Object Completion through Imagined Touch
- Link: Open Access
213. Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking
- Link: Open Access
214. TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures
- Link: Open Access
- arXiv: 2602.19679
215. PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning
- Link: Open Access
- arXiv: 2512.10840
216. VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection
- Link: Open Access
- arXiv: 2603.00912
217. UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pair
- Link: Open Access
- arXiv: 2603.19616
218. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers
- Link: Open Access
- arXiv: 2505.21541
219. Fast Spatial Tracking with Visual Geometry Transformer
- Link: Open Access
220. EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation
- Link: Open Access
- arXiv: 2604.01421
221. 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects
- Link: Open Access
- arXiv: 2605.10204
222. MetaSpectra+: A Compact Broadband Metasurface Camera for Snapshot Hyperspectral+ Imaging
- Link: Open Access
- arXiv: 2603.09116
223. TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection
- Link: Open Access
224. Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objects
- Link: Open Access
225. From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking
- Link: Open Access
- arXiv: 2512.02392
226. Representing 3D Faces with Learnable B-Spline Volumes
- Link: Open Access
- arXiv: 2604.12894
227. Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object Detection
- Link: Open Access
228. Gamba: Mamba-based graph convolutional network with dynamic graph topology learning for action recognition
- Link: Open Access
229. RARE: Learn to RAnk and REtrieve for Monocular 3D Object Detection
- Link: Open Access
230. Prospective Dynamic 3D MRI Reconstruction via Latent-Space Motion Tracking from Single Measurement
- Link: Open Access
231. R-4B: Incentivizing General-Purpose Auto-Thinking in MLLMs via Bi-Mode Annealing and Reinforce Learning
- Link: Open Access
232. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization
- Link: Open Access
- arXiv: 2603.03967
233. SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
- Link: Open Access
234. Frequency-domain Manipulation for Face Obfuscation
- Link: Open Access
235. Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking
- Link: Open Access
236. Rethinking Pose Refinement in 3D Gaussian Splatting under Pose Prior and Geometric Uncertainty
- Link: Open Access
- arXiv: 2603.16538
237. Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models
- Link: Open Access
- arXiv: 2603.24721
238. Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation
- Link: Open Access
- arXiv: 2603.12845
239. ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors
- Link: Open Access
- arXiv: 2512.09056
240. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection
- Link: Open Access
- arXiv: 2603.07142
241. Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting
- Link: Open Access
- arXiv: 2603.16129
242. RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
- Link: Open Access
- arXiv: 2603.03617
243. PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning
- Link: Open Access
- arXiv: 2605.13467
244. PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation
- Link: Open Access
- arXiv: 2603.22193
245. Condensed Test-Time Adaptation of VLMs for Action Recognition
- Link: Open Access
246. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Images
- Link: Open Access
- arXiv: 2512.24074
247. Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
- Link: Open Access
- arXiv: 2510.04225
248. TGTrack: Temporal Generative Learning for Unified Single Object Tracking
- Link: Open Access
249. OrienPose: Orientation-Guided Novel View Synthesis for Single-Image Unseen Object Pose Estimation
- Link: Open Access
250. VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement
- Link: Open Access
- arXiv: 2512.22351
251. ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions
- Link: Open Access
- arXiv: 2603.25791
252. MV-TAP: Tracking Any Point in Multi-View Videos
- Link: Open Access
- arXiv: 2512.02006
253. ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning
- Link: Open Access
- arXiv: 2508.03088
254. Humanoid Generative Pre-Training for Zero-Shot Motion Tracking
- Link: Open Access
255. Revisiting Unknowns: Towards Effective and Efficient Open-Set Active Learning
- Link: Open Access
- arXiv: 2603.07898
256. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection
- Link: Open Access
- arXiv: 2604.04444
257. InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models
- Link: Open Access
- arXiv: 2504.05662
258. Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping
- Link: Open Access
- arXiv: 2601.15288
259. Portable Active Learning for Object Detection
- Link: Open Access
- arXiv: 2605.10349
260. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation
- Link: Open Access
- arXiv: 2601.03054
261. PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection
- Link: Open Access
- arXiv: 2603.06917
262. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection
- Link: Open Access
263. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models
- Link: Open Access
- arXiv: 2603.25250
264. FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty Modeling
- Link: Open Access
265. DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognition
- Link: Open Access
266. PAF: Perturbation-Aware Filtering for Open-Set Semi-Supervised Learning
- Link: Open Access
267. M3Grounder: Mask-Based Multi-Span and Multi-Granular Grounding for Document QA
- Link: Open Access
268. Towards Persistence: Learning Topological Constraints for Event-based Small Object Detection
- Link: Open Access
269. ShadowDraw: From Any Object to Shadow-Drawing Compositional Art
- Link: Open Access
- arXiv: 2512.05110
270. MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior Dynamics
- Link: Open Access
271. Geometry-Aligned and Anomaly-Aware Reconstruction for 3D Anomaly Detection
- Link: Open Access
272. WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2602.23029
273. CVA: Context-aware Video-text Alignment for Video Temporal Grounding
- Link: Open Access
- arXiv: 2603.24934
274. APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation
- Link: Open Access
- arXiv: 2602.00551
275. Enhancing Part-Level Point Grounding for Any Open-Source MLLMs
- Link: Open Access
276. Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic Segmentation
- Link: Open Access
277. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection
- Link: Open Access
- arXiv: 2603.11521
278. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models
- Link: Open Access
279. Captain Safari: A World Engine with Pose-Aligned 3D Memory
- Link: Open Access
- arXiv: 2511.22815
280. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection
- Link: Open Access
- arXiv: 2603.23276
281. FedCART: Tackling Long-Tailed Distributions in Federated Adversarial Training via Classifier Refinement
- Link: Open Access
282. OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding
- Link: Open Access
- arXiv: 2604.25276
283. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel
- Link: Open Access
284. SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion
- Link: Open Access
- arXiv: 2603.12764
285. TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models
- Link: Open Access
- arXiv: 2604.15756
286. Fine-Grained Multi Image Object Hallucination Benchmark
- Link: Open Access
287. MS^2Gait: A Multi-Scale Spatio-Temporal Fusion Network for LiDAR-based Gait Recognition
- Link: Open Access
288. Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation
- Link: Open Access
- arXiv: 2604.23604
289. Object-Generalized Re-Identification: A Step Towards Universal Instance Perception
- Link: Open Access
290. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
- Link: Open Access
- arXiv: 2603.00512
291. Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection
- Link: Open Access
- arXiv: 2603.04337
292. Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection
- Link: Open Access
- arXiv: 2603.24030
293. Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition
- Link: Open Access
- arXiv: 2604.07884
294. Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface
- Link: Open Access
- arXiv: 2511.04029
295. Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images
- Link: Open Access
- arXiv: 2604.19257
296. Distribution-Aligned Multimodal Fusion for Robust Object Detection
- Link: Open Access
297. SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names
- Link: Open Access
298. BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detection
- Link: Open Access
- arXiv: 2603.16645
299. Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection
- Link: Open Access
- arXiv: 2603.25203
300. AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence
- Link: Open Access
- arXiv: 2604.10579
301. QuCNet: Quantum Deep Learning Driven Multi-Circuit Network for Remote Sensing Image Classification
- Link: Open Access
302. ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation
- Link: Open Access
- arXiv: 2605.16080
303. GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- Link: Open Access
- arXiv: 2512.02715
304. Ground Reaction Inertial Poser: Physics-based Human Motion Capture from Sparse IMUs and Insole Pressure Sensors
- Link: Open Access
- arXiv: 2603.16233
305. Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes
- Link: Open Access
- arXiv: 2602.22667
306. Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2505.04594
307. Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection
- Link: Open Access
308. Revisiting Pose Sensitivity in Splat-based Computed Tomography under Sparse-view Reconstruction
- Link: Open Access
309. Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface
- Link: Open Access
- arXiv: 2512.19402
310. MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding
- Link: Open Access
- arXiv: 2603.28120
311. Particulate: Feed-Forward 3D Object Articulation
- Link: Open Access
- arXiv: 2512.11798
312. CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data Classification
- Link: Open Access
313. Lyapunov Probes for Hallucination Detection in Large Foundation Models
- Link: Open Access
- arXiv: 2603.06081
314. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection
- Link: Open Access
- arXiv: 2604.18476
315. Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis
- Link: Open Access
316. UIKA: Fast Universal Head Avatar from Pose-Free Images
- Link: Open Access
- arXiv: 2601.07603
317. VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer
- Link: Open Access
- arXiv: 2603.07952
318. Streamlined Open-Vocabulary Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2603.27500
319. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection
- Link: Open Access
320. BulletTime: Decoupled Control of Time and Camera Pose for Video Generation
- Link: Open Access
- arXiv: 2512.05076
321. Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness
- Link: Open Access
- arXiv: 2603.08309
322. OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments
- Link: Open Access
- arXiv: 2603.02390
323. No Way To Steal My Face: Proactive Defense Against Identity-Preserving Personalized Generation
- Link: Open Access
324. From 3D Pose to Prose: Biomechanics-Grounded Vision-Language Coaching
- Link: Open Access
- arXiv: 2603.26938
325. Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective
- Link: Open Access
- arXiv: 2508.06878
326. Beyond Reassembly: Fractured Object Recovery with Missing Parts
- Link: Open Access
327. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation
- Link: Open Access
328. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination
- Link: Open Access
329. Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos
- Link: Open Access
- arXiv: 2603.00938
330. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentation
- Link: Open Access
331. Expanding mmWave Datasets for Human Pose Estimation with Unlabeled Data and LiDAR Datasets
- Link: Open Access
- arXiv: 2603.14507
332. Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraints
- Link: Open Access
333. ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models
- Link: Open Access
- arXiv: 2509.15695
334. Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision
- Link: Open Access
335. Hyperbolic Defect Feature Synthesis for Few-Shot Defect Classification
- Link: Open Access
336. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis
- Link: Open Access
- arXiv: 2604.03305
337. Beyond Appearance: Camouflaged Object Detection via Geometric Structure
- Link: Open Access
338. Learning to Focus and Precise Cropping:A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs
- Link: Open Access
- arXiv: 2603.27494
339. Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification
- Link: Open Access
- arXiv: 2508.03081
340. First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models
- Link: Open Access
- arXiv: 2604.00455
341. OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios
- Link: Open Access
- arXiv: 2511.16937
342. UETrack: A Unified and Efficient Framework for Single Object Tracking
- Link: Open Access
- arXiv: 2603.01412
343. More Natural, More Real: Object-aware Gaussian Splatting for 3D Visual Decoding from Human Brain
- Link: Open Access
344. Prototype-based Causal Intervention for Multi-Label Image Classification
- Link: Open Access
345. PartGS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2506.17212
346. Phantom: Physical Object Interactions as Dynamic Triggers for NMS-Exploited Backdoors
- Link: Open Access
347. Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition
- Link: Open Access
- arXiv: 2603.03827
348. HyperGait: Unleashing the Power of Parsing for Gait Recognition in the Wild via Hypergraph
- Link: Open Access
349. YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detection
- Link: Open Access
350. Hypergraph-State Collaborative Reasoning for Multi-Object Tracking
- Link: Open Access
- arXiv: 2604.12665
351. Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection
- Link: Open Access
- arXiv: 2603.02286
352. HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling
- Link: Open Access
- arXiv: 2510.23043
353. WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Link: Open Access
- arXiv: 2512.12309
354. Spike-driven Discrete Aggregation for Event-based Object Detection
- Link: Open Access
355. Precise Object and Effect Removal with Adaptive Target-Aware Attention
- Link: Open Access
- arXiv: 2505.22636
356. TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection
- Link: Open Access
357. UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression
- Link: Open Access
- arXiv: 2509.25934
358. PoseAnything: General Pose-guided Video Generation with Part-aware Temporal Coherence
- Link: Open Access
359. Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection
- Link: Open Access
- arXiv: 2604.26409
360. Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection
- Link: Open Access
- arXiv: 2603.18541
361. Revisiting F-measure Optimization in Multi-Label Classification: A Sampling-based Approach
- Link: Open Access
362. PoseD-Flow: Versatile and Guided Flow Matching Model of Human Pose
- Link: Open Access
363. MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment
- Link: Open Access
- arXiv: 2602.19004
364. HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding
- Link: Open Access
- arXiv: 2603.02329
365. Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes
- Link: Open Access
- arXiv: 2601.02356
366. Learning Scene Coordinate Reconstruction from Unposed Images via Pose Graph Optimization
- Link: Open Access
367. Memory-Augmented Scene Understanding and Exploration for Open-World Aerial Object-Goal Navigation
- Link: Open Access
368. Text-guided Feature Disentanglement for Cross-modal Gait Recognition
- Link: Open Access
369. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Prompt
- Link: Open Access
370. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation
- Link: Open Access
- arXiv: 2604.04490
371. Generative Video Motion Editing with 3D Point Tracks
- Link: Open Access
- arXiv: 2512.02015
372. Event6D: Event-based Novel Object 6D Pose Tracking
- Link: Open Access
- arXiv: 2603.28045
373. MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding
- Link: Open Access
374. Hearing the Room Through the Shape of the Drum: Modal-Guided Sound Recovery from Multi-Point Surface Vibrations
- Link: Open Access
- arXiv: 2604.26678
375. Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation
- Link: Open Access
- arXiv: 2512.01677
376. Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classification
- Link: Open Access
377. Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods
- Link: Open Access
- arXiv: 2603.25767
378. Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization
- Link: Open Access
- arXiv: 2603.14750
379. From Pixel to Precision: Enhancing Handwritten Mathematical Expression Recognition with Image-Level Reward
- Link: Open Access
380. GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking
- Link: Open Access
- arXiv: 2407.01007
381. BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognition
- Link: Open Access
- arXiv: 2604.12221
382. PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation
- Link: Open Access
- arXiv: 2503.14295
383. R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection
- Link: Open Access
- arXiv: 2603.11566
384. Enhancing Out-of-Distribution Detection with Extended Logit Normalization
- Link: Open Access
- arXiv: 2504.11434
385. Adaptive Confidence Regularization for Multimodal Failure Detection
- Link: Open Access
- arXiv: 2603.02200
386. Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routing
- Link: Open Access
387. ProgTrack: A Multi-Object Tracking Algorithm with Progressive Matching Strategy
- Link: Open Access
388. HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face Avatars
- Link: Open Access
- arXiv: 2507.02803
389. FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection
- Link: Open Access
- arXiv: 2601.03928
390. Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video
- Link: Open Access
- arXiv: 2601.05251
391. EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstruction
- Link: Open Access
392. V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence
- Link: Open Access
- arXiv: 2511.20886
393. CHAL: Causal-guided Hierarchical Anomaly-aware Learning for Moving Infrared Small Target Detection
- Link: Open Access
394. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation
- Link: Open Access
395. Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression
- Link: Open Access
- arXiv: 2603.27383
396. Zoo3D: Zero-Shot 3D Object Detection at Scene Level
- Link: Open Access
- arXiv: 2511.20253
397. Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs
- Link: Open Access
- arXiv: 2512.18897
398. What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution
- Link: Open Access
399. H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image Classifiers
- Link: Open Access
- arXiv: 2604.22045
400. Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods
- Link: Open Access
- arXiv: 2603.15026
401. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation
- Link: Open Access
- arXiv: 2605.01700
402. A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps
- Link: Open Access
- arXiv: 2603.28182
403. Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object Detection
- Link: Open Access
404. MVP: Multiple View Prediction Improves GUI Grounding
- Link: Open Access
- arXiv: 2512.08529
405. Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding
- Link: Open Access
- arXiv: 2603.26211
406. Cross-Modal Guided Visual Synthesis for Data-Efficient Multimodal Depression Recognition
- Link: Open Access
407. Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors
- Link: Open Access
- arXiv: 2507.12336
408. Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning
- Link: Open Access
- arXiv: 2508.01603
409. Universal-to-Specific: Dynamic Knowledge-Guided Multiple Instance Learning for Few-Shot Whole Slide Image Classification
- Link: Open Access
410. VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding
- Link: Open Access
- arXiv: 2507.13353
411. MEMO: Human-like Crisp Edge Detection Using Masked Edge Prediction
- Link: Open Access
- arXiv: 2603.20782
412. PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image Segmentation
- Link: Open Access
413. High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning
- Link: Open Access
- arXiv: 2503.22179
414. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning
- Link: Open Access
- arXiv: 2603.00550
415. Parameterized Prompt for Incremental Object Detection
- Link: Open Access
- arXiv: 2510.27316
416. VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection
- Link: Open Access
417. Compositional Transformation Reasoning for Composed Video Retrieval
- Link: Open Access
418. RAID: Retrieval-Augmented Anomaly Detection
- Link: Open Access
- arXiv: 2602.19611
419. Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning
- Link: Open Access
- arXiv: 2603.22758
420. Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classification
- Link: Open Access
421. LAM: Language Articulated Object Modelers
- Link: Open Access
422. TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio Classification
- Link: Open Access
423. IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding
- Link: Open Access
- arXiv: 2508.09456
424. Homaloidal parametrization for detecting critical two-view configurations
- Link: Open Access
425. Detect Anything via Next Point Prediction
- Link: Open Access
- arXiv: 2510.12798
426. Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs
- Link: Open Access
- arXiv: 2511.20158
427. From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection
- Link: Open Access
- arXiv: 2602.20630
428. Opti-NeuS: Neural Reconstruction for Dual-Layered Transparent and Opaque Objects
- Link: Open Access
429. EventGait: Towards Robust Gait Recognition with Event Streams
- Link: Open Access
- arXiv: 2605.22139
430. AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects
- Link: Open Access
- arXiv: 2605.12845
431. The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification
- Link: Open Access
- arXiv: 2511.15622
432. CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation
- Link: Open Access
- arXiv: 2603.09418
433. Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification
- Link: Open Access
- arXiv: 2510.24078
434. Hermite Radial Basis Function for Surface Reconstruction via Differentiable Rendering
- Link: Open Access
435. Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images
- Link: Open Access
- arXiv: 2604.10573
436. PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image Detection
- Link: Open Access
437. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection
- Link: Open Access
438. Scene Reconstruction as Mapping Priors for 3D Detection
- Link: Open Access
439. AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment
- Link: Open Access
440. BOP-ASK: Object-Interaction Reasoning for Vision-Language Models
- Link: Open Access
- arXiv: 2511.16857
441. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection
- Link: Open Access
442. 3D-Object Perception Transformer (3PT)
- Link: Open Access
443. InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions
- Link: Open Access
- arXiv: 2602.06035
444. Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection
- Link: Open Access
445. OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery
- Link: Open Access
- arXiv: 2603.27645
446. A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimation
- Link: Open Access
447. DARC: Dual Adjustment Reasoning with Counterfactuals for Trustworthy Chest X-ray Classification
- Link: Open Access
448. Explaining Object Detectors via Collective Contribution of Pixels
- Link: Open Access
- arXiv: 2412.00666
449. AntiStyler: Defending Object Detection Models Against Adversarial Patch Attacks Using Style Removal
- Link: Open Access
450. ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.20358
451. VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network
- Link: Open Access
- arXiv: 2605.07552
452. COPYLENS: Towards Copyrighted Characters Infringement Detection via Copyright-Aware Prompt Learning
- Link: Open Access
453. AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos
- Link: Open Access
- arXiv: 2603.07758
454. Partial Weakly-Supervised Oriented Object Detection
- Link: Open Access
455. FMPose3D: monocular 3D pose estimation via flow matching
- Link: Open Access
- arXiv: 2602.05755
456. A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World
- Link: Open Access
- arXiv: 2512.04837
457. Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation
- Link: Open Access
- arXiv: 2501.05264
458. Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments
- Link: Open Access
- arXiv: 2604.07997
459. SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video Grounding
- Link: Open Access
460. EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR
- Link: Open Access
- arXiv: 2603.04090
461. EV-CGNet: Co-visible Focused 3D-guided 2D Event Keypoint Detection Network
- Link: Open Access
462. Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection
- Link: Open Access
- arXiv: 2604.03972
463. UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection
- Link: Open Access
- arXiv: 2604.21904
464. SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors
- Link: Open Access
- arXiv: 2505.22499
465. Animator-Centric Skeleton Generation on Objects with Fine-Grained Details
- Link: Open Access
- arXiv: 2604.20539
466. 4DSurf: High-Fidelity Dynamic Scene Surface Reconstruction
- Link: Open Access
- arXiv: 2603.28064
467. BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modeling
- Link: Open Access
- arXiv: 2603.01163
468. DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera Conditions
- Link: Open Access
469. UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking
- Link: Open Access
- arXiv: 2602.23734
470. Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modeling
- Link: Open Access
471. Generative Point Tracking and Forecasting
- Link: Open Access
472. LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World
- Link: Open Access
- arXiv: 2605.05390
473. Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition
- Link: Open Access
474. Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuition
- Link: Open Access
- arXiv: 2603.21629
475. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors
- Link: Open Access
- arXiv: 2604.14563
476. CD-Buffer: Complementary Dual-Buffer Framework for Test-Time Adaptation in Adverse Weather Object Detection
- Link: Open Access
- arXiv: 2603.26092
477. SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models
- Link: Open Access
- arXiv: 2506.13723
478. SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition
- Link: Open Access
- arXiv: 2603.29692
479. GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection
- Link: Open Access
- arXiv: 2603.06048
480. Learning to Track Instance from Single Nature Language Description
- Link: Open Access
- arXiv: 2605.07064
481. ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation
- Link: Open Access
482. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection
- Link: Open Access
- arXiv: 2511.18385
483. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection
- Link: Open Access
- arXiv: 2512.23273
484. ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object Recognition
- Link: Open Access
485. MOGeo: Beyond One-to-One Cross-View Object Geo-localization
- Link: Open Access
- arXiv: 2603.13843
486. AdaDexTrack: Dynamic Modulation for Adaptive and Generalizable Dexterous Manipulation Tracking
- Link: Open Access
487. SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model
- Link: Open Access
- arXiv: 2512.10957
488. ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splatting
- Link: Open Access
- arXiv: 2605.04730
489. Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence
- Link: Open Access
- arXiv: 2605.01450
490. Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition
- Link: Open Access
491. Differentially Private 2D Human Pose Estimation
- Link: Open Access
- arXiv: 2504.10190
492. OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring
- Link: Open Access
493. SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling
- Link: Open Access
494. Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition
- Link: Open Access
495. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification
- Link: Open Access
496. Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction
- Link: Open Access
497. SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling
- Link: Open Access
498. 3D Gaussian Splatting with Self-Constrained Priors for High Fidelity Surface Reconstruction
- Link: Open Access
- arXiv: 2603.19682
499. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal
- Link: Open Access
- arXiv: 2603.28224
500. Detecting Unknown Objects via Energy-based Separation for Open World Object Detection
- Link: Open Access
- arXiv: 2603.29954
501. PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing
- Link: Open Access
- arXiv: 2603.19731
502. MoVie: Broaden Your Views with Human Motion for Action Detection
- Link: Open Access
503. D^3FER: Dual Channel and Dual Branch Network for Robust Facial Expression Recognition under Dual Challenges
- Link: Open Access
504. Choreographing a World of Dynamic Objects
- Link: Open Access
- arXiv: 2601.04194
505. Anomaly-Related Residual Fields for Cross-domain Anomaly Detection
- Link: Open Access
506. Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity Learning
- Link: Open Access
507. ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding
- Link: Open Access
508. SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2511.06702
509. Detecting Compressed AI-Generated Images via Phase Spectrum Robustness
- Link: Open Access
510. Exposing and Evaluating Hallucinations for GUI Grounding
- Link: Open Access
511. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video
- Link: Open Access
- arXiv: 2511.18814
512. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects
- Link: Open Access
513. MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection
- Link: Open Access
514. Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing
- Link: Open Access
- arXiv: 2506.01783
515. Pose-guided Enriched Feature Learning for Federated-by-camera Person Re-identification
- Link: Open Access
516. Goldilocks Test Sets for Face Verification
- Link: Open Access
- arXiv: 2405.15965
517. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection
- Link: Open Access
- arXiv: 2511.17181
518. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection
- Link: Open Access
- arXiv: 2604.00549
519. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images
- Link: Open Access
520. CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs
- Link: Open Access
- arXiv: 2511.14072
521. MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures
- Link: Open Access
- arXiv: 2603.28550
522. MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection
- Link: Open Access
- arXiv: 2603.03101
523. Toward Low-Cost yet Effective Temporal Learning for UAV Tracking
- Link: Open Access
524. MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models
- Link: Open Access
- arXiv: 2604.10971
525. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors
- Link: Open Access
526. Learning Latent Concepts for Detecting Out-of-Distribution Objects
- Link: Open Access
527. PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories
- Link: Open Access
528. Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognition
- Link: Open Access
529. YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal
- Link: Open Access
- arXiv: 2604.27322
530. CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection
- Link: Open Access
- arXiv: 2603.05042
531. VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction
- Link: Open Access
- arXiv: 2512.11099
532. ExPose: Reinforcing Video Generation Models for Extreme Pose Estimation
- Link: Open Access
533. DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
- Link: Open Access
- arXiv: 2604.12812
534. DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance
- Link: Open Access
- arXiv: 2512.14266
535. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification
- Link: Open Access
- arXiv: 2512.12887
536. InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation
- Link: Open Access
- arXiv: 2603.05898
537. Progressive Multi-cue Alignment for Unaligned RGBT Tracking
- Link: Open Access
538. MPL: Match-guided Prototype Learning for Few-shot Action Recognition
- Link: Open Access
539. Prompt-Free Unknown Label Generation for Open World Detection in Remote Sensing
- Link: Open Access
540. Cross-modal Representation Learning for Diffusion-generated Image Detection
- Link: Open Access
541. Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation
- Link: Open Access
- arXiv: 2603.05929
542. HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps
- Link: Open Access
- arXiv: 2601.02730
543. OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection
- Link: Open Access
- arXiv: 2511.21064
544. Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learning
- Link: Open Access
- arXiv: 2603.26179
545. NeuROK: Generative 4D Neural Object Kinematics
- Link: Open Access
546. Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification
- Link: Open Access
- arXiv: 2602.18842
547. Tracking through Severe Occlusion via Event-Derived Transient Cues
- Link: Open Access
548. Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation
- Link: Open Access
- arXiv: 2604.20336
549. CLEP: Contrastive Language-Pose Pretraining
- Link: Open Access
550. AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Models
- Link: Open Access
551. Chain-of-Thought Guided Multi-Modal Object Re-Identification
- Link: Open Access
552. PP-Brep: Few-Shot B-rep Classification with Hybrid Graph Representation
- Link: Open Access
553. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection
- Link: Open Access
554. Tracking by Predicting 3-D Gaussians Over Time
- Link: Open Access
- arXiv: 2512.22489
555. TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free Reconstruction
- Link: Open Access
- arXiv: 2603.00697
556. PAMotion: Physics-Aware Motion Generation for Full-Body Interaction with Multiple Objects
- Link: Open Access
557. RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework
- Link: Open Access
- arXiv: 2504.10018
558. G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.14710
559. VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation
- Link: Open Access
- arXiv: 2603.12918
560. Trust-calibrated Collaborative Learning for Long-Tailed Visual Recognition
- Link: Open Access
561. Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong Again
- Link: Open Access
- arXiv: 2503.07516
562. PriVi: Towards a General-Purpose Video Model for Primate Behavior in the Wild
- Link: Open Access
- arXiv: 2511.09675
563. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection
- Link: Open Access
564. Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.05393
565. Global-Aware Edge Prioritization for Pose Graph Initialization
- Link: Open Access
- arXiv: 2602.21963
566. Geometry-driven OOD Detectors Are Class-Incremental Learners
- Link: Open Access
567. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection
- Link: Open Access
- arXiv: 2603.25159
568. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs
- Link: Open Access
- arXiv: 2511.21251
569. 3D Gaussian Splatting from Unposed Spike Stream
- Link: Open Access
570. Post-training Feature Pruning for Fundus Images Classification
- Link: Open Access
571. Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos
- Link: Open Access
- arXiv: 2604.17749
572. Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generation
- Link: Open Access
- arXiv: 2505.19459
573. Progressive Cross-Modal Causal Intervention for Long-Term Action Recognition
- Link: Open Access
574. OpenFS: Multi-Hand-Capable Fingerspelling Recognition with Implicit Signing-Hand Detection and Frame-Wise Letter-Conditioned Synthesis
- Link: Open Access
- arXiv: 2602.22949
575. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection
- Link: Open Access
576. EXOTIC: External Vision-driven Incomplete Multi-view Classification
- Link: Open Access
577. A Difference-in-Difference Approach to Detecting AI-Generated Images
- Link: Open Access
- arXiv: 2602.23732
578. Ego-Grounding for Personalized Question-Answering in Egocentric Videos
- Link: Open Access
- arXiv: 2604.01966
579. SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent Recognition
- Link: Open Access
580. LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly Detection
- Link: Open Access
581. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection
- Link: Open Access
- arXiv: 2604.02071
582. CoWTracker: Tracking by Warping instead of Correlation
- Link: Open Access
- arXiv: 2602.04877
583. Transition Models: Rethinking the Generative Learning Objective
- Link: Open Access
- arXiv: 2509.04394
584. Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial Transcriptomics
- Link: Open Access
585. CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective Video
- Link: Open Access
586. ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection
- Link: Open Access
- arXiv: 2603.19776
587. ViHOI: Human-Object Interaction Synthesis with Visual Priors
- Link: Open Access
- arXiv: 2603.24383
588. Confusion-Aware Spectral Regularizer for Long-Tailed Recognition
- Link: Open Access
- arXiv: 2603.16732
589. From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing
- Link: Open Access
- arXiv: 2603.01038
590. EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions
- Link: Open Access
- arXiv: 2603.25135
591. Specificity-aware reinforcement learning for fine-grained open-world classification
- Link: Open Access
- arXiv: 2603.03197
592. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning
- Link: Open Access
- arXiv: 2602.19206
593. Rethinking BCE Loss for Multi-Label Image Recognition with Fine-Tuning
- Link: Open Access
594. SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose Estimation
- Link: Open Access
595. POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling
- Link: Open Access
- arXiv: 2512.13192
596. Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs
- Link: Open Access
- arXiv: 2603.02618
597. TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Tracking
- Link: Open Access
- arXiv: 2512.01329
598. Machine Unlearning via Adaptive Gradient Reweighting and Multi-stage Objective Optimization
- Link: Open Access
599. LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight
- Link: Open Access
- arXiv: 2511.20648
600. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection
- Link: Open Access
- arXiv: 2605.10130
601. No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection
- Link: Open Access
- arXiv: 2602.19248
602. Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields
- Link: Open Access
603. GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection
- Link: Open Access
604. MVLM: Template-Free Tracking via Vision-Language Margin Confidence and Memory-Gated Tracking
- Link: Open Access
605. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation
- Link: Open Access
- arXiv: 2602.19944
606. Decoupling Bias, Aligning Distributions: Synergistic Fairness Optimization for Deepfake Detection
- Link: Open Access
- arXiv: 2511.10150
607. Anti-I2V: Safeguarding your Photos from Malicious Image-to-video Generation
- Link: Open Access
- arXiv: 2603.24570
608. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation
- Link: Open Access
609. Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach
- Link: Open Access
- arXiv: 2508.15778
610. Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding
- Link: Open Access
611. Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimation
- Link: Open Access
612. BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation
- Link: Open Access
- arXiv: 2603.05921
613. PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding
- Link: Open Access
614. Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images
- Link: Open Access
- arXiv: 2508.03643
615. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection
- Link: Open Access
- arXiv: 2605.01638
616. Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning
- Link: Open Access
- arXiv: 2605.06092
617. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
- Link: Open Access
- arXiv: 2506.21076
618. Finding Distributed Object-Centric Properties in Self-Supervised Transformers
- Link: Open Access
- arXiv: 2603.26127
619. KLIP: Localized Distribution Shift Detection via KL-Divergence with Diffusion Priors in Inverse Problems
- Link: Open Access
620. Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes
- Link: Open Access
- arXiv: 2512.12982
621. Object-WIPER: Training-Free Object and Associated Effect Removal in Videos
- Link: Open Access
- arXiv: 2601.06391
622. The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition
- Link: Open Access
- arXiv: 2505.24840
623. SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker
- Link: Open Access
- arXiv: 2604.12502
624. Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding
- Link: Open Access
- arXiv: 2601.10611
625. MonoVLM: Monocular 3D Visual Grounding with Vision Language Models
- Link: Open Access
626. Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection
- Link: Open Access
627. From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection
- Link: Open Access
628. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2505.12702
629. MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid Cameras
- Link: Open Access
630. Detect Any AI-Counterfeited Text Image
- Link: Open Access
631. Diversity over Uniformity: Rethinking Representation in Generated Image Detection
- Link: Open Access
632. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding
- Link: Open Access
- arXiv: 2605.18018
633. GOR-IS: 3D Gaussian Object Removal In the Intrinsic Space
- Link: Open Access
- arXiv: 2605.00498
634. IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion
- Link: Open Access
635. Graph Attention Prototypical Network for Robust Few-Shot Classification
- Link: Open Access
636. MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision
- Link: Open Access
- arXiv: 2602.20689
637. Robust Promptable Video Object Segmentation
- Link: Open Access
- arXiv: 2605.12006
638. Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision
- Link: Open Access
- arXiv: 2603.27179
639. UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection
- Link: Open Access
- arXiv: 2603.17492
640. RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation
- Link: Open Access
- arXiv: 2603.14880
641. LA-Pose: Latent Action Pretraining Meets Pose Estimation
- Link: Open Access
- arXiv: 2604.27448
642. TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events
- Link: Open Access
643. OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control
- Link: Open Access
- arXiv: 2511.02483
644. Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detection
- Link: Open Access
645. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection
- Link: Open Access
646. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection
- Link: Open Access
647. SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images
- Link: Open Access
- arXiv: 2602.20412
648. FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning
- Link: Open Access
- arXiv: 2603.05506
649. IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation
- Link: Open Access
650. EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature Enhancement
- Link: Open Access
651. H^2A^2: Homogeneity-Aware and Heterogeneity-Aware Feature Perception for Unified Indoor 3D Object Detection
- Link: Open Access
652. D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation
- Link: Open Access
- arXiv: 2512.12622
653. IDperturb: Enhancing Variation in Synthetic Face Generation via Angular Perturbations
- Link: Open Access
- arXiv: 2602.18831
654. TESO: Online Tracking of Essential Matrix by Stochastic Optimization
- Link: Open Access
- arXiv: 2604.19420
655. DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge
- Link: Open Access
656. Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph Generation
- Link: Open Access
657. Exploring 6D Object Pose Estimation with Deformation
- Link: Open Access
- arXiv: 2604.06720
658. DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding
- Link: Open Access
- arXiv: 2605.15542
659. Towards Intrinsic-Aware Monocular 3D Object Detection
- Link: Open Access
- arXiv: 2603.27059
660. UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition
- Link: Open Access
- arXiv: 2404.15254
661. Fourier Angle Alignment for Oriented Object Detection in Remote Sensing
- Link: Open Access
- arXiv: 2602.23790
662. AeroGS: Scale-Aware Gaussian Splatting for Pose-Free Dynamic UAV Scene Reconstruction
- Link: Open Access
663. InsCal: Calibrated Multi-Source Fully Test-Time Prompt Tuning for Object Detection
- Link: Open Access
664. Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection
- Link: Open Access
- arXiv: 2408.13516
665. ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
- Link: Open Access
- arXiv: 2602.01639
666. MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image Classification
- Link: Open Access
- arXiv: 2602.20873
667. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection
- Link: Open Access
- arXiv: 2603.24139
668. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation
- Link: Open Access
- arXiv: 2603.10703
669. DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Prompting
- Link: Open Access
670. PhysHO: Physics-Based Dynamic 3D Gaussian Human and Object from Monocular Video
- Link: Open Access
671. EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision
- Link: Open Access
- arXiv: 2605.13152
672. Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection
- Link: Open Access
673. Grounding Everything in Tokens for Multimodal Large Language Models
- Link: Open Access
- arXiv: 2512.10554
674. Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection
- Link: Open Access
675. ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications
- Link: Open Access
- arXiv: 2510.10113
676. Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection
- Link: Open Access
- arXiv: 2512.17514
677. DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classification
- Link: Open Access
678. Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models
- Link: Open Access
- arXiv: 2603.00431
679. Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation
- Link: Open Access
- arXiv: 2512.06158
680. DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imaging
- Link: Open Access
681. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection
- Link: Open Access
- arXiv: 2605.14110
682. FastRef: Fast Prototype Refinement for Few-shot Industrial Anomaly Detection
- Link: Open Access
683. Generalizable Co-Salient Object Detection via Mixed Content-Style Modulation
- Link: Open Access
684. Region-Aware Instance Consistency Learning for Micro-Expression Recognition
- Link: Open Access
685. TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis
- Link: Open Access
- arXiv: 2506.20380
686. BDNet:Bio-Inspired Dual-Backbone Small Object Detection Network
- Link: Open Access
687. RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph
- Link: Open Access
688. RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation
- Link: Open Access
- arXiv: 2604.03454
689. FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement
- Link: Open Access
- arXiv: 2603.19608
690. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
- Link: Open Access
- arXiv: 2604.19432
691. Neu-PiG: Neural Preconditioned Grids for Fast Dynamic Surface Reconstruction on Long Sequences
- Link: Open Access
692. Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking
- Link: Open Access
- arXiv: 2603.06034
693. AnthroTAP: Learning Point Tracking with Real-World Motion
- Link: Open Access
- arXiv: 2507.06233
694. Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations
- Link: Open Access
- arXiv: 2604.04863
695. See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditions
- Link: Open Access
696. WildPose: A Unified Framework for Robust Pose Estimation in the Wild
- Link: Open Access
- arXiv: 2605.12774
697. Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding
- Link: Open Access
- arXiv: 2603.03762
698. Seeing Motion Through Polarity for Event-based Action Recognition
- Link: Open Access
699. BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object Extraction
- Link: Open Access
700. FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition
- Link: Open Access
- arXiv: 2603.26908
701. ZINA: Multimodal Fine-grained Hallucination Detection and Editing
- Link: Open Access
- arXiv: 2506.13130
702. Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal Transport
- Link: Open Access
703. TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment
- Link: Open Access
- arXiv: 2603.22819
704. TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition
- Link: Open Access
- arXiv: 2512.01248
705. DialogueVPR: Towards Conversational Visual Place Recognition
- Link: Open Access
706. RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation
- Link: Open Access
- arXiv: 2603.11106
707. Why Not Hyperparameter-Friendly Optimisation? A Monotonic Adaptive Norm Rescaling Approach For Long-Tailed Recognition
- Link: Open Access
708. Free-Grained Hierarchical Visual Recognition
- Link: Open Access
- arXiv: 2510.14737
709. KV-Tracker: Real-Time Pose Tracking with Transformers
- Link: Open Access
- arXiv: 2512.22581
710. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation
- Link: Open Access
711. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos
- Link: Open Access
- arXiv: 2602.22091
712. SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splatting
- Link: Open Access
- arXiv: 2604.19202
713. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection
- Link: Open Access
714. Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building Labels
- Link: Open Access
715. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training
- Link: Open Access
716. Changes in Real Time: Online Scene Change Detection with Multi-View Fusion
- Link: Open Access
- arXiv: 2511.12370
717. HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos
- Link: Open Access
- arXiv: 2603.06732
718. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision
- Link: Open Access
- arXiv: 2602.13195
719. Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented Adaptation
- Link: Open Access
- arXiv: 2604.01974
720. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement
- Link: Open Access
- arXiv: 2602.23120
721. DIMOS: Disentangling Instance-level Moving Object Segmentation
- Link: Open Access
722. BiGain: Unified Token Compression for Joint Generation and Classification
- Link: Open Access
- arXiv: 2603.12240
723. Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-View
- Link: Open Access
724. Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence
- Link: Open Access
- arXiv: 2510.20470
725. ORD: Object-Relation Decoupling for Generalized 3D Visual Grounding
- Link: Open Access
726. SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networks
- Link: Open Access
- arXiv: 2503.08703
727. SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning
- Link: Open Access
- arXiv: 2604.27596
728. Mechanisms of Object Localization in Vision-Language Models
- Link: Open Access
729. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation
- Link: Open Access
- arXiv: 2602.03595
730. FSLoRA: Harmonizing Detection and Re-Identification via Freq-Spatial Low-Rank Adapter for One-Stage Person Search
- Link: Open Access
731. Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision
- Link: Open Access
- arXiv: 2511.14197
732. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection
- Link: Open Access
733. Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection
- Link: Open Access
- arXiv: 2603.10598
734. RankOOD - Class Ranking-based Out-of-Distribution Detection
- Link: Open Access
- arXiv: 2511.19996
735. BAMI: Training-Free Bias Mitigation in GUI Grounding
- Link: Open Access
- arXiv: 2605.06664
736. Scene Grounding in the Wild
- Link: Open Access
737. Zero-shot Detection of AI-Generated Image via RAW-RGB Alignment
- Link: Open Access
738. Hunting Normality from Query Sample via Residual Learning for Generalist Anomaly Detection
- Link: Open Access
739. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size
- Link: Open Access
- arXiv: 2603.07988
740. Making the Classification Explanation Faithful to the Confidence Score
- Link: Open Access
741. Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World
- Link: Open Access
- arXiv: 2603.12746
742. C^2FG: Control Classifier-Free Guidance via Score Discrepancy Analysis
- Link: Open Access
743. Semantic Alignment for Pose-Invariant Identity Preserving Diffusion
- Link: Open Access
744. UniChange: Unifying Change Detection with Multimodal Large Language Model
- Link: Open Access
- arXiv: 2511.02607
745. Adapting In-context Generation for Enhanced Composed Image Retrieval
- Link: Open Access
746. Unleashing Vision-Language Semantics for Deepfake Video Detection
- Link: Open Access
- arXiv: 2603.24454
747. Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Tracking
- Link: Open Access
748. EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence
- Link: Open Access
- arXiv: 2512.15160
749. ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild
- Link: Open Access
- arXiv: 2604.22202
750. Target-Aware Invertible Encoder with Reconstruction Guidance for Infrared Small Target Detection
- Link: Open Access
751. The Invisible Gorilla Effect in Out-of-distribution Detection
- Link: Open Access
- arXiv: 2602.20068
752. Instance-level Visual Active Tracking with Occlusion-Aware Planning
- Link: Open Access
- arXiv: 2604.21453
753. Pointing at Parts: Training-Free Few-Shot Grounding in Multimodal LLMs
- Link: Open Access
754. ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision Transformer
- Link: Open Access
755. Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection
- Link: Open Access
756. Good Can Sometimes be Bad: A Unified Attack against 3D Point Cloud Classifier by a Flexible Isotropic Resampling
- Link: Open Access
757. Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality
- Link: Open Access
- arXiv: 2512.07951
758. Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval
- Link: Open Access