R
Published on

CVPR 2026 — Detection & Recognition

Detection & Recognition

758 papers

1. Ultra Diffusion Poser: Diffusion-Based Human Motion Tracking from Sparse Inertial Sensors and Ranging-based Between-sensor Distances

2. An Efficient Token Compression Framework for Visual Object Tracking

3. Skyra: AI-Generated Video Detection via Grounded Artifact Reasoning

4. SRGCD: Stability-Driven Region Growth Framework for 3D Change Detection

5. GuardTrace-VL: Detecting Unsafe Multimodel Reasoning via Iterative Safety Supervision

6. MultiAnimate: Pose-Guided Image Animation Made Extensible

7. Does YOLO Really Need to See Every Training Image in Every Epoch?

8. MonoSAOD: Monocular 3D Object Detection with Sparsely Annotated Label

9. CrossHOI: Learning Cross-View Representations for Monocular 3D Human-Object Interaction Reconstruction

10. ANTS: Adaptive Negative Textual Space Shaping for OOD Detection via Test-Time MLLM Understanding and Reasoning

11. Rethinking Occlusion Modeling for UAV Tracking

12. ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos

13. Make it SING: Analyzing Semantic Invariants in Classifiers

14. PRISM: Prototype-based Reasoning with Inter-modal Semantic Mining for Interpretable Image Recognition

15. JRM: Joint Reconstruction Model for Multiple Objects without Alignment

16. EG-3DVG: Expression and Geometry Aware Grounding Decoder for 3D Visual Grounding

17. Synergistic Bleeding Region and Point Detection in Laparoscopic Surgical Videos

18. Refracting Reality: Generating Images with Realistic Transparent Objects

19. Tri-Modal Fusion Transformers for UAV-based Object Detection

20. RetFormer: Multimodal Retrieval for Enhancing Image Recognition

21. Globally Optimal Pose from Orthographic Silhouettes

22. Mind the Gap: Transferring Labels to Align Object Detection Datasets

23. Decoupled Generative Modeling for Human-Object Interaction Synthesis

24. Neural Field-Based 3D Surface Reconstruction of Microstructures from Multi-Detector Signals in Scanning Electron Microscopy

25. Paparazzo: Active Mapping of Moving 3D Objects

26. QueryMe: Query-Driven Open-Vocabulary 3D Object Affordances Grounding from Multimodal Evidence

27. DREAM: Document Recognition with Explicit Adaptive Memory

28. RHINO: Reconstructing Human Interactions with Novel Objects from Monocular Videos

29. Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video

30. Harnessing the Power of Foundation Models for Accurate Material Classification

31. HamiPose: Hamiltonian Optimization for Unsupervised Domain Adaptive Pose Estimation

32. AKCMamba-YOLO: Selective State Space Models For Real-Time Object Detection

33. Refacade: Editing Object with Given Reference Texture

34. COG: Confidence-aware Optimal Geometric Correspondence for Unsupervised Single-reference Novel Object Pose Estimation

35. WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition

36. Breaking Smooth-Motion Assumptions: A UAV Benchmark for Multi-Object Tracking in Complex and Adverse Conditions

37. Breaking Spurious Correlations: Uncertainty-Driven Causal Transformers for AU Detection

38. VGGDrive: Empowering Vision-Language Models with Cross-View Geometric Grounding for Autonomous Driving

39. FBTA: Enabling Single-GPU End-to-End Gigapixel WSI Classification with Feature Bridging and Translation Alignment

40. Agile Deliberation: Concept Deliberation for Subjective Visual Classification

41. UAST: Unified Active Search and Tracking for Arbitrary Targets with UAVs

42. DiG: Differential Grounding for Enhancing Fine-Grained Perception in Multimodal Large Language Models

43. From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level Grounding

44. ViTPrompt: Training-Free Prompt Refinement with Visual Tokens for Open-Vocabulary Detection

45. Beyond Heuristic Prompting: A Concept-Guided Bayesian Framework for Zero-Shot Image Recognition

46. SDDF: Specificity-Driven Dynamic Focusing for Open-Vocabulary Camouflaged Object Detection

47. InterRVOS: Interaction-Aware Referring Video Object Segmentation

48. ICTPolarReal: A Polarized Reflection and Material Dataset of Real World Objects

49. Pano3DComposer: Feed-Forward Compositional 3D Scene Generation from Single Panoramic Image

50. MGDHand: Multi-Granularity Prior-to-Inertial Distillation Framework for Sequential 3D Hand Pose Estimation from Sparse IMUs

51. Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video

52. T2SGrid: Temporal-to-Spatial Gridification for Video Temporal Grounding

53. All in One: Unifying Deepfake Detection, Tampering Localization, and Source Tracing with a Robust Landmark-Identity Watermark

54. Learning Cross-View Object Correspondence via Cycle-Consistent Mask Prediction

55. DyFCLT: Dynamic Frequency-Decoupled Cross-Modal Learning Transformer for Multimodal Tiny Object Detection

56. FedSDR: Federated Graph Learning with Structural Noise Detection and Reconstruction

57. Glove2Hand: Synthesizing Natural Hand-Object Interaction from Multi-Modal Sensing Gloves

58. DICArt: Advancing Category-level Articulated Object Pose Estimation in Discrete State-Spaces

59. Reinforcing Video Object Segmentation to Think before it Segments

60. Face2Scene: Using Facial Degradation as an Oracle for Diffusion-Based Scene Restoration

61. FHAvatar: Fast and High-Fidelity Reconstruction of Face-and-Hair Composable 3D Head Avatar from Few Casual Captures

62. TVHighlights: LLM-Guided Human-Free Collaborative Training for Video Highlight Detection in Movies and TV Dramas

63. Physically Ground Commonsense Knowledge for Articulated Object Manipulation with Analytic Concepts

64. UZ3DVG: Unaided Zero-Shot 3D Visual Grounding with Generated Language Conditions

65. Your One-Stop Solution for AI-Generated Video Detection

66. Common Inpainted Objects In-N-Out of Context

67. One Model, Many Budgets: Elastic Latent Interfaces for Diffusion Transformers

68. Hierarchical Concept Embedding & Pursuit for Interpretable Image Classification

69. Visual Prototype Conditioned Focal Region Generation for UAV-Based Object Detection

70. X-AVDT: Audio-Visual Cross-Attention for Robust Deepfake Detection

71. CLEX: Complementary Label Exchange Learning for Noisy Facial Expression Recognition

72. OneHOI: Unifying Human-Object Interaction Generation and Editing

73. FACE: A Face-based Autoregressive Representation for High-Fidelity and Efficient Mesh Generation

74. GenMatter: Perceiving Physical Objects with Generative Matter Models

75. MedMO: Grounding and Understanding Multimodal Large Language Model for Medical Images

76. UNI-OOD: Unified Object- and Image-level Out-of-Distribution Detection via Cross-Context Attentive Vision-Language Modeling

77. Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval

78. CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling

79. Back to Point: Exploring Point-Language Models for Zero-Shot 3D Anomaly Detection

80. GroundingME: Exposing the Visual Grounding Gap in MLLMs through Multi-Dimensional Evaluation

81. VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition

82. Synthetic Object Compositions for Scalable and Accurate Learning in Detection, Segmentation, and Grounding

83. SMV-EAR: Bring Spatiotemporal Multi-View Representation Learning into Efficient Event-Based Action Recognition

84. FVBench: Benchmarking Deepfake Video Detection Capability of Large Multimodal Models

85. Exploring Adaptive Masked Reconstruction for Self-Supervised Skeleton-Based Action Recognition

86. One-to-All Animation: Alignment-Free Character Animation and Image Pose Transfer

87. Out of Sight, Out of Track: Adversarial Attacks on Propagation-based Multi-Object Trackers via Query State Manipulation

88. Multi-view Crowd Tracking Transformer with View-Ground Interactions Under Large Real-World Scenes

89. UST-Hand: An Uncertainty-aware Spatiotemporal Point Cloud Interaction Network for 3D Self-supervised Hand Pose Estimation

90. CoCoVideo: The High-Quality Commercial-Model-Based Contrastive Benchmark for AI-Generated Video Detection

91. Learnability-Driven Submodular Optimization for Active Roadside 3D Detection

92. TRCoRSurg: Temporal-Relational Co-Reasoning for Surgical Video Triplet Recognition

93. Clothe and Pose

94. RHCNet: Residual-Guided Hierarchical Calibration Network for Robust Underwater Object Detection

95. TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs

96. HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition

97. Pose-Free Omnidirectional Gaussian Splatting for 360-Degree Videos with Consistent Depth Priors

98. ObjectMorpher: 3D-Aware Image Editing via Deformable 3DGS

99. C-LaV: Conditional Latent Velocity Field Denoising for Weather-Robust LiDAR Place Recognition

100. Enhancing Accuracy of Uncertainty Estimation in Appearance-based Gaze Tracking with Probabilistic Evaluation and Calibration

101. UDAPose: Unsupervised Domain Adaptation for Low-Light Human Pose Estimation

102. Back to Source: Open-Set Continual Test-Time Adaptation via Domain Compensation

103. ManifoldNeuS: Manifold-aware View Optimizability for Pose-Free Neural Surface Reconstruction

104. Unlearning without Forgetting: Securely Removing Targeted Concepts from Large-Scale Vision-Language Open-Vocabulary Detectors

105. MatMart: Material Reconstruction of 3D Objects via Diffusion

106. Phrase-Grounding-Aware Supervised Fine-Tuning for Chart Recognition via Side-Masked Attention

107. E2EGS: Event-to-Edge Gaussian Splatting for Pose-Free 3D Reconstruction

108. TrackMAE: Video Representation Learning via Track Mask and Predict

109. Translating Signals to Languages for sEMG-Based Activity Recognition

110. HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics

111. Composite-Attribute Person Re-Identification via Pose-Guided Disentanglement

112. SpikeTrack: High-performance and Energy-efficient Event-Based Object Tracking with Spiking Neural Network

113. PoseGaussian: 6D Pose Estimation for Unseen Objects via Sparse-View Object-Level 3D Gaussian Splatting

114. Neural Distribution Prior for LiDAR Out-of-Distribution Detection

115. Physical Adversarial Clothing Evades Visible-Thermal Detectors via Non-Overlapping RGB-T Pattern

116. Complementary Prototype Mapping for Efficient Multimodal Anomaly Detection

117. OSMO: Open-vocabulary Self-eMOtion Tracking

118. SymphoMotion: Joint Control of Camera Motion and Object Dynamics for Coherent Video Generation

119. Boosting Vision-Language Models Towards Cross-Domain Incremental Object Detection

120. Black-Box Domain Adaptation for Object Detection with Retention-Driven Knowledge Compression

121. SpikeTrack: A Spike-driven Framework for Efficient Visual Tracking

122. Adaptive Data Augmentation with Multi-armed Bandit: Sample-Efficient Embedding Calibration for Implicit Pattern Recognition

123. Energy-GS: Image Energy-guided Pose Alignment Gaussian Splatting with redesigned pose gradient flow

124. OMG-Bench: A New Challenging Benchmark for Skeleton-based Online Micro Hand Gesture Recognition

125. Learning to Refuse: Refusal-Aware Reinforcement Fine-Tuning for Hard-Irrelevant Queries in Video Temporal Grounding

126. Neural Gabor Splatting: Enhanced Gaussian Splatting with Neural Gabor for High-frequency Surface Reconstruction

127. Let Your Image Move with Your Motion! -- Implicit Multi-Object Multi-Motion Transfer

128. E-3DPSM: A State Machine for Event-based Egocentric 3D Human Pose Estimation

129. SAIDO: Generalizable Detection of AI-Generated Images via Scene-Aware and Importance-Guided Dynamic Optimization in Continual Learning

130. Drift-Resilient Temporal Priors for Visual Tracking

131. Recovering Physically Plausible Human-Object Interactions from Monocular Videos

132. Local Motion Matters: A Deconstruct-Recompose Paradigm for Reinforcement Learning Pre-training from Videos

133. EW-DETR: Evolving World Object Detection via Incremental Low-Rank DEtection TRansformer

134. Back to the Feature: Explaining Video Classifiers with Video Counterfactual Explanations

135. CFG-Ctrl: Control-Based Classifier-Free Diffusion Guidance

136. FlowComposer: Composable Flows for Compositional Zero-Shot Learning

137. NoOVD: Novel Category Discovery and Embedding for Open-Vocabulary Object Detection

138. TACO: Task-Aware Contrastive Learning for Joint LiDAR Localization and 3D Object Detection

139. Egocentric Visibility-Aware Human Pose Estimation

140. Visual Grounding for Object Questions

141. Agent4FaceForgery: Multi-Agent LLM Framework for Realistic Face Forgery Detection

142. D2FANet: Enhancing Video Object Detection with Dual-Domain Feature Aggregation Network

143. Towards an Incremental Unified Multimodal Anomaly Detection: Augmenting Multimodal Denoising From an Information Bottleneck Perspective

144. PhaseWin Search Framework Enable Efficient Object-Level Interpretation

145. Plug-and-Play Incomplete Multi-View Clustering via Janus-Faced Affinity Learning with Topology Harmonization

146. SHOW3D: Capturing Scenes of 3D Hands and Objects in the Wild

147. CARI4D: Category Agnostic 4D Reconstruction of Human-Object Interaction

148. Real-World Point Tracking with Verifier-Guided Pseudo-Labeling

149. Dual-Prototype-Guided Multi-task Learning for Unsupervised Anomaly Detection and Classification

150. AXG-Reasoner: Error Detection and Explanation in Long Task Videos with Vision-Language Models

151. MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction

152. Breaking the 3D Dataset Bottleneck: Fast Scalable Generation of Aligned 3D Assets from Scratch for Category 6D Pose Estimation and Robotic Grasping

153. The Road Less Seen: Segment Exploration for Weakly Supervised Video Anomaly Detection

154. Query2Uncertainty: Robust Uncertainty Quantification and Calibration for 3D Object Detection under Distribution Shift

155. PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal Consistency

156. DiffusionFF: A Diffusion-based Framework for Joint Face Forgery Detection and Fine-Grained Artifact Localization

157. Omni-AD: A Large-scale and Versatile Benchmark for Industrial Anomaly Detection

158. DeepProtect: Proactive Face-Swapping Defense using Identity Blending and Attribute Distortion

159. OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image Generation

160. Protego: User-Centric Pose-Invariant Privacy Protection Against Face Recognition-Induced Digital Footprint Exposure

161. MeshWeaver: Sparse-Voxel-Guided Surface Weaving for Autoregressive Mesh Generation

162. Adaptive Capacity Autoregressive Visual Tracking

163. EffectErase: Joint Video Object Removal and Insertion for High-Quality Effect Erasing

164. Heuristic-inspired Reasoning Priors Facilitate Data-Efficient Referring Object Detection

165. Enhancing Hands in 3D Whole-Body Pose Estimation with Conditional Hands Modulator

166. DA-Mamba: Learning Domain-Aware State Space Model for Global-Local Alignment in Domain Adaptive Object Detection

167. Beyond Duality: A Hybrid Framework of Leveraging Shared and Private Features for RGB-Event Object Detection

168. HierUQ: Hierarchical Uncertainty Quantification with Adaptive Granularity Reconciliation for Degraded Image Classification

169. Off The Grid: Detection of Primitives for Feed-Forward 3D Gaussian Splatting

170. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval

171. EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses

172. RaGS: Unleashing 3D Gaussian Splatting from 4D Radar and Monocular Cue for 3D Object Detection

173. From Few-way to Many-way: Rethinking Few-shot Fine-grained Image Classification

174. A Debiased Reconstruction-based Framework for Training-Free Detection of AI-Generated Images

175. SeeGroup: Multi-Layer Depth Estimation of Transparent Surfaces via Self-Determined Grouping

176. CrossVL: Complexity-Aware Feature Routing and Paired Curriculum for Cross-View Vision-Language Detection

177. Affostruction: 3D Affordance Grounding with Generative Reconstruction

178. RegFormer: Transferable Relational Grounding for Efficient Weakly-Supervised Human-Object Interaction Detection

179. Real-Time Multimodal Fingertip Contact Detection via Depth and Motion Fusion for Vision-Based Human-Computer Interaction

180. MMGait: Towards Multi-Modal Gait Recognition

181. PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing

182. Physical Object Understanding with a Physically Controllable World Model

183. GroundVTS: Visual Token Sampling in Multimodal Large Language Models for Video Temporal Grounding

184. DFD-HR: Generalizable Deepfake Detection via Hierarchical Routing Learning

185. Data Leakage Detection and De-duplication in Large Scale Geospatial Image Datasets

186. ConsistCompose: Unified Multimodal Layout Control for Image Composition

187. MMFace-DiT: A Dual-Stream Diffusion Transformer for High-Fidelity Multimodal Face Generation

188. KASALv2: Fully Automatic 3D Rotational Symmetry Classification and Axis Localization

189. Look Before You Fuse: 2D-Guided Cross-Modal Alignment for Robust 3D Detection

190. Similarity-Consistent Likelihood Diffusion enables Hidden Person Detection from Wall Reflections

191. Batman: Benign Knowledge Alignment Through Malicious Null Space in Federated Backdoor Attack

192. ArtPro: Self-Supervised Articulated Object Reconstruction with Adaptive Integration of Mobility Proposals

193. Enabling Supervised Learning of Generative Signatures for Generalized AI-Generated Images Detection

194. TTP: Test-Time Padding for Adversarial Detection and Robust Adaptation on Vision-Language Models

195. DETACH : Decomposed Spatio-Temporal Alignment for Exocentric Video and Ambient Sensors with Staged Learning

196. Think Before You Drive: World Model-Inspired Multimodal Grounding

197. S2^2-MLLM: Boosting Spatial Reasoning Capability of MLLMs for 3D Visual Grounding with Structural Guidance

198. UniSpector: Towards Universal Open-set Defect Recognition via Spectral-Contrastive Visual Prompting

199. MimiCAT: Mimic with Correspondence-Aware Cascade-Transformer for Category-Free 3D Pose Transfer

200. Advancing Image Classification with Discrete Diffusion Classification Modeling

201. RecoverMark: Robust Watermarking for Localization and Recovery of Manipulated Faces

202. Spectral-Geometric Neural Fields for Pose-Free LiDAR View Synthesis

203. Active Inference for Micro-Gesture Recognition: EFE-Guided Temporal Sampling and Adaptive Learning

204. DGGT: Feedforward 4D Reconstruction of Dynamic Driving Scenes using Unposed Images

205. RPGFusion: 4D Radar Prior-Guided Multi-Modal Fusion for 3D Detection

206. Rounded or Streamlined Head? Bridging Concept Bottleneck Models and Attribute-Described Object Parts

207. MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding

208. PET-DINO: Unifying Visual Cues into Grounding DINO with Prompt-Enriched Training

209. UniRefiner: Teaching Pre-trained ViTs to Self-Dispose Dross via Contrastive Register

210. Spatio-Temporal Conditional Denoising Transformer for Modality-Missing RGBT Tracking

211. Distilling Unsigned Distance Function for Surface Reconstruction from 3D Gaussian Splatting

212. TouchDream: 3D Object Completion through Imagined Touch

213. Dual-branch Distilled Transformer for Efficient Asymmetric UAV Tracking

214. TeHOR: Text-Guided 3D Human and Object Reconstruction with Textures

215. PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View Reasoning

216. VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

217. UniPR: Unified Object-level Real-to-Sim Perception and Reconstruction from a Single Stereo Pair

218. DiffDecompose: Layer-Wise Decomposition of Alpha-Composited Images via Diffusion Transformers

219. Fast Spatial Tracking with Visual Geometry Transformer

220. EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation

221. 3DReflecNet: A Large-Scale Dataset for 3D Reconstruction of Reflective, Transparent, and Low-Texture Objects

222. MetaSpectra+: A Compact Broadband Metasurface Camera for Snapshot Hyperspectral+ Imaging

223. TF-CADE: Foreground-Concentrated Text-Video Alignment for Zero-Shot Temporal Action Detection

224. Artiverse: A Diverse and Physically Grounded Dataset for Articulated Objects

225. From Detection to Association: Learning Discriminative Object Embeddings for Multi-Object Tracking

226. Representing 3D Faces with Learnable B-Spline Volumes

227. Rotation Invariant and Symmetry Aware Pixel Difference Network for Remote Sensing Object Detection

228. Gamba: Mamba-based graph convolutional network with dynamic graph topology learning for action recognition

229. RARE: Learn to RAnk and REtrieve for Monocular 3D Object Detection

230. Prospective Dynamic 3D MRI Reconstruction via Latent-Space Motion Tracking from Single Measurement

231. R-4B: Incentivizing General-Purpose Auto-Thinking in MLLMs via Bi-Mode Annealing and Reinforce Learning

232. UniRain: Unified Image Deraining with RAG-based Dataset Distillation and Multi-objective Reweighted Optimization

233. SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation

234. Frequency-domain Manipulation for Face Obfuscation

235. Generalizable Structure-Aware Keypoint Correspondence for Category-Unified 3D Single Object Tracking

236. Rethinking Pose Refinement in 3D Gaussian Splatting under Pose Prior and Geometric Uncertainty

237. Scalable Object Relation Encoding for Better 3D Spatial Reasoning in Large Language Models

238. Multimodal Protein Language Models for Enzyme Kinetic Parameters: From Substrate Recognition to Conformational Adaptation

239. ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept Vectors

240. PDD: Manifold-Prior Diverse Distillation for Medical Anomaly Detection

241. Boosting Quantitive and Spatial Awareness for Zero-Shot Object Counting

242. RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation

243. PDCR: Perception-Decomposed Confidence Reward for Vision-Language Reasoning

244. PAM: A Pose-Appearance-Motion Engine for Sim-to-Real HOI Video Generation

245. Condensed Test-Time Adaptation of VLMs for Action Recognition

246. Balanced Hierarchical Contrastive Learning with Decoupled Queries for Fine-grained Object Detection in Remote Sensing Images

247. Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images

248. TGTrack: Temporal Generative Learning for Unified Single Object Tracking

249. OrienPose: Orientation-Guided Novel View Synthesis for Single-Image Unseen Object Pose Estimation

250. VULCAN: Tool-Augmented Multi Agents for Iterative 3D Object Arrangement

251. ArtHOI: Taming Foundation Models for Monocular 4D Reconstruction of Hand-Articulated-Object Interactions

252. MV-TAP: Tracking Any Point in Multi-View Videos

253. ADSeeker: A Knowledge-Grounded Reasoning Framework for Industry Anomaly Detection and Reasoning

254. Humanoid Generative Pre-Training for Zero-Shot Motion Tracking

255. Revisiting Unknowns: Towards Effective and Efficient Open-Set Active Learning

256. Parameter-Efficient Semantic Augmentation for Enhancing Open-Vocabulary Object Detection

257. InvAD: Inversion-based Reconstruction-Free Anomaly Detection with Diffusion Models

258. Attribute-Preserving Pseudo-Labeling for Diffusion-Based Face Swapping

259. Portable Active Learning for Object Detection

260. IBISAgent: Reinforcing Pixel-Level Visual Reasoning in MLLMs for Universal Biomedical Object Referring and Segmentation

261. PaQ-DETR: Learning Pattern and Quality-Aware Dynamic Queries for Object Detection

262. When Transformers Meet Mamba: A Hybrid Transformer-Mamba Network for Video Object Detection

263. Activation Matters: Test-time Activated Negative Labels for OOD Detection with Vision-Language Models

264. FisherPoser: Human Motion Estimation from Sparse Observations with Hierarchical Region-Wise Fisher-Matrix Uncertainty Modeling

265. DarkAct: A RGB-Thermal Dataset and Fusion Framework for Multimodal Low-Light Action Recognition

266. PAF: Perturbation-Aware Filtering for Open-Set Semi-Supervised Learning

267. M3Grounder: Mask-Based Multi-Span and Multi-Granular Grounding for Document QA

268. Towards Persistence: Learning Topological Constraints for Event-based Small Object Detection

269. ShadowDraw: From Any Object to Shadow-Drawing Compositional Art

270. MooCap: A Multi-View Benchmark for Cow-Object-Human Interaction and Behavior Dynamics

271. Geometry-Aligned and Anomaly-Aware Reconstruction for 3D Anomaly Detection

272. WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval

273. CVA: Context-aware Video-text Alignment for Video Temporal Grounding

274. APEX: A Decoupled Memory-based Explorer for Asynchronous Aerial Object Goal Navigation

275. Enhancing Part-Level Point Grounding for Any Open-Source MLLMs

276. Mitigating Objectness Bias and Region-to-Text Misalignment for Open-Vocabulary Panoptic Segmentation

277. EReCu: Pseudo-label Evolution Fusion and Refinement with Multi-Cue Learning for Unsupervised Camouflage Detection

278. PAS: Prelim Attention Score for Detecting Object Hallucinations in Large Vision-Language Models

279. Captain Safari: A World Engine with Pose-Aligned 3D Memory

280. CCF: Complementary Collaborative Fusion for Domain Generalized Multi-Modal 3D Object Detection

281. FedCART: Tackling Long-Tailed Distributions in Federated Adversarial Training via Classifier Refinement

282. OmniVTG: A Large-Scale Dataset and Training Paradigm for Open-World Video Temporal Grounding

283. Efficient Video Object Segmentation and Tracking with Recurrent Dynamic Submodel

284. SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion

285. TTL: Test-time Textual Learning for OOD Detection with Pretrained Vision-Language Models

286. Fine-Grained Multi Image Object Hallucination Benchmark

287. MS^2Gait: A Multi-Scale Spatio-Temporal Fusion Network for LiDAR-based Gait Recognition

288. Learning to Identify Out-of-Distribution Objects for 3D LiDAR Anomaly Segmentation

289. Object-Generalized Re-Identification: A Step Towards Universal Instance Perception

290. Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding

291. Pointer-CAD: Unifying B-Rep and Command Sequences via Pointer-based Edges & Faces Selection

292. Decompose and Transfer: CoT-Prompting Enhanced Alignment for Open-Vocabulary Temporal Action Detection

293. Reinforcement-Guided Synthetic Data Generation for Privacy-Sensitive Identity Recognition

294. Faithful Contouring: Near-Lossless 3D Voxel Representation Free from Iso-surface

295. Unposed-to-3D: Learning Simulation-Ready Vehicles from Real-World Images

296. Distribution-Aligned Multimodal Fusion for Robust Object Detection

297. SRA-Det: Learning Omni-Grained Open-Vocabulary Detection Beyond Category Names

298. BUSSARD: Normalizing Flows for Bijective Universal Scene-Specific Anomalous Relationship Detection

299. Probabilistic Concept Graph Reasoning for Multimodal Misinformation Detection

300. AffordGen: Generating Diverse Demonstrations for Generalizable Object Manipulation with Affordance Correspondence

301. QuCNet: Quantum Deep Learning Driven Multi-Circuit Network for Remote Sensing Image Classification

302. ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned Representation

303. GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding

304. Ground Reaction Inertial Poser: Physics-based Human Motion Capture from Sparse IMUs and Insole Pressure Sensors

305. Monocular Open Vocabulary Occupancy Prediction for Indoor Scenes

306. Unleashing the Power of Chain-of-Prediction for Monocular 3D Object Detection

307. Learning from Noisy Supervision: A Denoising-Debiasing Framework for Weakly Supervised Video Anomaly Detection

308. Revisiting Pose Sensitivity in Splat-based Computed Tomography under Sparse-view Reconstruction

309. Real2Edit2Real: Generating Robotic Demonstrations via a 3D Control Interface

310. MedLoc-R1: Performance-Aware Curriculum Reward Scheduling for GRPO-Based Medical Visual Grounding

311. Particulate: Feed-Forward 3D Object Articulation

312. CF-IPT: Cross-Modal Fusion Interactive Prompt Tuning of Vision-Language Pre-Trained Model for Multisource Remote Sensing Data Classification

313. Lyapunov Probes for Hallucination Detection in Large Foundation Models

314. SemLT3D: Semantic-Guided Expert Distillation for Camera-only Long-Tailed 3D Object Detection

315. Mitigating Simplicity Bias in OOD Detection through Object Co-occurrence Analysis

316. UIKA: Fast Universal Head Avatar from Pose-Free Images

317. VisualAD: Language-Free Zero-Shot Anomaly Detection via Vision Transformer

318. Streamlined Open-Vocabulary Human-Object Interaction Detection

319. Multi-Prototype Compactness and Boundary-Aware Synthesis for Unsupervised Anomaly Detection

320. BulletTime: Decoupled Control of Time and Camera Pose for Video Generation

321. Concept-Guided Fine-Tuning: Steering ViTs away from Spurious Correlations to Improve Robustness

322. OpenMarcie: Dataset for Multimodal Action Recognition in Industrial Environments

323. No Way To Steal My Face: Proactive Defense Against Identity-Preserving Personalized Generation

324. From 3D Pose to Prose: Biomechanics-Grounded Vision-Language Coaching

325. Seeing Through the Noise: Improving Infrared Small Target Detection and Segmentation from Noise Suppression Perspective

326. Beyond Reassembly: Fractured Object Recovery with Missing Parts

327. VMD-FACT: A New Video Dataset and MLLM-based method for Detecting Realistic AI-Generated Video Misinformation

328. Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination

329. Seeing Beyond 8bits: Subjective and Objective Quality Assessment of HDR-UGC Videos

330. Seeing Both Sides: Towards Bidirectional Semantic Alignment for Open-Vocabulary Camouflaged Object Segmentation

331. Expanding mmWave Datasets for Human Pose Estimation with Unlabeled Data and LiDAR Datasets

332. Copy-Transform-Paste: Zero-Shot Object-Object Alignment Guided by Vision-Language and Geometric Constraints

333. ORIC: Benchmarking Object Recognition under Contextual Incongruity in Large Vision-Language Models

334. Pixels Don't Lie (But Your Detector Might): Bootstrapping MLLM-as-a-Judge for Trustworthy Deepfake Detection and Reasoning Supervision

335. Hyperbolic Defect Feature Synthesis for Few-Shot Defect Classification

336. HVG-3D: Bridging Real and Simulation Domains for 3D-Conditional Hand-Object Interaction Video Synthesis

337. Beyond Appearance: Camouflaged Object Detection via Geometric Structure

338. Learning to Focus and Precise Cropping:A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs

339. Contrastive Cross-Bag Augmentation for Multiple Instance Learning-based Whole Slide Image Classification

340. First Logit Boosting: Visual Grounding Method to Mitigate Object Hallucination in Large Vision-Language Models

341. OmniGround: A Comprehensive Spatio-Temporal Grounding Benchmark for Real-World Complex Scenarios

342. UETrack: A Unified and Efficient Framework for Single Object Tracking

343. More Natural, More Real: Object-aware Gaussian Splatting for 3D Visual Decoding from Human Brain

344. Prototype-based Causal Intervention for Multi-Label Image Classification

345. Part2^{2}GS: Part-aware Modeling of Articulated Objects using 3D Gaussian Splatting

346. Phantom: Physical Object Interactions as Dynamic Triggers for NMS-Exploited Backdoors

347. Evolutionary Multimodal Reasoning via Hierarchical Semantic Representation for Intent Recognition

348. HyperGait: Unleashing the Power of Parsing for Gait Recognition in the Wild via Hypergraph

349. YOLO-ULM: Ultra-Lightweight Models for Real-Time Object Detection

350. Hypergraph-State Collaborative Reasoning for Multi-Object Tracking

351. Beyond Prompt Degradation: Prototype-guided Dual-pool Prompting for Incremental Object Detection

352. HieraMamba: Video Temporal Grounding via Hierarchical Anchor-Mamba Pooling

353. WeDetect: Fast Open-Vocabulary Object Detection as Retrieval

354. Spike-driven Discrete Aggregation for Event-based Object Detection

355. Precise Object and Effect Removal with Adaptive Target-Aware Attention

356. TriDF: Evaluating Perception, Detection, and Hallucination for Interpretable DeepFake Detection

357. UniMMAD: Unified Multi-Modal and Multi-Class Anomaly Detection via MoE-Driven Feature Decompression

358. PoseAnything: General Pose-guided Video Generation with Part-aware Temporal Coherence

359. Sparsity as a Key: Unlocking New Insights from Latent Structures for Out-of-Distribution Detection

360. Remedying Target-Domain Astigmatism for Cross-Domain Few-Shot Object Detection

361. Revisiting F-measure Optimization in Multi-Label Classification: A Sampling-based Approach

362. PoseD-Flow: Versatile and Guided Flow Matching Model of Human Pose

363. MoBind: Motion Binding for Fine-Grained IMU-Video Pose Alignment

364. HAMMER: Harnessing MLLMs via Cross-Modal Integration for Intention-Driven 3D Affordance Grounding

365. Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in Scenes

366. Learning Scene Coordinate Reconstruction from Unposed Images via Pose Graph Optimization

367. Memory-Augmented Scene Understanding and Exploration for Open-World Aerial Object-Goal Navigation

368. Text-guided Feature Disentanglement for Cross-modal Gait Recognition

369. Training-Free Open-Vocabulary Camouflaged Object Segmentation via Fine-Grained Object Binding and Adaptive Hybrid Prompt

370. RAVEN: Radar Adaptive Vision Encoders for Efficient Chirp-wise Object Detection and Segmentation

371. Generative Video Motion Editing with 3D Point Tracks

372. Event6D: Event-based Novel Object 6D Pose Tracking

373. MMLandmarks: a Cross-View Instance-Level Benchmark for Geo-Spatial Understanding

374. Hearing the Room Through the Shape of the Drum: Modal-Guided Sound Recovery from Multi-Point Surface Vibrations

375. Open-world Hand-Object Interaction Video Generation Based on Structure and Contact-aware Representation

376. Tackling Model Bias via Game-theoretic Multi-agent Collaboration Framework for Hateful Meme Classification

377. Unlocking Strong Supervision: A Data-Centric Study of General-Purpose Audio Pre-Training Methods

378. Face-Guided Sentiment Boundary Enhancement for Weakly-Supervised Temporal Sentiment Localization

379. From Pixel to Precision: Enhancing Handwritten Mathematical Expression Recognition with Image-Level Reward

380. GMT: Effective Global Framework for Multi-Camera Multi-Target Tracking

381. BarbieGait: An Identity-Consistent Synthetic Human Dataset with Versatile Cloth-Changing for Gait Recognition

382. PC-Talk: Precise Facial Animation Control for Audio-Driven Talking Face Generation

383. R4Det: 4D Radar-Camera Fusion for High-Performance 3D Object Detection

384. Enhancing Out-of-Distribution Detection with Extended Logit Normalization

385. Adaptive Confidence Regularization for Multimodal Failure Detection

386. Adaptive Depth Lightweight RGB-T Tracking with Holistic Token Routing

387. ProgTrack: A Multi-Object Tracking Algorithm with Progressive Matching Strategy

388. HyperGaussians: High-Dimensional Gaussian Splatting for High-Fidelity Animatable Face Avatars

389. FocusUI: Efficient UI Grounding via Position-Preserving Visual Token Selection

390. Mesh4D: 4D Mesh Reconstruction and Tracking from Monocular Video

391. EchoPOSE: 6D Pose Estimation of Sparse Echocardiograms for Left-Ventricular 3D Shape Reconstruction

392. V^2-SAM: Marrying SAM2 with Multi-Prompt Experts for Cross-View Object Correspondence

393. CHAL: Causal-guided Hierarchical Anomaly-aware Learning for Moving Infrared Small Target Detection

394. DeRVOS: Decoupling Consistent Trajectory Generation and Multimodal Understanding for Referring Video Object Segmentation

395. Decompose, Mix, Adapt: A Unified Framework for Parameter-Efficient Neural Network Recombination and Compression

396. Zoo3D: Zero-Shot 3D Object Detection at Scene Level

397. Thinking Beyond Labels: Vocabulary-Free Fine-Grained Recognition using Reasoning-Augmented LMMs

398. What's Wrong with Synthetic Data for Scene Text Recognition? A Strong Synthetic Engine with Diverse Simulations and Self-Evolution

399. H-Sets: Hessian-Guided Discovery of Set-Level Feature Interactions in Image Classifiers

400. Training-free Detection of Generated Videos via Spatial-Temporal Likelihoods

401. TrajRAG: Retrieving Geometric-Semantic Experience for Zero-Shot Object Navigation

402. A Closer Look at Cross-Domain Few-Shot Object Detection: Fine-Tuning Matters and Parallel Decoder Helps

403. Spe-BEVHead: Rethinking the Detection Head Design for Bird's-Eye-View Object Detection

404. MVP: Multiple View Prediction Improves GUI Grounding

405. Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding

406. Cross-Modal Guided Visual Synthesis for Data-Efficient Multimodal Depression Recognition

407. Unsupervised Monocular 3D Keypoint Discovery from Multi-View Diffusion Priors

408. Towards Generalizable AI-Generated Image Detection via Image-Adaptive Prompt Learning

409. Universal-to-Specific: Dynamic Knowledge-Guided Multiple Instance Learning for Few-Shot Whole Slide Image Classification

410. VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding

411. MEMO: Human-like Crisp Edge Detection Using Masked Edge Prediction

412. PIX-TAB: Efficient PIXel-Precise TABle Structure Recognition Approach with Speculative Decoding and Region-Based Image Segmentation

413. High-Fidelity Diffusion Face Swapping with ID-Constrained Facial Conditioning

414. Weakly Supervised Video Anomaly Detection with Anomaly-Connected Components and Intention Reasoning

415. Parameterized Prompt for Incremental Object Detection

416. VLM4RSDet: Collaborative Optimization with Vision-Language Model for Enhancing Remote Sensing Object Detection

417. Compositional Transformation Reasoning for Composed Video Retrieval

418. RAID: Retrieval-Augmented Anomaly Detection

419. Reconstruction-Guided Slot Curriculum: Addressing Object Over-Fragmentation in Video Object-Centric Learning

420. Dual-Level Hypergraph Generation for Addressing Feature Scarcity in Whole-Slide Image Classification

421. LAM: Language Articulated Object Modelers

422. TAPE: Task-Adaptive Prototype Evolution in Audio-Language Models for Fully Few-shot Class-incremental Audio Classification

423. IAG: Input-aware Backdoor Attack on VLM-based Visual Grounding

424. Homaloidal parametrization for detecting critical two-view configurations

425. Detect Anything via Next Point Prediction

426. Harmonious Parameter Adaptation in Continual Visual Instruction Tuning for Safety-Aligned MLLMs

427. From Pairs to Sequences: Track-Aware Policy Gradients for Keypoint Detection

428. Opti-NeuS: Neural Reconstruction for Dual-Layered Transparent and Opaque Objects

429. EventGait: Towards Robust Gait Recognition with Event Streams

430. AssemblyBench: Physics-Aware Assembly of Complex Industrial Objects

431. The SA-FARI Dataset: Segment Anything in Footage of Animals for Recognition and Identification

432. CIGPose: Causal Intervention Graph Neural Network for Whole-Body Pose Estimation

433. Beyond Objects: Contextual Synthetic Data Generation for Fine-Grained Classification

434. Hermite Radial Basis Function for Surface Reconstruction via Differentiable Rendering

435. Learning 3D Representations for Spatial Intelligence from Unposed Multi-View Images

436. PPM-CLIP: Probabilistic Prompt Modeling for Generalizable AI-Generated Image Detection

437. Beyond Weak Supervision: MLLMs-Guided Graded Knowledge Distillation for Unsupervised Camouflaged Object Detection

438. Scene Reconstruction as Mapping Priors for 3D Detection

439. AlignPose: Generalizable 6D Pose Estimation via Multi-view Feature-metric Alignment

440. BOP-ASK: Object-Interaction Reasoning for Vision-Language Models

441. M4-SAM: Multi-Modal Mixture-of-Experts with Memory-Augmented SAM for RGB-D Video Salient Object Detection

442. 3D-Object Perception Transformer (3PT)

443. InterPrior: Scaling Generative Control for Physics-Based Human-Object Interactions

444. Expert-Teacher-Student Collaborative Learning for Domain Adaptive Object Detection

445. OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery

446. A Cross-view Fusion Framework for Robust 6-DoF Grasp Pose Estimation

447. DARC: Dual Adjustment Reasoning with Counterfactuals for Trustworthy Chest X-ray Classification

448. Explaining Object Detectors via Collective Contribution of Pixels

449. AntiStyler: Defending Object Detection Models Against Adversarial Patch Attacks Using Style Removal

450. ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval

451. VIMCAN: Visual-Inertial 3D Human Pose Estimation with Hybrid Mamba-Cross-Attention Network

453. AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos

454. Partial Weakly-Supervised Oriented Object Detection

455. FMPose3D: monocular 3D pose estimation via flow matching

456. A Sanity Check for Multi-In-Domain Face Forgery Detection in the Real World

457. Towards Balanced Multi-Modal Learning in 3D Human Pose Estimation

458. Few-Shot Incremental 3D Object Detection in Dynamic Indoor Environments

459. SARL-STG: A Spatially Aware Reinforcement Learning Framework for Refining MLLMs in Spatio-Temporal Video Grounding

460. EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR

461. EV-CGNet: Co-visible Focused 3D-guided 2D Event Keypoint Detection Network

462. Hierarchical Point-Patch Fusion with Adaptive Patch Codebook for 3D Shape Anomaly Detection

463. UniGenDet: A Unified Generative-Discriminative Framework for Co-Evolutionary Image Generation and Generated Image Detection

464. SABER: Spatially Consistent 3D Universal Adversarial Objects for BEV Detectors

465. Animator-Centric Skeleton Generation on Objects with Fine-Grained Details

466. 4DSurf: High-Fidelity Dynamic Scene Surface Reconstruction

467. BeautyGRPO: Aesthetic Alignment for Face Retouching via Dynamic Path Guidance and Fine-Grained Preference Modeling

468. DarkShake-DVS: Event-based Human Action Recognition under Low-light and Shaking Camera Conditions

469. UTPTrack: Towards Simple and Unified Token Pruning for Visual Tracking

470. Clay-to-Stone: Phase-wise 3D Gaussian Splatting for Monocular Articulated Hand-Object Manipulation Modeling

471. Generative Point Tracking and Forecasting

472. LAMP: Localization Aware Multi-camera People Tracking in Metric 3D World

473. Unlocking Motion from Large Vision Models with a Semantic and Kinematic Duality for Gait Recognition

474. Dual-level Adaptation for Multi-Object Tracking: Building Test-Time Calibration from Experience and Intuition

475. Revisiting Token Compression for Accelerating ViT-based Sparse Multi-View 3D Object Detectors

476. CD-Buffer: Complementary Dual-Buffer Framework for Test-Time Adaptation in Adverse Weather Object Detection

477. SOTA: Self-adaptive Optimal Transport for Zero-Shot Classification with Multiple Foundation Models

478. SkeletonContext: Skeleton-side Context Prompt Learning for Zero-Shot Skeleton-based Action Recognition

479. GenHOI: Towards Object-Consistent Hand-Object Interaction with Temporally Balanced and Spatially Selective Object Injection

480. Learning to Track Instance from Single Nature Language Description

481. ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose Estimation

482. Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection

483. YOLO-Master: MOE-Accelerated with Specialized Transformers for Enhanced Real-time Detection

484. ORSATR-X: A Foundation Model based on Differential-and-Excitation Networks for Optical Remote Sensing Object Recognition

485. MOGeo: Beyond One-to-One Cross-View Object Geo-localization

486. AdaDexTrack: Dynamic Modulation for Adaptive and Generalizable Dexterous Manipulation Tracking

487. SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation Model

488. ULF-Loc: Unbiased Landmark Feature for Robust Visual Localization with 3D Gaussian Splatting

489. Registration-Free Learnable Multi-View Capture of Faces in Dense Semantic Correspondence

490. Protect to Adapt: Orthogonal Subspace Control with Ranked Negative-Prompt Curriculum for Few-Shot Action Recognition

491. Differentially Private 2D Human Pose Estimation

492. OMoBlur: An Object Motion Blur Dataset and Benchmark for Real-World Local Motion Deblurring

493. SpatialDiff: 3D-Aware Object Movement via Implicit Spatial Modeling

494. Beyond Missing Modalities: Hypergraph Conditioned Diffusion for Uncertainty-Aware Multimodal Emotion Recognition

495. Cross-View Distillation and Adaptive Masking for Incomplete Multi-View Multi-Label Classification

496. Anatomical Domain Shifts: Test-time Heterogeneous Adaptation for 3D Human Pose Prediction

497. SubspaceAD: Training-Free Few-Shot Anomaly Detection via Subspace Modeling

498. 3D Gaussian Splatting with Self-Constrained Priors for High Fidelity Surface Reconstruction

499. Ghost-FWL: A Large-Scale Full-Waveform LiDAR Dataset for Ghost Detection and Removal

500. Detecting Unknown Objects via Energy-based Separation for Open World Object Detection

501. PerformRecast: Expression and Head Pose Disentanglement for Portrait Video Editing

502. MoVie: Broaden Your Views with Human Motion for Action Detection

503. D^3FER: Dual Channel and Dual Branch Network for Robust Facial Expression Recognition under Dual Challenges

504. Choreographing a World of Dynamic Objects

506. Fine-VAD: Towards Fine-Grained Video Anomaly Detection via Progressive Cross-Granularity Learning

507. ReGenHOI: Unifying Reconstruction and Generation for 3D Human-Object Interaction Understanding

508. SPAN: Spatial-Projection Alignment for Monocular 3D Object Detection

509. Detecting Compressed AI-Generated Images via Phase Spectrum Robustness

510. Exposing and Evaluating Hallucinations for GUI Grounding

511. DetAny4D: Detect Anything 4D Temporally in a Streaming RGB Video

512. CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects

513. MMSD3.0: A Multi-Image Benchmark for Real-World Multimodal Sarcasm Detection

514. Harnessing Chain-of-Thought Reasoning in Multimodal Large Language Models for Face Anti-Spoofing

515. Pose-guided Enriched Feature Learning for Federated-by-camera Person Re-identification

516. Goldilocks Test Sets for Face Verification

517. Investigating Self-Supervised Representations for Audio-Visual Deepfake Detection

518. TF-SSD: A Strong Pipeline via Synergic Mask Filter for Training-free Co-salient Object Detection

519. BEV-SLD: Self-Supervised Scene Landmark Detection for Global Localization with LiDAR Bird's-Eye View Images

520. CORE: Compact Object-centric REpresentations as a New Paradigm for Token Merging in LVLMs

521. MarkushGrapher-2: End-to-end Multimodal Recognition of Chemical Structures

522. MoECLIP: Patch-Specialized Experts for Zero-shot Anomaly Detection

523. Toward Low-Cost yet Effective Temporal Learning for UAV Tracking

524. MMR-AD: A Large-Scale Multimodal Dataset for Benchmarking General Anomaly Detection with Multimodal Large Language Models

525. AnomalyVFM -- Transforming Vision Foundation Models into Zero-Shot Anomaly Detectors

526. Learning Latent Concepts for Detecting Out-of-Distribution Objects

527. PLACID: Identity-Preserving Multi-Object Compositing via Video Diffusion with Synthetic Trajectories

528. Dynamic Label Noise Suppression with Optimal Teacher Pool for Facial Expression Recognition

529. YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object Removal

530. CoIn3D: Revisiting Configuration-Invariant Multi-Camera 3D Object Detection

531. VGent: Visual Grounding via Modular Design for Disentangling Reasoning and Prediction

532. ExPose: Reinforcing Video Generation Models for Extreme Pose Estimation

533. DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding

534. DriverGaze360: OmniDirectional Driver Attention with Object-Level Guidance

535. Revisiting 2D Foundation Models for Scalable 3D Medical Image Classification

536. InnoAds-Composer: Efficient Condition Composition for E-Commerce Poster Generation

537. Progressive Multi-cue Alignment for Unaligned RGBT Tracking

538. MPL: Match-guided Prototype Learning for Few-shot Action Recognition

539. Prompt-Free Unknown Label Generation for Open World Detection in Remote Sensing

540. Cross-modal Representation Learning for Diffusion-generated Image Detection

541. Beyond Static Frames: Temporal Aggregate-and-Restore Vision Transformer for Human Pose Estimation

542. HOLO: Homography-Guided Pose Estimator Network for Fine-Grained Visual Localization on SD Maps

543. OVOD-Agent: A Markov-Bandit Framework for Proactive Visual Reasoning and Self-Evolving Detection

544. Consistency Beyond Contrast: Enhancing Open-Vocabulary Object Detection Robustness via Contextual Consistency Learning

545. NeuROK: Generative 4D Neural Object Kinematics

546. Detecting AI-Generated Forgeries via Iterative Manifold Deviation Amplification

547. Tracking through Severe Occlusion via Event-Derived Transient Cues

548. Stability-Driven Motion Generation for Object-Guided Human-Human Co-Manipulation

549. CLEP: Contrastive Language-Pose Pretraining

550. AutoDebias: An Automated Framework for Detecting and Mitigating Backdoor Biases in Text-to-Image Models

551. Chain-of-Thought Guided Multi-Modal Object Re-Identification

552. PP-Brep: Few-Shot B-rep Classification with Hybrid Graph Representation

553. STUR3D: Spatio-Temporal Unified Representation Learning for 3D Object Detection

554. Tracking by Predicting 3-D Gaussians Over Time

555. TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free Reconstruction

556. PAMotion: Physics-Aware Motion Generation for Full-Body Interaction with Multiple Objects

557. RGB-Event based Pedestrian Attribute Recognition: A Benchmark Dataset and An Asymmetric RWKV Fusion Framework

558. G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval

559. VIRD: View-Invariant Representation through Dual-Axis Transformation for Cross-View Pose Estimation

560. Trust-calibrated Collaborative Learning for Long-Tailed Visual Recognition

561. Rethinking Two-Stage Referring-by-Tracking in Referring Multi-Object Tracking: Make it Strong Again

562. PriVi: Towards a General-Purpose Video Model for Primate Behavior in the Wild

563. SFR-Net: Steering-Fusion-Refining Network in Multi-label Zero-Shot Sewer Defect Detection

564. Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval

565. Global-Aware Edge Prioritization for Pose Graph Initialization

566. Geometry-driven OOD Detectors Are Class-Incremental Learners

567. A Semantically Disentangled Unified Model for Multi-category 3D Anomaly Detection

568. AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMs

569. 3D Gaussian Splatting from Unposed Spike Stream

570. Post-training Feature Pruning for Fundus Images Classification

571. Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos

572. Your Classifier Can Do More: Towards Balancing the Gaps in Classification, Robustness, and Generation

573. Progressive Cross-Modal Causal Intervention for Long-Term Action Recognition

574. OpenFS: Multi-Hand-Capable Fingerspelling Recognition with Implicit Signing-Hand Detection and Frame-Wise Letter-Conditioned Synthesis

575. TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly Detection

576. EXOTIC: External Vision-driven Incomplete Multi-view Classification

577. A Difference-in-Difference Approach to Detecting AI-Generated Images

578. Ego-Grounding for Personalized Question-Answering in Egocentric Videos

579. SeD-UD: An Influence-Driven and Hierarchically-Decoupled Information Bottleneck for Multimodal Intent Recognition

580. LayoutAD: Exploring Semantic-Geometric Misalignment Reasoning for Scene Layout Anomaly Detection

581. Mining Instance-Centric Vision-Language Contexts for Human-Object Interaction Detection

582. CoWTracker: Tracking by Warping instead of Correlation

583. Transition Models: Rethinking the Generative Learning Objective

584. Bulk RNA-seq Guided Multi-modal Detection of Anomalous Regions in Human Cancer via Spatial Transcriptomics

585. CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective Video

586. ReManNet: A Riemannian Manifold Network for Monocular 3D Lane Detection

587. ViHOI: Human-Object Interaction Synthesis with Visual Priors

588. Confusion-Aware Spectral Regularizer for Long-Tailed Recognition

589. From Intuition to Investigation: A Tool-Augmented Reasoning MLLM Framework for Generalizable Face Anti-Spoofing

590. EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions

591. Specificity-aware reinforcement learning for fine-grained open-world classification

592. GS-CLIP: Zero-shot 3D Anomaly Detection by Geometry-Aware Prompt and Synergistic View Representation Learning

593. Rethinking BCE Loss for Multi-Label Image Recognition with Fine-Tuning

594. SE(3)-Equivariance with Geometric and Topological Guidance for Category-Level Object Pose Estimation

595. POLAR: A Portrait OLAT Dataset and Generative Framework for Illumination-Aware Face Modeling

596. Mind the Way You Select Negative Texts: Pursuing the Distance Consistency in OOD Detection with VLMs

597. TagSplat: Topology-Aware Gaussian Splatting for Dynamic Mesh Modeling and Tracking

598. Machine Unlearning via Adaptive Gradient Reweighting and Multi-stage Objective Optimization

599. LocateAnything3D: Vision-Language 3D Detection with Chain-of-Sight

600. Thermal-Det: Language-Guided Cross-Modal Distillation for Open-Vocabulary Thermal Object Detection

601. No Need For Real Anomaly: MLLM Empowered Zero-Shot Video Anomaly Detection

602. Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields

603. GPFlow: Gaussian Prototype Probability Flow for Unsupervised Multi-Modal Anomaly Detection

604. MVLM: Template-Free Tracking via Vision-Language Margin Confidence and Memory-Gated Tracking

605. Discover, Segment, and Select: A Progressive Mechanism for Zero-shot Camouflaged Object Segmentation

606. Decoupling Bias, Aligning Distributions: Synergistic Fairness Optimization for Deepfake Detection

607. Anti-I2V: Safeguarding your Photos from Malicious Image-to-video Generation

608. Structure-Aware Representation Distillation for Tiny-Dense Object Segmentation

609. Towards Stealthy and Effective Backdoor Attacks on Lane Detection: A Naturalistic Data Poisoning Approach

610. Small Object, Great Challenge: A Benchmark for Small Object Visual Grounding

611. Cov2Pose: Leveraging Spatial Covariance for Direct Manifold-aware 6-DoF Object Pose Estimation

612. BlackMirror: Black-Box Backdoor Detection for Text-to-Image Models via Instruction-Response Deviation

613. PV-Ground: Text-Guided Point-Voxel Interaction for 3D Visual Grounding

614. Uni3R: Unified 3D Reconstruction and Semantic Understanding via Generalizable Gaussian Splatting from Unposed Multi-View Images

615. Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection

616. Boosting Self-Supervised Tracking with Contextual Prompts and Noise Learning

617. PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation

618. Finding Distributed Object-Centric Properties in Self-Supervised Transformers

619. KLIP: Localized Distribution Shift Detection via KL-Divergence with Diffusion Priors in Inverse Problems

620. Scaling Up AI-Generated Image Detection with Generator-Aware Prototypes

621. Object-WIPER: Training-Free Object and Associated Effect Removal in Videos

622. The LLM Bottleneck: Why Open-Source Vision LLMs Struggle with Hierarchical Visual Recognition

623. SEATrack: Simple, Efficient, and Adaptive Multimodal Tracker

624. Molmo2: Open Weights and Data for Vision-Language Models with Video Understanding and Grounding

625. MonoVLM: Monocular 3D Visual Grounding with Vision Language Models

626. Learning to Diversify and Focus: A Reinforcement Framework for Open-Vocabulary HOI Detection

627. From Attraction to Equilibrium: Physics-Inspired Semantic Gravitons for Zero-Shot Anomaly Detection

628. Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation

629. MER-Tracker: Towards High-Speed 3D Point Tracking via Multi-View Event-RGB Hybrid Cameras

630. Detect Any AI-Counterfeited Text Image

631. Diversity over Uniformity: Rethinking Representation in Generated Image Detection

632. See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding

633. GOR-IS: 3D Gaussian Object Removal In the Intrinsic Space

634. IMU-HOI: A Symbiotic Framework for Coherent Human-Object Interaction and Motion Capture via Contact-Conscious Inertial Fusion

635. Graph Attention Prototypical Network for Robust Few-Shot Classification

636. MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision

637. Robust Promptable Video Object Segmentation

638. Reasoning-Driven Anomaly Detection and Localization with Image-Level Supervision

639. UAV-CB: A Complex-Background RGB-T Dataset and Local Frequency Bridge Network for UAV Detection

640. RealVLG-R1: A Large-Scale Real-World Visual-Language Grounding Benchmark for Robotic Perception and Manipulation

641. LA-Pose: Latent Action Pretraining Meets Pose Estimation

642. TTAPFormer: Robust Arbitrary Point Tracking via Transient Asynchronous Fusion of Frames and Events

643. OLATverse: A Large-scale Real-world Object Dataset with Precise Lighting Control

644. Beyond [CLS] Token: Query-Driven Token-Level Forgery Purification for Generalizable Deepfake Detection

645. RDF-MIG: A Robust Diffusion Framework for Masked Image Generation to Augment Semantic Segmentation and Change Detection

646. AgentDet: A Shared-Blackboard Multi-Agent Framework for Zero-/Few-Shot Object Detection

647. SimLBR: Learning to Detect Fake Images by Learning to Detect Real Images

648. FaceCam: Portrait Video Camera Control via Scale-Aware Conditioning

649. IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation

650. EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature Enhancement

651. H^2A^2: Homogeneity-Aware and Heterogeneity-Aware Feature Perception for Unified Indoor 3D Object Detection

652. D3D-VLP: Dynamic 3D Vision-Language-Planning Model for Embodied Grounding and Navigation

653. IDperturb: Enhancing Variation in Synthetic Face Generation via Angular Perturbations

654. TESO: Online Tracking of Essential Matrix by Stochastic Optimization

655. DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge

656. Mixture-of-Experts based Feature Decoupling for Open Vocabulary Scene Graph Generation

657. Exploring 6D Object Pose Estimation with Deformation

658. DRS-GUI: Dynamic Region Search for Training-Free GUI Grounding

659. Towards Intrinsic-Aware Monocular 3D Object Detection

660. UniMERNet: A Universal Network for Real-World Mathematical Expression Recognition

661. Fourier Angle Alignment for Oriented Object Detection in Remote Sensing

662. AeroGS: Scale-Aware Gaussian Splatting for Pose-Free Dynamic UAV Scene Reconstruction

663. InsCal: Calibrated Multi-Source Fully Test-Time Prompt Tuning for Object Detection

664. Bidirectional Multimodal Prompt Learning with Scale-Aware Training for Few-Shot Multi-Class Anomaly Detection

665. ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval

666. MUSE: Harnessing Precise and Diverse Semantics for Few-Shot Whole Slide Image Classification

667. Tutor-Student Reinforcement Learning: A Dynamic Curriculum for Robust Deepfake Detection

668. WalkGPT: Grounded Vision-Language Conversation with Depth-Aware Segmentation for Pedestrian Navigation

669. DLVP-CLIP: Enhancing Fine-Grained Zero-Shot Anomaly Detection via Dynamic Local Visual Prompting

670. PhysHO: Physics-Based Dynamic 3D Gaussian Human and Object from Monocular Video

671. EvObj: Learning Evolving Object-centric Representations for 3D Instance Segmentation without Scene Supervision

672. Joint Learning of General and Diverse Patterns with Mixture of Memory Experts for Weakly-Supervised Video Anomaly Detection

673. Grounding Everything in Tokens for Multimodal Large Language Models

674. Uncertainty-Aware Modality Fusion for Unaligned RGB-T Salient Object Detection

675. ImmerIris: A Large-Scale Dataset and Benchmark for Off-Axis and Unconstrained Iris Recognition in Immersive Applications

676. Foundation Model Priors Enhance Object Focus in Feature Space for Source-Free Object Detection

677. DF^2-VB: Dual-level Fuzzy Fusion with View-specific Boosting for Multi-view Multi-label Classification

678. Taxonomy-Aware Representation Alignment for Hierarchical Visual Recognition with Large Multimodal Models

679. Tracking-Guided 4D Generation: Foundation-Tracker Motion Priors for 3D Model Animation

680. DetectSCI: Toward Object-Guided ROI Reconstruction for High-Resolution Video Snapshot Compressive Imaging

681. SToRe3D: Sparse Token Relevance in ViTs for Efficient Multi-View 3D Object Detection

682. FastRef: Fast Prototype Refinement for Few-shot Industrial Anomaly Detection

683. Generalizable Co-Salient Object Detection via Mixed Content-Style Modulation

684. Region-Aware Instance Consistency Learning for Micro-Expression Recognition

685. TESSERA: Temporal Embeddings of Surface Spectra for Earth Representation and Analysis

686. BDNet:Bio-Inspired Dual-Backbone Small Object Detection Network

687. RoboTAG: End-to-end Robot Pose Estimation via Topological Alignment Graph

688. RDFace: A Benchmark Dataset for Rare Disease Facial Image Analysis under Extreme Data Scarcity and Phenotype-Aware Synthetic Generation

689. FB-CLIP: Fine-Grained Zero-Shot Anomaly Detection with Foreground-Background Disentanglement

690. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval

691. Neu-PiG: Neural Preconditioned Grids for Fast Dynamic Surface Reconstruction on Long Sequences

692. Occlusion-Aware SORT: Observing Occlusion for Robust Multi-Object Tracking

693. AnthroTAP: Learning Point Tracking with Real-World Motion

694. Beyond the Global Scores: Fine-Grained Token Grounding as a Robust Detector of LVLM Hallucinations

695. See What We Cannot See: A Geo-guided Reasoning Benchmark for Object Counting under Adverse Earth Observation Conditions

696. WildPose: A Unified Framework for Robust Pose Estimation in the Wild

697. Seeing as Experts Do: A Knowledge-Augmented Agent for Open-Set Fine-Grained Visual Understanding

698. Seeing Motion Through Polarity for Event-based Action Recognition

699. BEA-GS: BEyond RAdiance Supervision in 3DGS for Precise Object Extraction

700. FusionAgent: A Multimodal Agent with Dynamic Model Selection for Human Recognition

701. ZINA: Multimodal Fine-grained Hallucination Detection and Editing

702. Bypassing the Transport Plan: Dynamic Reweighting for Out-of-Distribution Detection with Optimal Transport

703. TDATR: Improving End-to-End Table Recognition via Table Detail-Aware Learning and Cell-Level Visual Alignment

704. TRivia: Self-supervised Fine-tuning of Vision-Language Models for Table Recognition

705. DialogueVPR: Towards Conversational Visual Place Recognition

706. RC-NF: Robot-Conditioned Normalizing Flow for Real-Time Anomaly Detection in Robotic Manipulation

707. Why Not Hyperparameter-Friendly Optimisation? A Monotonic Adaptive Norm Rescaling Approach For Long-Tailed Recognition

708. Free-Grained Hierarchical Visual Recognition

709. KV-Tracker: Real-Time Pose Tracking with Transformers

710. Incremental Object Detection via Future-Aware Decoupled Cross-Head Distillation

711. Learning to Drive is a Free Gift: Large-Scale Label-Free Autonomy Pretraining from Unposed In-The-Wild Videos

712. SketchFaceGS: Real-Time Sketch-Driven Face Editing and Generation with Gaussian Splatting

713. DeepfakeImpact: A Two-Stage Benchmark with Real-World Impact in Deepfake Detection

714. Revisiting the Necessity of Full Accuracy: Weakly Supervised Object-Level Offset Correction for Misaligned Building Labels

715. SHands: A Multi-View Dataset and Benchmark for Surgical Hand-Gesture and Error Recognition Toward Medical Training

716. Changes in Real Time: Online Scene Change Detection with Multi-View Fusion

717. HERO: Hierarchical Embedding-Refinement for Open-Vocabulary Temporal Sentence Grounding in Videos

718. Conversational Image Segmentation: Grounding Abstract Concepts with Scalable Supervision

719. Interactive Tracking: A Human-in-the-Loop Paradigm with Memory-Augmented Adaptation

720. TriLite: Efficient Weakly Supervised Object Localization with Universal Visual Features and Tri-Region Disentanglement

721. DIMOS: Disentangling Instance-level Moving Object Segmentation

722. BiGain: Unified Token Compression for Joint Generation and Classification

723. Wavelet-Driven 3D Anomaly Detection under Pose-Agnostic and Sparse-View

724. Conan: Progressive Learning to Reason Like a Detective over Multi-Scale Visual Evidence

725. ORD: Object-Relation Decoupling for Generalized 3D Visual Grounding

726. SDTrack: A Baseline for Event-based Tracking via Spiking Neural Networks

727. SECOS: Semantic Capture for Rigorous Classification in Open-World Semi-Supervised Learning

728. Mechanisms of Object Localization in Vision-Language Models

729. Refer-Agent: A Collaborative Multi-Agent System with Reasoning and Reflection for Referring Video Object Segmentation

731. Online Data Curation for Object Detection via Marginal Contributions to Dataset-level Average Precision

732. Defect Cue-Preserved Structural Feature Refinement for Few-Shot Anomaly Detection

733. Layer Consistency Matters: Elegant Latent Transition Discrepancy for Generalizable Synthetic Image Detection

734. RankOOD - Class Ranking-based Out-of-Distribution Detection

735. BAMI: Training-Free Bias Mitigation in GUI Grounding

736. Scene Grounding in the Wild

737. Zero-shot Detection of AI-Generated Image via RAW-RGB Alignment

738. Hunting Normality from Query Sample via Residual Learning for Generalist Anomaly Detection

739. TeamHOI: Learning a Unified Policy for Cooperative Human-Object Interactions with Any Team Size

740. Making the Classification Explanation Faithful to the Confidence Score

741. Thinking in Dynamics: How Multimodal Large Language Models Perceive, Track, and Reason Dynamics in Physical 4D World

742. C^2FG: Control Classifier-Free Guidance via Score Discrepancy Analysis

743. Semantic Alignment for Pose-Invariant Identity Preserving Diffusion

744. UniChange: Unifying Change Detection with Multimodal Large Language Model

745. Adapting In-context Generation for Enhanced Composed Image Retrieval

746. Unleashing Vision-Language Semantics for Deepfake Video Detection

747. Beyond Explicit Language: Plug-and-Play Visual-to-Linguistic Modeling Toward General Object Tracking

748. EagleVision: A Dual-Stage Framework with BEV-grounding-based Chain-of-Thought for Spatial Intelligence

749. ArchSym: Detecting 3D-Grounded Architectural Symmetries in the Wild

750. Target-Aware Invertible Encoder with Reconstruction Guidance for Infrared Small Target Detection

751. The Invisible Gorilla Effect in Out-of-distribution Detection

752. Instance-level Visual Active Tracking with Occlusion-Aware Planning

753. Pointing at Parts: Training-Free Few-Shot Grounding in Multimodal LLMs

754. ElasticFormer: Detecting Objects in HRW Shots via Elastic Computing Vision Transformer

755. Alert-CLIP: Abnormality-aware Latent-Enhanced Representation Tuning of CLIP for Video Anomaly Detection

756. Good Can Sometimes be Bad: A Unified Attack against 3D Point Cloud Classifier by a Flexible Isotropic Resampling

757. Preserving Source Video Realism: High-Fidelity Face Swapping for Cinematic Quality

758. Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval