- Published on
CVPR 2026 — Image Retrieval & Matching
Image Retrieval & Matching
163 papers
1. Efficient and High-Fidelity Omni Modality Retrieval
- Link: Open Access
- arXiv: 2603.02098
2. Tackling Alignment Ambiguity in Person Retrieval through Conversational Attribute Mining
- Link: Open Access
3. Beyond Matching to Tiles: Bridging Unaligned Aerial and Satellite Views for Vision-Only UAV Navigation
- Link: Open Access
- arXiv: 2603.22153
4. RetFormer: Multimodal Retrieval for Enhancing Image Recognition
- Link: Open Access
5. FlowFM: Advancing Dark Optical Flow Estimation with Flow Matching
- Link: Open Access
6. Can We Build Scene Graphs, Not Classify Them? FlowSG: Progressive Image-Conditioned Scene Graph Generation with Flow Matching
- Link: Open Access
- arXiv: 2604.18623
7. Robust Remote Sensing Image-Text Retrieval with Noisy Correspondence
- Link: Open Access
- arXiv: 2603.28134
8. GuideFlow: Constraint-Guided Flow Matching for Planning in End-to-End Autonomous Driving
- Link: Open Access
- arXiv: 2511.18729
9. Beyond Soft Label: Dataset Distillation via Orthogonal Gradient Matching
- Link: Open Access
10. GazeOnce360: Fisheye-Based 360deg Multi-Person Gaze Estimation with Global-Local Feature Fusion
- Link: Open Access
11. Retrieve-to-Restore: Efficient All-in-One Image Restoration with a Retrieval-Based Degradation Bank
- Link: Open Access
12. MotionHiFlow: Text-to-Motion via Hierarchical Flow Matching
- Link: Open Access
- arXiv: 2604.23264
13. EvoGraph-R1: Self-Evolving Multimodal Knowledge Hypergraphs for Agentic Retrieval
- Link: Open Access
14. SAR2Net: Learning Spatially Anchored Representations for Retrieval-Guided Cross-Stain Alignment
- Link: Open Access
15. Language-driven Fine-grained Retrieval
- Link: Open Access
- arXiv: 2512.06255
16. Air-Know: Arbiter-Calibrated Knowledge-Internalizing Robust Network for Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.19386
17. Transition Matching Distillation for Fast Video Generation
- Link: Open Access
- arXiv: 2601.09881
18. Pip-Stereo: Progressive Iterations Pruner for Iterative Optimization based Stereo Matching
- Link: Open Access
- arXiv: 2602.20496
19. Towards Cross-Modal Preservation, Consistency and Alignment for Privacy-Preserving Visible-Infrared Person Re-Identification
- Link: Open Access
20. HypeVPR: Exploring Hyperbolic Space for Perspective to Equirectangular Visual Place Recognition
- Link: Open Access
- arXiv: 2506.04764
21. C-LaV: Conditional Latent Velocity Field Denoising for Weather-Robust LiDAR Place Recognition
- Link: Open Access
22. Intra-class Distribution-guided Generative Hashing with Neighbor Refinement for Cross-modal Retrieval
- Link: Open Access
23. Composite-Attribute Person Re-Identification via Pose-Guided Disentanglement
- Link: Open Access
24. Attribution as Retrieval: Model-Agnostic AI-Generated Image Attribution
- Link: Open Access
- arXiv: 2603.10583
25. R4-CGQA: Retrieval-based Vision Language Models for Computer Graphics Image Quality Assessment
- Link: Open Access
- arXiv: 2603.10578
26. Scaling Multi-Identity Consistency for Image Customization via Multi-to-Multi Matching Paradigm
- Link: Open Access
27. ProjFlow: Projection Sampling with Flow Matching for Zero-Shot Exact Spatial Motion Control
- Link: Open Access
28. Beyond Global Similarity: Multi-Conditional Retrieval for Fine-Grained Cross-Modal Understanding
- Link: Open Access
29. POGA: Paraphrased and Oppositional Graph Alignment for Fine-Grained Cross-Modal Retrieval
- Link: Open Access
30. PatchAlign3D: Local Feature Alignment for Dense 3D Shape Understanding
- Link: Open Access
- arXiv: 2601.02457
31. SSM-Aware Token-Efficient VMamba via Adaptive Patch Pruning and Merging for Person Re-Identification
- Link: Open Access
32. Scalable Feature Matching via State Space Modeling and Sparse Correlation
- Link: Open Access
33. MV-RoMa: From Pairwise Matching into Multi-View Track Reconstruction
- Link: Open Access
- arXiv: 2603.27542
34. RMIR: A Benchmark Dataset for Reasoning-Intensive Multimodal Image Retrieval
- Link: Open Access
35. Stable Mean Flow: Lyapunov-Inspired One-Step Flow Matching
- Link: Open Access
36. GeniNav: Generative Model Driven Image-Goal Navigation via Imagination-Guided Consistency Flow Matching
- Link: Open Access
37. DMGD: Train-Free Dataset Distillation with Semantic-Distribution Matching in Diffusion Models
- Link: Open Access
- arXiv: 2605.03877
38. STiTch: Semantic Transition and Transportation in Collaboration for Training-Free Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2605.21261
39. PinPoint: Evaluation of Composed Image Retrieval with Explicit Negatives, Multi-Image Queries, and Paraphrase Testing
- Link: Open Access
- arXiv: 2603.04598
40. COT-FM: Cluster-wise Optimal Transport Flow Matching
- Link: Open Access
- arXiv: 2603.13395
41. PromptStereo: Zero-Shot Stereo Matching via Structure and Motion Prompts
- Link: Open Access
- arXiv: 2603.01650
42. CogniVerse: Revolutionizing Multi-Modal Retrieval-Augmented Generation with Cognitive Reflection and Geometric Reasoning
- Link: Open Access
43. Frequency-Aware Flow Matching for High-Quality Image Generation
- Link: Open Access
- arXiv: 2604.15521
44. Universal 3D Shape Matching via Coarse-to-Fine Language Guidance
- Link: Open Access
- arXiv: 2602.19112
45. MRD: Multi-resolution Retrieval-Detection Fusion for High-Resolution Image Understanding
- Link: Open Access
- arXiv: 2512.02906
46. Spatial-Frequency Collaborative Learning for Occluded Visible-Infrared Person Re-Identification
- Link: Open Access
47. COPE: Consistent Occlusion and Prompt Enhancement Network for Occluded Person Re-identification
- Link: Open Access
48. Imagine Before Concentration: Diffusion-Guided Registers Enhance Partially Relevant Video Retrieval
- Link: Open Access
- arXiv: 2604.03653
49. EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation
- Link: Open Access
- arXiv: 2604.01421
50. Fast-FoundationStereo: Real-Time Zero-Shot Stereo Matching
- Link: Open Access
- arXiv: 2512.11130
51. GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated Clipping
- Link: Open Access
- arXiv: 2510.22319
52. View-Aware Semantic Alignment for Aerial-Ground Person Re-Identification
- Link: Open Access
- arXiv: 2605.18192
53. ImageRAGTurbo: Towards One-step Text-to-Image Generation with Retrieval-Augmented Diffusion Models
- Link: Open Access
- arXiv: 2602.12640
54. Prompt-Anchored Vision-Text Distillation for Lifelong Person Re-identification
- Link: Open Access
- arXiv: 2605.05027
55. RAGTrack: Language-aware RGBT Tracking with Retrieval-Augmented Generation
- Link: Open Access
- arXiv: 2603.03617
56. RI-Mamba: Rotation-Invariant Mamba for Robust Text-to-Shape Retrieval
- Link: Open Access
- arXiv: 2602.11673
57. MV3DIS: Multi-View Mask Matching via 3D Guides for Zero-Shot 3D Instance Segmentation
- Link: Open Access
- arXiv: 2604.08916
58. Camouflage-aware Image-Text Retrieval via Expert Collaboration
- Link: Open Access
- arXiv: 2604.01251
59. Few-shot Acoustic Synthesis with Multimodal Flow Matching
- Link: Open Access
- arXiv: 2603.19176
60. WISER: Wider Search, Deeper Thinking, and Adaptive Fusion for Training-Free Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2602.23029
61. Thinking in 360deg: Humanoid Visual Search in the Wild
- Link: Open Access
62. Beyond Caption-Based Queries in Video Moment Retrieval
- Link: Open Access
63. Dataset Distillation by Influence Matching
- Link: Open Access
64. SATTC: Structure-Aware Label-Free Test-Time Calibration for Cross-Subject EEG-to-Image Retrieval
- Link: Open Access
- arXiv: 2603.20738
65. EagleNet: Energy-Aware Fine-Grained Relationship Learning Network for Text-Video Retrieval
- Link: Open Access
- arXiv: 2603.25267
66. Object-Generalized Re-Identification: A Step Towards Universal Instance Perception
- Link: Open Access
67. AsymLoc: Towards Asymmetric Feature Matching for Efficient Visual Localization
- Link: Open Access
- arXiv: 2604.09445
68. GeoViS: Geospatially Rewarded Visual Search for Remote Sensing Visual Grounding
- Link: Open Access
- arXiv: 2512.02715
69. MR-RAG: Multimodal Relevance-Aware Retrieval-Augmented Generation for Medical Visual Question Answering
- Link: Open Access
70. RTUA: Reconstruction-residual Based Targeted and Untargeted Attack Against Text-Image Person Re-Identification
- Link: Open Access
71. TextFM: Robust Semi-dense Feature Matching with Language Guidance
- Link: Open Access
72. ReMatch: Boosting Representation through Matching for Multimodal Retrieval
- Link: Open Access
- arXiv: 2511.19278
73. RAG-TP: A General Framework for Vehicle Trajectory Prediction via Retrieval-Augmented Generation
- Link: Open Access
74. Fast Markov Random Field Optimisation for Topologically Noisy 3D Shape Matching
- Link: Open Access
75. Gravitation-Driven Semantic Alignment for Text Video Retrieval
- Link: Open Access
76. WHU-MARS: A Multispectral Aerial-Ground Benchmark Towards Any-Scenario Person Re-Identification
- Link: Open Access
77. Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline Matching
- Link: Open Access
78. M^3KG-RAG: Multi-hop Multimodal Knowledge Graph-enhanced Retrieval-Augmented Generation
- Link: Open Access
79. WeDetect: Fast Open-Vocabulary Object Detection as Retrieval
- Link: Open Access
- arXiv: 2512.12309
80. PoseD-Flow: Versatile and Guided Flow Matching Model of Human Pose
- Link: Open Access
81. Multimodal Distribution Matching for Vision-Language Dataset Distillation
- Link: Open Access
82. Flow Matching for Multimodal Distributions
- Link: Open Access
83. ProgTrack: A Multi-Object Tracking Algorithm with Progressive Matching Strategy
- Link: Open Access
84. MicroFM: Physics-guided Flow Matching for Isotropic Microscopy Reconstruction
- Link: Open Access
85. Flowception: Temporally Expansive Flow Matching for Video Generation
- Link: Open Access
- arXiv: 2512.11438
86. Spatial Retrieval Augmented Autonomous Driving
- Link: Open Access
- arXiv: 2512.06865
87. UniPixie: Unified and Probabilistic 3D Physics Learning via Flow Matching
- Link: Open Access
88. FlowHijack: A Dynamics-Aware Backdoor Attack on Flow-Matching Vision-Language-Action Models
- Link: Open Access
- arXiv: 2604.09651
89. WAM-Flow: Parallel Coarse-to-Fine Motion Planning via Discrete Flow Matching for Autonomous Driving
- Link: Open Access
- arXiv: 2512.06112
90. Unified Number-Free Text-to-Motion Generation Via Flow Matching
- Link: Open Access
- arXiv: 2603.27040
91. PlanaReLoc: Camera Relocalization in 3D Planar Primitives via Region-Based Structure Matching
- Link: Open Access
- arXiv: 2603.20818
92. Compositional Transformation Reasoning for Composed Video Retrieval
- Link: Open Access
93. RAID: Retrieval-Augmented Anomaly Detection
- Link: Open Access
- arXiv: 2602.19611
94. BiFM: Bidirectional Flow Matching for Few-Step Image Editing and Generation
- Link: Open Access
95. Lite Any Stereo: Efficient Zero-Shot Stereo Matching
- Link: Open Access
- arXiv: 2511.16555
96. VRCLIP: Multimodal Canonical Correlation Alignment for CLIP-Driven Vision-Radio Person Re-Identification
- Link: Open Access
97. TIGER: A Unified Framework for Time, Images and Geo-location Retrieval
- Link: Open Access
- arXiv: 2603.24749
98. OpenDPR: Open-Vocabulary Change Detection via Vision-Centric Diffusion-Guided Prototype Retrieval for Remote Sensing Imagery
- Link: Open Access
- arXiv: 2603.27645
99. ConeSep: Cone-based Robust Noise-Unlearning Compositional Network for Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.20358
100. Text-Phase Synergy Network with Dual Priors for Unsupervised Cross-Domain Image Retrieval
- Link: Open Access
- arXiv: 2603.12711
101. AR2-4FV: Anchored Referring and Re-identification for Long-Term Grounding in Fixed-View Videos
- Link: Open Access
- arXiv: 2603.07758
102. FMPose3D: monocular 3D pose estimation via flow matching
- Link: Open Access
- arXiv: 2602.05755
103. CARLoS: Retrieval via Concise Assessment Representation of LoRAs at Scale
- Link: Open Access
- arXiv: 2512.08826
104. GS-ASM: 2DGS-Supervised Active Stereo Matching
- Link: Open Access
105. LeapAlign: Post-training Flow Matching Models at Any Generation Step by Building Two-Step Trajectories
- Link: Open Access
- arXiv: 2604.15311
106. TriSim: Tri-Dimensional Similarity Modeling with Extreme Value Theory for False-Negative Mitigation in Remote Sensing Image-Text Retrieval
- Link: Open Access
107. FAST: Topology-Aware Frequency-Domain Distribution Matching for Coreset Selection
- Link: Open Access
- arXiv: 2511.19476
108. MFEN: Multi-Frequency Expert Network for Visible-Infrared Person Re-ID
- Link: Open Access
109. Pose-guided Enriched Feature Learning for Federated-by-camera Person Re-identification
- Link: Open Access
110. StreamRAG: Enhancing Real-Time Video Understanding with Retrieval Augmentation
- Link: Open Access
111. CAST: Context-Aware Dynamic Latent Space Transformation for Interactive Text-to-Image Retrieval
- Link: Open Access
112. Chain-of-Thought Guided Multi-Modal Object Re-Identification
- Link: Open Access
113. R4: Retrieval-Augmented Reasoning for Vision-Language Models in 4D Spatio-Temporal Space
- Link: Open Access
- arXiv: 2512.15940
114. Evo-Retriever: LLM-Guided Curriculum Evolution with Viewpoint-Pathway Collaboration for Multimodal Document Retrieval
- Link: Open Access
- arXiv: 2603.16455
115. G-MIXER: Geodesic Mixup-based Implicit Semantic Expansion and Explicit Semantic Re-ranking for Zero-Shot Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.14710
116. Beyond Semantic Search: Towards Referential Anchoring in Composed Image Retrieval
- Link: Open Access
- arXiv: 2604.05393
117. What Makes Good Synthetic Training Data for Zero-Shot Stereo Matching?
- Link: Open Access
- arXiv: 2504.16930
118. Love Me, Love My Label: Rethinking the Role of Labels in Prompt Retrieval for Visual In-Context Learning
- Link: Open Access
- arXiv: 2604.03657
119. From Feature Learning to Spectral Basis Learning: A Unifying and Flexible Framework for Efficient and Robust Shape Matching
- Link: Open Access
- arXiv: 2603.23383
120. Red-teaming Retrieval-Augmented Diffusion Models via Poisoning Knowledge Bases
- Link: Open Access
121. SAVE: Speech-Aware Video Representation Learning for Video-Text Retrieval
- Link: Open Access
- arXiv: 2603.08224
122. Stepwise Credit Assignment for GRPO on Flow-Matching Models
- Link: Open Access
- arXiv: 2603.28718
123. Matching Every Pair to Track Every Point: PairFormer for All-Pairs Tracking and Video Trajectory Fields
- Link: Open Access
124. SANER: Switchable Adapter with Non-parametric Enhanced Routing for Person De-Reidentification
- Link: Open Access
125. Quota-Calibrated Fine-Grained Alignment with Context-Aware Marginals for Text-based Person Retrieval
- Link: Open Access
126. Cross-modal Fuzzy Alignment Network for Text-Aerial Person Retrieval and A Large-scale Benchmark
- Link: Open Access
- arXiv: 2603.20721
127. BluRef: Unsupervised Image Deblurring with Dense-Matching References
- Link: Open Access
- arXiv: 2603.14176
128. Follow the Saliency: Supervised Saliency for Retrieval-augmented Dense Video Captioning
- Link: Open Access
- arXiv: 2603.11460
129. Forging a Dynamic Memory: Retrieval-Guided Continual Learning for Generalist Medical Foundation Models
- Link: Open Access
- arXiv: 2512.13072
130. FedBPrompt: Federated Domain Generalization Person Re-Identification via Body Distribution Aware Visual Prompts
- Link: Open Access
- arXiv: 2603.12912
131. Dynamic Magic: Unleashing Restricted Knowledge for Lifelong Person Re-Identification
- Link: Open Access
132. MatchED: Crisp Edge Detection Using End-to-End, Matching-based Supervision
- Link: Open Access
- arXiv: 2602.20689
133. SAG-GNN: Semantic-Aware Guided GNN for Descriptor-Free 2D-3D Matching
- Link: Open Access
134. SearchAD: Large-Scale Rare Image Retrieval Dataset for Autonomous Driving
- Link: Open Access
- arXiv: 2604.08008
135. RenderFlow: Single-Step Neural Rendering via Flow Matching
- Link: Open Access
- arXiv: 2601.06928
136. URICA: A Uniformity Region Affine Identifier Capture Algorithm for Arbitrary Region Retrieval in Pathology Images
- Link: Open Access
137. Towards Highly-Constrained Human Motion Generation with Retrieval-Guided Diffusion Noise Optimization
- Link: Open Access
- arXiv: 2605.08054
138. Phased DMD: Few-step Distribution Matching Distillation via Score Matching within Subintervals
- Link: Open Access
- arXiv: 2510.27684
139. EfficientVPR: Toward Efficient Visual Place Recognition via Scene-Aware Prompt Tuning and Adaptive Feature Enhancement
- Link: Open Access
140. Spatiotemporal Pyramid Flow Matching for Climate Emulation
- Link: Open Access
- arXiv: 2512.02268
141. Factorized Context Aggregation for Robust Cancer Risk Estimation via Soft Re-Ranked Retrieval and Hierarchical Anchors
- Link: Open Access
142. DiT-Distill: Open-Set Fine-Grained Retrieval via Generative Curriculum Knowledge
- Link: Open Access
143. GeodesicNVS: Probability Density Geodesic Flow Matching for Novel View Synthesis
- Link: Open Access
- arXiv: 2603.01010
144. ReCALL: Recalibrating Capability Degradation for MLLM-based Composed Image Retrieval
- Link: Open Access
- arXiv: 2602.01639
145. Learning Straight Flows: Variational Flow Matching for Efficient Generation
- Link: Open Access
- arXiv: 2511.17583
146. RobustVisRAG: Causality-Aware Vision-Based Retrieval-Augmented Generation under Visual Degradations
- Link: Open Access
- arXiv: 2602.22013
147. Multi-Patch Global-to-Local Transformer Architecture For Efficient Flow Matching and Diffusion Model
- Link: Open Access
- arXiv: 2603.26357
148. AdvFM: Lookahead Flow-Matching Velocity-Field Attacks for Imperceptible and Transferable Adversarial Examples
- Link: Open Access
149. Optical Flow Matching: Reframing Optical Flow as Continuous Transport Dynamics
- Link: Open Access
150. MOS: Mitigating Optical-SAR Modality Gap for Cross-Modal Ship Re-Identification
- Link: Open Access
- arXiv: 2512.03404
151. DINO Eats CLIP: Adapting Beyond Knowns for Open-set 3D Object Retrieval
- Link: Open Access
- arXiv: 2604.19432
152. Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching Distillation
- Link: Open Access
- arXiv: 2512.04678
153. 2ndMatch: Finetuning Pruned Diffusion Models via Second-Order Jacobian Matching
- Link: Open Access
- arXiv: 2506.05398
154. Fixed Anchors Are Not Enough: Dynamic Retrieval and Persistent Homology for Dataset Distillation
- Link: Open Access
- arXiv: 2602.24144
155. DialogueVPR: Towards Conversational Visual Place Recognition
- Link: Open Access
156. FlowDIS: Language-Guided Dichotomous Image Segmentation with Flow Matching
- Link: Open Access
- arXiv: 2605.05077
157. FSLoRA: Harmonizing Detection and Re-Identification via Freq-Spatial Low-Rank Adapter for One-Stage Person Search
- Link: Open Access
158. GM-R^2: Generative Matching Learning for Unsupervised Geometric Representation and Registration
- Link: Open Access
159. BIT: Matching-based Bi-directional Interaction Transformation Network for Visible-Infrared Person Re-Identification
- Link: Open Access
- arXiv: 2603.14243
160. Adapting In-context Generation for Enhanced Composed Image Retrieval
- Link: Open Access
161. CodeMMR: Bridging Natural Language, Code, and Image for Unified Retrieval
- Link: Open Access
- arXiv: 2604.15663
162. Self-guided Semantic Inspection for Zero-Shot Composed Image Retrieval
- Link: Open Access
163. Vision-Language Attribute Disentanglement and Reinforcement for Lifelong Person Re-Identification
- Link: Open Access
- arXiv: 2603.19678