- Published on
ICLR 2026 — Computer Vision
Computer Vision
816 papers (0 oral)
Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)
- Link: OpenReview
Neon: Negative Extrapolation From Self-Training Improves Image Generation
- Link: OpenReview
Reasoning as Representation: Rethinking Visual Reinforcement Learning in Image Quality Assessment
- Link: OpenReview
Depth Anything 3: Recovering the Visual Space from Any Views
- Link: OpenReview
Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator
- Link: OpenReview
DepthLM: Metric Depth from Vision Language Models
- Link: OpenReview
Multimodal Aligned Semantic Knowledge for Unpaired Image-text Matching
- Link: OpenReview
Vid-LLM: A Compact Video-based 3D Multimodal LLM with Reconstruction–Reasoning Synergy
- Link: OpenReview
MomaGraph: State-Aware Unified Scene Graphs with Vision-Language Models for Embodied Task Planning
- Link: OpenReview
Locality-aware Parallel Decoding for Efficient Autoregressive Image Generation
- Link: OpenReview
Visual Planning: Let's Think Only with Images
- Link: OpenReview
NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale
- Link: OpenReview
Learning to See Before Seeing: Demystifying LLM Visual Priors from Language Pre-training
- Link: OpenReview
Through the Lens of Contrast: Self-Improving Visual Reasoning in VLMs
- Link: OpenReview
Seeing Through the Brain: New Insights from Decoding Visual Stimuli with fMRI
- Link: OpenReview
WAVE: Learning Unified & Versatile Audio-Visual Embeddings with Multimodal LLM
- Link: OpenReview
Visual symbolic mechanisms: Emergent symbol processing in Vision Language Models
- Link: OpenReview
A Scalable Distributed Framework for Multimodal GigaVoxel Image Registration
- Link: OpenReview
FlashWorld: High-quality 3D Scene Generation within Seconds
- Link: OpenReview
EditVerse: Unifying Image and Video Editing and Generation with In-Context Learning
- Link: OpenReview
: Mixed Autoregressive and Diffusion Transformers for Continuous Image Generation
- Link: OpenReview
10. ARINBEV: Bird's-Eye View Layout Estimation with Conditional Autoregressive Model
- Topics: LLMs & Foundation Models
I-DRUID: Layout to image generation via instance-disentangled representation and unpaired data
- Link: OpenReview
SatDreamer360: Multiview-Consistent Generation of Ground-Level Scenes from Satellite Imagery
- Link: OpenReview
Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization
- Link: OpenReview
FlowAlign: Trajectory-Regularized, Inversion-Free Flow-based Image Editing
- Link: OpenReview
Long-Text-to-Image Generation via Compositional Prompt Decomposition
- Link: OpenReview
Easier Painting Than Thinking: Can Text-to-Image Models Set the Stage, but Not Direct the Play?
- Link: OpenReview
Turbo-DDCM: Fast and Flexible Zero-Shot Diffusion-Based Image Compression
- Link: OpenReview
Features Emerge as Discrete States: The First Application of SAEs to 3D Representations
- Link: OpenReview
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks
- Link: OpenReview
LearnIR: Learnable Posterior Sampling for Real-World Image Restoration
- Link: OpenReview
Learning Unified Representation of 3D Gaussian Splatting
- Link: OpenReview
CIAR: Interval-based Collaborative Decoding for Image Generation Acceleration
- Link: OpenReview
Analyzing the Training Dynamics of Image Restoration Transformers: A Revisit to Layer Normalization
- Link: OpenReview
ChronoEdit: Towards Temporal Reasoning for In-Context Image Editing and World Simulation
- Link: OpenReview
Pixel-Perfect Puppetry: Precision-Guided Enhancement for Face Image and Video Editing
- Link: OpenReview
Cross-ControlNet: Training-Free Fusion of Multiple Conditions for Text-to-Image Generation
- Link: OpenReview
Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-Distillation
- Link: OpenReview
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
- Link: OpenReview
SurfSplat: Conquering Feedforward 2D Gaussian Splatting with Surface Continuity Priors
- Link: OpenReview
Dens3R: A Foundation Model for 3D Geometry Prediction
- Link: OpenReview
AssetFormer: Modular 3D Assets Generation with Autoregressive Transformer
- Link: OpenReview
UniSplat: Unified Spatio-Temporal Fusion via 3D Latent Scaffolds for Dynamic Driving Scene Reconstruction
- Link: OpenReview
: Permutation-Equivariant Visual Geometry Learning
- Link: OpenReview
Mono4DGS-HDR: High Dynamic Range 4D Gaussian Splatting from Alternating-exposure Monocular Videos
- Link: OpenReview
VideoJudge: Bootstrapping Enables Scalable Supervision of MLLM-as-a-Judge for Video Understanding
- Link: OpenReview
GuirlVG: Incentivize GUI Visual Grounding via Empirical Exploration on Reinforcement Learning
- Link: OpenReview
AdaReasoner: Dynamic Tool Orchestration for Iterative Visual Reasoning
- Link: OpenReview
MergeMix: A Unified Augmentation Paradigm for Visual and Multi-Modal Understanding
- Link: OpenReview
TimeSearch-R: Adaptive Temporal Search for Long-Form Video Understanding via Self-Verification Reinforcement Learning
- Link: OpenReview
MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
- Link: OpenReview
GHOST: Hallucination-Inducing Image Generation for Multimodal LLMs
- Link: OpenReview
Provably Accelerated Imaging with Restarted Inertia and Score-based Image Priors
- Link: OpenReview
Discrete Diffusion for Reflective Vision-Language-Action Models in Autonomous Driving
- Link: OpenReview
VisionReasoner: Unified Reasoning-Integrated Visual Perception via Reinforcement Learning
- Link: OpenReview
Unified Multi-Modal Interactive and Reactive 3D Motion Generation via Rectified Flow
- Link: OpenReview
Revisual-R1: Advancing Multimodal Reasoning From Optimized Cold Start to Staged Reinforcement Learning
- Link: OpenReview
MergeTune: Continued Fine-Tuning of Vision-Language Models
- Link: OpenReview
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
- Link: OpenReview
Self-Evolving Vision-Language Models for Image Quality Assessment via Voting and Ranking
- Link: OpenReview
VisualPRM400K: An Effective Dataset for Training Multimodal Process Reward Models
- Link: OpenReview
QLIP: A Dynamic Quadtree Vision Prior Enhances MLLM Performance Without Retraining
- Link: OpenReview
UniFlow: A Unified Pixel Flow Tokenizer for Visual Understanding and Generation
- Link: OpenReview
SPRQ: Static Priority-based Rectifier Routing Quantization for Image Super-Resolution
- Link: OpenReview
Enhancing Multi-Image Understanding through Delimiter Token Scaling
- Link: OpenReview
PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
- Link: OpenReview
Linear Mechanisms for Spatiotemporal Reasoning in Vision Language Models
- Link: OpenReview
A Comprehensive Information-Decomposition Analysis of Large Vision-Language Models
- Link: OpenReview
Visual Self-Refine: A Pixel-Guided Paradigm for Accurate Chart Parsing
- Link: OpenReview
Exploring the Potential of Encoder-free Architectures in 3D LMMs
- Link: OpenReview
ImageDoctor: Diagnosing Text-to-Image Generation via Grounded Image Reasoning
- Link: OpenReview
You Point, I Learn: Online Adaptation of Interactive Segmentation Models for Handling Distribution Shifts in Medical Imaging
- Link: OpenReview
Detective SAM: Adaptive AI-Image Forgery Localization
- Link: OpenReview
Unveiling Perceptual Artifacts: A Fine-Grained Benchmark for Interpretable AI-Generated Image Detection
- Link: OpenReview
Towards Reliable Detection of Empty Space: Conditional Marked Point Processes for Object Detection
- Link: OpenReview
DVLA-RL: Dual-Level Vision–Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
- Link: OpenReview
Vulcan: Crafting Compact Class-Specific Vision Transformers For Edge Intelligence
- Link: OpenReview
OmniText: A Training-Free Generalist for Controllable Text-Image Manipulation
- Link: OpenReview
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
- Link: OpenReview
gen2seg: Generative Models Enable Generalizable Instance Segmentation
- Link: OpenReview
Content-Aware Mamba for Learned Image Compression
- Link: OpenReview
SceneTransporter: Optimal Transport-Guided Compositional Latent Diffusion for Single-Image Structured 3D Scene Generation
- Link: OpenReview
From Sparse to Dense: Spatio-Temporal Fusion for Multi-View 3D Human Pose Estimation with DenseWarper
- Link: OpenReview
Rethinking Expressivity and Degradation-Awareness in Attention for All-in-One Blind Image Restoration
- Link: OpenReview
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
- Link: OpenReview
Follow-Your-Preference: Towards Preference-Aligned Image Inpainting
- Link: OpenReview
Motion-R1: Enhancing Motion Generation with Decomposed Chain-of-Thought and RL Binding
- Link: OpenReview
Pixel-Level Residual Diffusion Transformer: Scalable 3D CT Volume Generation
- Link: OpenReview
Landscape of Thoughts: Visualizing the Reasoning Process of Large Language Models
- Link: OpenReview
Continual Unlearning for Text-to-Image Diffusion Models: A Regularization Perspective
- Link: OpenReview
VisCoder2: Building Multi-Language Visualization Coding Agents
- Link: OpenReview
Vision Language Models are Biased
- Link: OpenReview
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
- Link: OpenReview
Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
- Link: OpenReview
Assessing Robustness via Score-Based Adversarial Image Generation
- Link: OpenReview
414. Towards Anomaly-Aware Pre-Training and Fine-Tuning for Graph Anomaly Detection
- Topics: LLMs & Foundation Models, Graph Neural Networks, Graphs & Combinatorial
GeoDiv: Framework for Measuring Geographical Diversity in Text-to-Image Models
- Link: OpenReview
STEDiff: Revealing the Spatial and Temporal Redundancy of Backdoor Attacks in Text-to-Image Diffusion Models
- Link: OpenReview
VEAttack: Downstream-agnostic Vision Encoder Attack against Large Vision Language Models
- Link: OpenReview
Enhancing Image-Conditional Coverage in Segmentation: Adaptive Thresholding via Differentiable Miscoverage Loss
- Link: OpenReview
Do Vision-Language Models Respect Contextual Integrity in Location Disclosure?
- Link: OpenReview
CGSA: Class-Guided Slot-Aware Adaptation for Source-Free Object Detection
- Link: OpenReview
Test-time Domain Generalization for Image Super-resolution
- Link: OpenReview
Dual-Kernel Adapter: Expanding Spatial Horizons for Data-Constrained Medical Image Analysis
- Link: OpenReview
Reading Images Like Texts: Sequential Image Understanding in Vision-Language Models
- Link: OpenReview
CryoSplat: Gaussian Splatting for Cryo-EM Homogeneous Reconstruction
- Link: OpenReview
Synthesizing High-Quality Visual Question Answering from Medical Documents with Generator-Verifier LMMs
- Link: OpenReview
A Structured, Tagged, and Localized Visual Question Answering Dataset with Full Sentence Answers and Scene Graphs for Chest X-ray Images
- Link: OpenReview
FETAL-GAUGE: A BENCHMARK FOR ASSESSING VISION-LANGUAGE MODELS IN FETAL ULTRASOUND
- Link: OpenReview
villa-X: Enhancing Latent Action Modeling in Vision-Language-Action Models
- Link: OpenReview
Vision-Language-Action Instruction Tuning: From Understanding to Manipulation
- Link: OpenReview
PEERING INTO THE UNKNOWN: ACTIVE VIEW SELECTION WITH NEURAL UNCERTAINTY MAPS FOR 3D RECONSTRUCTION
- Link: OpenReview
RAP: 3D Rasterization Augmented End-to-End Planning
- Link: OpenReview
Interleave-VLA: Enhancing Robot Manipulation with Image-Text Interleaved Instructions
- Link: OpenReview
QVLA: Not All Channels Are Equal in Vision-Language-Action Model's Quantization
- Link: OpenReview
Unified Vision-Language-Action Model
- Link: OpenReview
Vlaser: Vision-Language-Action Model with Synergistic Embodied Reasoning
- Link: OpenReview
MemoryVLA: Perceptual-Cognitive Memory in Vision-Language-Action Models for Robotic Manipulation
- Link: OpenReview
Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow
- Link: OpenReview
Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Method
- Link: OpenReview
OmniNav: A Unified Framework for Prospective Exploration and Visual-Language Navigation
- Link: OpenReview
ORCaS: Unsupervised Depth Completion via Occluded Region Completion as Supervision
- Link: OpenReview
PerfGuard: A Performance-Aware Agent for Visual Content Generation
- Link: OpenReview
Taming Hierarchical Image Coding Optimization: A Spectral Regularization Perspective
- Link: OpenReview
PrismAudio: Decomposed Chain-of-Thought and Multi-dimensional Rewards for Video-to-Audio Generation
- Link: OpenReview
PTQ4ARVG: Post-Training Quantization for AutoRegressive Visual Generation Models
- Link: OpenReview
Bridging Degradation Discrimination and Generation for Universal Image Restoration
- Link: OpenReview
SpatiaLab: Can Vision–Language Models Perform Spatial Reasoning in the Wild?
- Link: OpenReview
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark
- Link: OpenReview
SELF-HARMONY: LEARNING TO HARMONIZE SELF-SUPERVISION AND SELF-PLAY IN TEST-TIME REINFORCEMENT LEARNING
- Link: OpenReview
Efficient Audio-Visual Speech Separation with Discrete Lip Semantics and Multi-Scale Global-Local Attention
- Link: OpenReview
Stroke3D: Lifting 2D strokes into rigged 3D model via latent diffusion models
- Link: OpenReview
Compose Your Policies! Improving Diffusion-based or Flow-based Robot Policies via Test-time Distribution-level Composition
- Link: OpenReview
Uncovering Semantic Selectivity of Latent Groups in Higher Visual Cortex with Mutual Information-Guided Diffusion
- Link: OpenReview
Stretching Beyond the Obvious: A Gradient-Free Framework to Unveil the Hidden Landscape of Visual Invariance
- Link: OpenReview
CoPRS: Learning Positional Prior from Chain-of-Thought for Reasoning Segmentation
- Link: OpenReview
Autoregressive Visual Decoding from EEG Signals
- Link: OpenReview
Model-Guided Microstimulation Steers Primate Visual Behavior
- Link: OpenReview
SEED: Towards More Accurate Semantic Evaluation for Visual Brain Decoding
- Link: OpenReview
Learning Brain Representation with Hierarchical Visual Embeddings
- Link: OpenReview
Inducing Dyslexia in Vision Language Models
- Link: OpenReview
Uncertainty-Aware 3D Reconstruction for Dynamic Underwater Scenes
- Link: OpenReview
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models
- Link: OpenReview
Enhancing Vision-Language Model with Unmasked Token Alignment
- Link: OpenReview
821. AB-UPT: Scaling Neural CFD Surrogates for High- Fidelity Automotive Aerodynamics Simulations via Anchored- Branched Universal Physics Transformers
- Topics: LLMs & Foundation Models
822. Seek-CAD: A Self-refined Generative Modeling for 3D Parametric CAD Using Local Inference via DeepSeek
- Topics: Diffusion Models & Generative AI, Computer Vision, Efficiency & Compression
Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation
- Link: OpenReview
FastVGGT: Fast Visual Geometry Transformer
- Link: OpenReview
Curvature-Guided Task Synergy for Skeleton based Temporal Action Segmentation
- Link: OpenReview
WithAnyone: Toward Controllable and ID Consistent Image Generation
- Link: OpenReview
SIGMA-Gen: Structure and Identity Guided Multi-Subject Assembly for Image Generation
- Link: OpenReview
Follow-Your-Shape: Shape-Aware Image Editing via Trajectory-Guided Region Control
- Link: OpenReview
VMDiff: Visual Mixing Diffusion for Limitless Cross-Object Synthesis
- Link: OpenReview
Autoregressive Image Generation with Randomized Parallel Decoding
- Link: OpenReview
Culture in Action: Evaluating Text-to-Image Models through Social Activities
- Link: OpenReview
AutoDV: An End-to-End Deep Learning Model for High-Dimensional Data Visualization
- Link: OpenReview
Weight Space Representation Learning on Diverse NeRF Architectures
- Link: OpenReview
VQ-Transplant: Efficient VQ-Module Integration for Pre-trained Visual Tokenizers
- Link: OpenReview
ImageRAG: Dynamic Image Retrieval for Reference-Guided Image Generation
- Link: OpenReview
CTRL&SHIFT: High-quality Geometry-Aware Object Manipulation in Visual Generation
- Link: OpenReview
DeLeaker: Dynamic Inference-Time Reweighting For Semantic Leakage Mitigation in Text-to-Image Models
- Link: OpenReview
SSG: Scaled Spatial Guidance for Multi-Scale Visual Autoregressive Generation
- Link: OpenReview
FreeAdapt: Unleashing Diffusion Priors for Ultra-High-Definition Image Restoration
- Link: OpenReview
LogiStory: A Logic-Aware Framework for Multi-Image Story Visualization
- Link: OpenReview
Dragging with Geometry: From Pixels to Geometry-Guided Image Editing
- Link: OpenReview
Does FLUX Already Know How to Perform Physically Plausible Image Composition?
- Link: OpenReview
SuperF: Neural Implicit Fields for Multi-Image Super-Resolution
- Link: OpenReview
Escaping Low-Rank Traps: Interpretable Visual Concept Learning via Implicit Vector Quantization
- Link: OpenReview
Learnable Sparsity for Vision Generative Models
- Link: OpenReview
Spatially Informed Autoencoders for Interpretable Visual Representation Learning
- Link: OpenReview
Meta-Adaptive Prompt Distillation for Few-Shot Visual Question Answering
- Link: OpenReview
NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting
- Link: OpenReview
One2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single Image
- Link: OpenReview
TTT3R: 3D Reconstruction as Test-Time Training
- Link: OpenReview
FullPart: Generating each 3D Part at Full Resolution
- Link: OpenReview
Interference-Isolated Elastic Weight Consolidation and Knowledge Calibration for Incremental Object Detection
- Link: OpenReview
Variation-aware Flexible 3D Gaussian Editing
- Link: OpenReview
Augmented Radiance Field: A General Framework for Enhanced Gaussian Splatting
- Link: OpenReview
Stylos: Multi-View 3D Stylization with Single-Forward Gaussian Splatting
- Link: OpenReview
Signal Structure-Aware Gaussian Splatting for Large-Scale Scene Reconstruction
- Link: OpenReview
ULTRA-360: Unconstrained Dataset for Large-scale Temporal 3D Reconstruction across Altitudes and Omnidirectional Views
- Link: OpenReview
InSight-o3: Empowering Multimodal Foundation Models with Generalized Visual Search
- Link: OpenReview
Hallucination-aware Intermediate Representation Edit in Large Vision-Language Models
- Link: OpenReview
ViPER: Empowering the Self-Evolution of Visual Perception Abilities in Vision-Language Models
- Link: OpenReview
Memory-Free Continual Learning with Null Space Adaptation for Zero-Shot Vision-Language Models
- Link: OpenReview
AVERE: Improving Audiovisual Emotion Reasoning with Preference Optimization
- Link: OpenReview
Seeing Across Views: Benchmarking Spatial Reasoning of Vision-Language Models in Robotic Scenes
- Link: OpenReview
Omni-IML: Towards Unified Interpretable Image Manipulation Localization
- Link: OpenReview
RewardMap: Tackling Sparse Rewards in Fine-grained Visual Reasoning via Multi-Stage Reinforcement Learning
- Link: OpenReview
Prune Redundancy, Preserve Essence: Vision Token Compression in VLMs via Synergistic Importance-Diversity
- Link: OpenReview
GTR-Bench: Evaluating Geo-Temporal Reasoning in Vision-Language Models
- Link: OpenReview
VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual Reasoning
- Link: OpenReview
GranViT: A Fine-Grained Vision Model For Autoregressive Multimodal Large Language Models
- Link: OpenReview
Spatial-DISE: A Unified Benchmark for Evaluating Spatial Reasoning in Vision-Language Models
- Link: OpenReview
Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language Translation
- Link: OpenReview
MoRA: Missing Modality Low-Rank Adaptation for Visual Recognition
- Link: OpenReview
Transductive Visual Programming: Evolving Tool Libraries from Experience for Spatial Reasoning
- Link: OpenReview
How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks
- Link: OpenReview
CapRL: Stimulating Dense Image Caption Capabilities via Reinforcement Learning
- Link: OpenReview
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
- Link: OpenReview
No Labels, No Problem: Training Visual Reasoners with Multimodal Verifiers
- Link: OpenReview
Composition-Grounded Data Synthesis for Visual Reasoning
- Link: OpenReview
EgoNight: Towards Egocentric Vision Understanding at Night with a Challenging Benchmark
- Link: OpenReview
GaussianFusion: Unified 3D Gaussian Representation for Multi-Modal Fusion Perception
- Link: OpenReview
From Narrow to Panoramic Vision: Attention-Guided Cold-Start Reshapes Multimodal Reasoning
- Link: OpenReview
HiDrop: Hierarchical Vision Token Reduction in MLLMs via Late Injection, Concave Pyramid Pruning, and Early Exit
- Link: OpenReview
ScaleCap: Scalable Image Captioning via Dual-Modality Debiasing
- Link: OpenReview
Hierarchical Prototype Learning for Semantic Segmentation
- Link: OpenReview
QPrompt-R1: Real-Time Reasoning for Domain-Generalized Semantic Segmentation via Group-Relative Query Alignment
- Link: OpenReview
CARL: Camera-Agnostic Representation Learning for Spectral Image Analysis
- Link: OpenReview
ELViS: Efficient Visual Similarity from Local Descriptors that Generalizes Across Domains
- Link: OpenReview
DeCo-DETR: Decoupled Cognition DETR for efficient Open-Vocabulary Object Detection
- Link: OpenReview
SpikeStereoNet: A Brain-Inspired Framework for Stereo Depth Estimation from Spike Streams
- Link: OpenReview
Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose Estimation
- Link: OpenReview
Pose-RFT: Aligning MLLMs for 3D Pose Generation via Hybrid Action Reinforcement Fine-Tuning
- Link: OpenReview
Benchmarking Open-ended Segmentation
- Link: OpenReview
JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- Link: OpenReview
GenCape: Structure-Inductive Generative Modeling for Category-Agnostic Pose Estimation
- Link: OpenReview
Inlier-Centric Post-Training Quantization for Object Detection Models
- Link: OpenReview
PointRePar : SpatioTemporal Point Relation Parsing for Robust Category-Unified 3D Tracking
- Link: OpenReview
FideDiff: Efficient Diffusion Model for High-Fidelity Image Motion Deblurring
- Link: OpenReview
Splat and Distill: Augmenting Teachers with Feed-Forward 3D Reconstruction For 3D-Aware Distillation
- Link: OpenReview
Unsupervised Representation Learning for 3D Mesh Parameterization with Semantic and Visibility Objectives
- Link: OpenReview
Consistent Text-to-Image Generation via Scene De-Contextualization
- Link: OpenReview
Aligning Visual Foundation Encoders to Tokenizers for Diffusion Models
- Link: OpenReview
DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies
- Link: OpenReview
HoloPart: Generative 3D Part Amodal Segmentation
- Link: OpenReview
From Sure" to Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons
- Link: OpenReview
Enabling True Global Perception in State Space Models for Visual Tasks
- Link: OpenReview
Seeing Through Deception: Uncovering Misleading Creator Intent in Multimodal News with Vision-Language Models
- Link: OpenReview
Preserve and Sculpt: Manifold-Aligned Fine-tuning of Vision-Language Models for Few-Shot Learning
- Link: OpenReview
Decomposition of Concept-Level Rules in Visual Scenes
- Link: OpenReview
Medical thinking with multiple images
- Link: OpenReview
Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMs
- Link: OpenReview
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Link: OpenReview
GAGA: Gaussianity-Aware Gaussian Approximation for Efficient 3D Molecular Generation
- Link: OpenReview
Exploring Interpretability for Visual Prompt Tuning with Cross-layer Concepts
- Link: OpenReview
Improving 2D Diffusion Models for 3D Medical Imaging with Inter‑Slice Consistent Stochasticity
- Link: OpenReview
VLBiMan: Vision-Language Anchored One-Shot Demonstration Enables Generalizable Bimanual Robotic Manipulation
- Link: OpenReview
SpikePingpong: Spike Vision-based Fast-Slow Pingpong Robot System
- Link: OpenReview
Verifier-free Test-Time Sampling for Vision-Language-Action Models
- Link: OpenReview
OpenFly: A COMPREHENSIVE PLATFORM FOR AERIAL VISION-LANGUAGE NAVIGATION
- Link: OpenReview
Hybrid Training for Vision-Language-Action Models
- Link: OpenReview
HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action Model
- Link: OpenReview
Unified Diffusion VLA: Vision-Language-Action Model via Joint Discrete Denosing Diffusion Process
- Link: OpenReview
VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding Models
- Link: OpenReview
CogFlow: Bridging Perception and Reasoning through Knowledge Internalization for Visual Mathematical Problem Solving
- Link: OpenReview
Entropy-Monitored Kernelized Token Distillation for Audio-Visual Compression
- Link: OpenReview
X-VLA: Soft-Prompted Transformer as Scalable Cross-Embodiment Vision-Language-Action Model
- Link: OpenReview
ST-SimDiff: Balancing Spatiotemporal Similarity and Difference for Efficient Video Understanding with MLLMs
- Link: OpenReview
Thyme: Think Beyond Images
- Link: OpenReview
Spatially Guided Training for Vision-Language-Action Model
- Link: OpenReview
SenseFlow: Scaling Distribution Matching for Flow-based Text-to-Image Distillation
- Link: OpenReview
Decoding Dynamic Visual Experience from Calcium Imaging via Cell-Pattern-Aware Pretraining
- Link: OpenReview
CloDS: Visual-Only Unsupervised Cloth Dynamics Learning in Unknown Conditions
- Link: OpenReview
LaVCa: LLM-assisted Visual Cortex Captioning
- Link: OpenReview
We-Math 2.0: A Versatile MathBook System for Incentivizing Visual Mathematical Reasoning
- Link: OpenReview
Unified Vision–Language Modeling via Concept Space Alignment
- Link: OpenReview
Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models
- Link: OpenReview
AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation
- Link: OpenReview
Nano3D: A Training-Free Approach for Efficient 3D Editing Without Masks
- Link: OpenReview
PAT3D: Physics-Augmented Text-to-3D Scene Generation
- Link: OpenReview
Hierarchical Value-Decomposed Offline Reinforcement Learning for Whole-Body Control
- Link: OpenReview
Frequency-aware Dynamic Gaussian Splatting
- Link: OpenReview
SceneCOT: Eliciting Grounded Chain-of-Thought Reasoning in 3D Scenes
- Link: OpenReview
MANZANO: A Simple and Scalable Unified Multimodal Model with a Hybrid Vision Tokenizer
- Link: OpenReview
CARPRT: Class-Aware Zero-Shot Prompt Reweighting for Vision-Language Model
- Link: OpenReview
HFSTI-Net: Hierarchical Frequency-spatial-temporal Interactions for Video Polyp Segmentation
- Link: OpenReview
Spatial Structure and Selective Text Jointly Facilitate Image Clustering
- Link: OpenReview
Durian: Dual Reference Image-Guided Portrait Animation with Attribute Transfer
- Link: OpenReview
ToProVAR: Efficient Visual Autoregressive Modeling via Tri-Dimensional Entropy-Aware Semantic Analysis and Sparsity Optimization
- Link: OpenReview
Data Provenance for Image Auto-Regressive Generation
- Link: OpenReview
Token-Efficient Item Representation via Images for LLM Recommender Systems
- Link: OpenReview
TINKER: Diffusion's Gift to 3D--Multi-View Consistent Editing From Sparse Inputs without Per-Scene Optimization
- Link: OpenReview
Everything in Its Place: Benchmarking Spatial Intelligence of Text-to-Image Models
- Link: OpenReview
EchoGen: Generating Visual Echoes in Any Scene via Feed-Forward Subject-Driven Auto-Regressive Model
- Link: OpenReview
SiNGER: A Clearer Voice Distills Vision Transformers Further
- Link: OpenReview
Purrception: Variational Flow Matching for Vector-Quantized Image Generation
- Link: OpenReview
Localized Concept Erasure in Text-to-Image Diffusion Models via High-Level Representation Misdirection
- Link: OpenReview
Locality-Attending Vision Transformer
- Link: OpenReview
Generate Any Scene: Scene Graph Driven Data Synthesis for Visual Generation Training
- Link: OpenReview
PQGAN: Product-Quantised Image Representation for High-Quality Image Synthesis
- Link: OpenReview
3D Scene Prompting for Scene-Consistent Camera-Controllable Video Generation
- Link: OpenReview
BAR: Refactor the Basis of Autoregressive Visual Generation
- Link: OpenReview
CLoD-GS: Continuous Level-of-Detail via 3D Gaussian Splatting
- Link: OpenReview
Active Learning of 3D Gaussian Splatting with Consistent Region Partition and Robust Pose Estimation
- Link: OpenReview
A Step to Decouple Optimization in 3DGS
- Link: OpenReview
VITA: Zero-Shot Value Functions via Test-Time Adaptation of Vision–Language Models
- Link: OpenReview
Uncertainty Matters in Dynamic Gaussian Splatting for Monocular 4D Reconstruction
- Link: OpenReview
CylinderSplat: 3D Gaussian Splatting with Cylindrical Triplanes for Panoramic Novel View Synthesis
- Link: OpenReview
G4Splat: Geometry-Guided Gaussian Splatting with Generative Prior
- Link: OpenReview
SkyEvents: A Large-Scale Event-enhanced UAV Dataset for Robust 3D Scene Reconstruction
- Link: OpenReview
PDGS: Part-Level Decoupling and Continuous Deformation of Articulated Objects via Gaussian Splatting
- Link: OpenReview
NOVA3R: Non-pixel-aligned Visual Transformer for Amodal 3D Reconstruction
- Link: OpenReview
ETGS: Explicit Thermodynamics Gaussian Splatting for Dynamic Thermal Reconstruction
- Link: OpenReview
Prompt-Robust Vision-Language Models via Meta-Finetuning
- Link: OpenReview
Implicit 4D Gaussian Splatting for Fast Motion with Large Inter-Frame Displacements
- Link: OpenReview
Condition Matters in Full-head 3D GANs
- Link: OpenReview
All That Glitters Is Not Gold: Key-Secured 3D Secrets within 3D Gaussian Splatting
- Link: OpenReview
COMPASS: Robust Feature Conformal Prediction for Medical Segmentation Metrics
- Link: OpenReview
UFO-4D: Unposed Feedforward 4D Reconstruction from Two Images
- Link: OpenReview
DGS: Depth-and-Density Guided Gaussian Splatting for Stable and Accurate Sparse-View Reconstruction
- Link: OpenReview
Post-hoc Probabilistic Vision-Language Models
- Link: OpenReview
Hyden: A Hybrid Dual-Path Encoder for Monocular Geometry of High-resolution Images
- Link: OpenReview
Secondary Motion-Aware 3D Clothed Gaussian Avatars from Monocular Videos
- Link: OpenReview
Kaleidoscope: In-language Exams for Massively Multilingual Vision Evaluation
- Link: OpenReview
CoMem: Compositional Concept-Graph Memory for Vision–Language Adaptation
- Link: OpenReview
Delving into Spectral Clustering with Vision-Language Representations
- Link: OpenReview
Efficient Test-Time Scaling for Small Vision-Language Models
- Link: OpenReview
SURGE: Surprise-Guided Token Reduction for Efficient Video Understanding with VLMs
- Link: OpenReview
Beyond Text-to-Image: Liberating Generation with a Unified Discrete Diffusion Model
- Link: OpenReview
Omni-View: Unlocking How Generation Facilitates Understanding in Unified 3D Model based on Multiview images
- Link: OpenReview
Automatic Image-Level Morphological Trait Annotation for Organismal Images
- Link: OpenReview
CityLens: Evaluating Large Vision-Language Models for Urban Socioeconomic Sensing
- Link: OpenReview
Long-tailed Test-Time Adaptation for Vision-Language Models
- Link: OpenReview
Imitating the Truth: Attention-aware Truth-Guided Enhancement for Hallucination Mitigation in Large Vision-Language Models
- Link: OpenReview
EdgeCape: Edge Weight Prediction For Category-Agnostic Pose Estimation
- Link: OpenReview
Boosting Medical Visual Understanding From Multi-Granular Language Learning
- Link: OpenReview
pFedMMA: Personalized Federated Fine-Tuning with Multi-Modal Adapter for Vision-Language Models
- Link: OpenReview
Mini-o3: Scaling Up Reasoning Patterns and Interaction Turns for Visual Search
- Link: OpenReview
Mordal: Automated Pretrained Model Selection for Vision Language Models
- Link: OpenReview
Reversible Primitive–Composition Alignment for Continual Vision–Language Learning
- Link: OpenReview
Efficient Discriminative Joint Encoders for Large Scale Vision-Language Reranking
- Link: OpenReview
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence
- Link: OpenReview
DaVinci: Reinforcing Visual-Structural Syntax in MLLMs for Generalized Scientific Diagram Parsing
- Link: OpenReview
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- Link: OpenReview
To View Transform or Not to View Transform: NeRF-based Pre-training Perspective
- Link: OpenReview
Mitigating Hallucination in Vision-Language Model with Depth and Spatial-aware Key-Value Refinement
- Link: OpenReview
Simulation to Rules: A Dual-VLM Framework for Formal Visual Planning
- Link: OpenReview
Uni-CoT: Towards Unified Chain-of-Thought Reasoning Across Text and Vision
- Link: OpenReview
MARC: Memory-Augmented RL Token Compression for Efficient Video Understanding
- Link: OpenReview
Query-Guided Spatial–Temporal–Frequency Interaction for Music Audio–Visual Question Answering
- Link: OpenReview
AgilePruner: An Empirical Study of Attention and Diversity for Adaptive Visual Token Pruning in Large Vision-Language Models
- Link: OpenReview
K-Prism: A Knowledge-Guided and Prompt Integrated Universal Medical Image Segmentation Model
- Link: OpenReview
Cross-Timestep: 3D Diffusion Model with Trans-temporal Memory LSTM and Adaptive Priori Decoding Strategy for Medical Segmentation
- Link: OpenReview
S3OD: Towards Generalizable Salient Object Detection with Synthetic Data
- Link: OpenReview
WOW-Seg: A Word-free Open World Segmentation Model
- Link: OpenReview
ASCIIEval: Benchmarking Models' Visual Perception in Text Strings via ASCII Art
- Link: OpenReview
Self-Refining Vision Language Model for Robotic Failure Detection and Reasoning
- Link: OpenReview
FSOD-VFM: Few-Shot Object Detection with Vision Foundation Models and Graph Diffusion
- Link: OpenReview
DiVE-k: DIFFERENTIAL VISUAL REASONING FOR FINE-GRAINED IMAGE RECOGNITION
- Link: OpenReview
VisionLaw: Inferring Interpretable Intrinsic Dynamics from Visual Observations via Bilevel Optimization
- Link: OpenReview
Customizing Visual Emotion Evaluation for MLLMs: An Open-vocabulary, Multifaceted, and Scalable Approach
- Link: OpenReview
3DSMT: A Hybrid Spiking Mamba-Transformer for Point Cloud Analysis
- Link: OpenReview
Point-Focused Attention Meets Context-Scan State Space: Robust Biological Visual Perception for Point Cloud Representation
- Link: OpenReview
Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
- Link: OpenReview
Divergence-Free Neural Networks with Application to Image Denoising
- Link: OpenReview
OD: Optimization-free Dataset Distillation for Object Detection
- Link: OpenReview
BioTamperNet: Affinity-Guided State-Space Model Detecting Tampered Biomedical Images
- Link: OpenReview
SPWOOD: Sparse Partial Weakly-Supervised Oriented Object Detection
- Link: OpenReview
Detecting Misbehaviors of Large Vision-Language Models by Evidential Uncertainty Quantification
- Link: OpenReview
Image Can Bring Your Memory Back: A Novel Multi-Modal Guided Attack against Image Generation Model Unlearning
- Link: OpenReview
Sample Reward Soups: Query-efficient Multi-Reward Guidance for Text-to-Image Diffusion Models
- Link: OpenReview
Diverse Text-to-Image Generation via Contrastive Noise Optimization
- Link: OpenReview
Scaling Reasoning Hop Exposes Weaknesses: Demystifying and Improving Hop Generalization in Large Language Models
- Link: OpenReview
VPI-Bench: Visual Prompt Injection Attacks for Computer-Use Agents
- Link: OpenReview
CompMarkGS: Robust Watermarking for Compressed 3D Gaussian Splatting
- Link: OpenReview
Bilateral Information-aware Test-time Adaptation for Vision-Language Models
- Link: OpenReview
Compose and Fuse: Revisiting the Foundational Bottlenecks in Multimodal Reasoning
- Link: OpenReview
M3CoTBench: Benchmark Chain-of-Thought of MLLMs in Medical Image Understanding
- Link: OpenReview
MAVEN: A Mesh-Aware Volumetric Encoding Network for Simulating 3D Flexible Deformation
- Link: OpenReview
ArtVIP: Articulated Digital Assets of Visual Realism, Modular Interaction, and Physical Fidelity for Robot Learning
- Link: OpenReview
OneTwoVLA: A Unified Vision-Language-Action Model with Adaptive Reasoning
- Link: OpenReview
Flash-Mono: Feed-Forward Accelerated Gaussian Splatting Monocular SLAM
- Link: OpenReview
MATA: A Trainable Hierarchical Automaton System for Multi-Agent Visual Reasoning
- Link: OpenReview
TwinVLA: Data-Efficient Bimanual Manipulation with Twin Single-Arm Vision-Language-Action Models
- Link: OpenReview
UniHM: Unified Dexterous Hand Manipulation with Vision Language Model
- Link: OpenReview
D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI
- Link: OpenReview
JanusVLN: Decoupling Semantics and Spatiality with Dual Implicit Memory for Vision-Language Navigation
- Link: OpenReview
HAMLET: Switch Your Vision-Language-Action Model into a History-Aware Policy
- Link: OpenReview
VITA: Vision-to-Action Flow Matching Policy
- Link: OpenReview
AutoFly: Vision-Language-Action Model for UAV Autonomous Navigation in the Wild
- Link: OpenReview
PixelVLA: Advancing Pixel-level Understanding in Vision-Language-Action Model
- Link: OpenReview
Unsupervised Learning of Efficient Exploration: Pre-training Adaptive Policies via Self-Imposed Goals
- Link: OpenReview
ViPO: Visual Preference Optimization at Scale
- Link: OpenReview
WMPO: World Model-based Policy Optimization for Vision-Language-Action Models
- Link: OpenReview
Latent Visual Reasoning
- Link: OpenReview
DispViT: Direct Stereo Disparity Regression with a Single-Stream Vision Transformer
- Link: OpenReview
Image Quality Assessment for Embodied AI
- Link: OpenReview
PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation Models
- Link: OpenReview
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
- Link: OpenReview
WavePolyp: Video Polyp Segmentation via Hierarchical Wavelet-Based Feature Aggregation and Inter-Frame Divergence Perception
- Link: OpenReview
Tug-of-War No More: Harmonizing Accuracy and Robustness in Vision-Language Models via Stability-Aware Task Vector Merging
- Link: OpenReview
3DCS: Datasets and Benchmark for Evaluating Conformational Sensitivity in Molecular Representations
- Link: OpenReview
Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models
- Link: OpenReview
Low-Pass Filtering Improves Behavioral Alignment of Vision Models
- Link: OpenReview
Bayesian Test-Time Adaptation via Dirichlet feature projection and GMM-Driven Inference for Motor Imagery EEG Decoding
- Link: OpenReview
Functional MRI Time Series Generation via Wavelet-Based Image Transform and Spectral Flow Matching for Brain Disorder Identification
- Link: OpenReview
Moving Beyond Diffusion: Hierarchy-to-Hierarchy Autoregression for fMRI-to-Image Reconstruction
- Link: OpenReview
No Detail Left Behind: Revisiting Self-Retrieval for Fine-Grained Image Captioning
- Link: OpenReview
2669. SyncTrack: Rhythmic Stability and Synchronization in Multi-Track Music Generation
- Topics: Audio & Speech
Product of Experts for Visual Generation
- Link: OpenReview
Zebra-CoT: A Dataset for Interleaved Vision-Language Reasoning
- Link: OpenReview
OCR-Reasoning Benchmark: Unveiling the True Capabilities of MLLMs in Complex Text-Rich Image Reasoning
- Link: OpenReview
Self-Improving Loops for Visual Robotic Planning
- Link: OpenReview
Enhancing Visual Token Representations for Video Large Language Models via Training-free Spatial-Temporal Pooling and Gridding
- Link: OpenReview
VaseVQA-3D: Benchmarking 3D VLMs on Ancient Greek Pottery
- Link: OpenReview
YoNoSplat: You Only Need One Model for Feedforward 3D Gaussian Splatting
- Link: OpenReview
WorldEdit: Towards Open-World Image Editing with a Knowledge-Informed Benchmark
- Link: OpenReview
ReLi3D: Relightable Multi-view 3D Reconstruction with Disentangled Illumination
- Link: OpenReview
Do 3D Large Language Models Really Understand 3D Spatial Relationships?
- Link: OpenReview
FaSTA*: Fast-Slow Toolpath Agent with Subroutine Mining for Efficient Multi-turn Image Editing
- Link: OpenReview
K-Sort Eval: Efficient Preference Evaluation for Visual Generation via Corrected VLM-as-a-Judge
- Link: OpenReview
Object Fidelity Diffusion for Remote Sensing Image Generation
- Link: OpenReview
Next Visual Granularity Generation
- Link: OpenReview
MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning
- Link: OpenReview
SketchingReality: From Freehand Scene Sketches to Photorealistic Images
- Link: OpenReview
RefAny3D: 3D Asset-Referenced Diffusion Models for Image Generation
- Link: OpenReview
Many-for-Many: Unify the Training of Multiple Video and Image Generation and Manipulation Tasks
- Link: OpenReview
VINCIE: Unlocking In-context Image Editing from Video
- Link: OpenReview
DragFlow: Unleashing DiT Priors with Region-Based Supervision for Drag Editing
- Link: OpenReview
Asynchronous Denoising Diffusion Models for Aligning Text-to-Image Generation
- Link: OpenReview
CLAP: Unsupervised 3D Representation Learning for Fusion 3D Perception via Curvature Sampling and Prototype Learning
- Link: OpenReview
DiffVax: Optimization-Free Image Immunization Against Diffusion-Based Editing
- Link: OpenReview
MotionWeaver: Holistic 4D-Anchored Framework for Multi-Humanoid Image Animation
- Link: OpenReview
Condition Errors Refinement in Autoregressive Image Generation with Diffusion Loss
- Link: OpenReview
Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing
- Link: OpenReview
MULTIMODALITY AS SUPERVISION: SELF-SUPERVISED SPECIALIZATION TO THE TEST ENVIRONMENT VIA MULTIMODALITY
- Link: OpenReview
MoCa: Modeling Object Consistency for 3D Camera Control in Video Generation
- Link: OpenReview
ACCORD: Alleviating Concept Coupling through Dependence Regularization for Text-to-Image Diffusion Personalization
- Link: OpenReview
From Prediction to Perfection: Introducing Refinement to Autoregressive Image Generation
- Link: OpenReview
reAR: Rethinking Visual Autoregressive Models via Token-wise Consistency Regularization
- Link: OpenReview
There and Back Again: On the relation between Noise and Image Inversions in Diffusion Models
- Link: OpenReview
LinearSR: Unlocking Linear Attention for Stable and Efficient Image Super-Resolution
- Link: OpenReview
Preserve and Personalize: Personalized Text-to-Image Diffusion Models without Distributional Drift
- Link: OpenReview
ReconViaGen: Towards Accurate Multi-view 3D Object Reconstruction via Generation
- Link: OpenReview
ComGS: Efficient 3D Object-Scene Composition via Surface Octahedral Probes
- Link: OpenReview
EgoHandICL: Egocentric 3D Hand Reconstruction with In-Context Learning
- Link: OpenReview
BigMaQ: A Big Macaque Motion and Animation Dataset Bridging Image and 3D Pose Representations
- Link: OpenReview
Direct Reward Fine-Tuning on Poses for Single Image to 3D Human in the Wild
- Link: OpenReview
Interp3D: Correspondence-aware Interpolation for Generative Textured 3D Morphing
- Link: OpenReview
MoE-GS: Mixture of Experts for Dynamic Gaussian Splatting
- Link: OpenReview
Beyond Visual Reconstruction Quality: Object Perception-aware 3D Gaussian Splatting for Autonomous Driving
- Link: OpenReview
PatchRefiner V2: Fast and Lightweight Real-Domain High-Resolution Metric Depth Estimation
- Link: OpenReview
CogniMap3D: Cognitive 3D Mapping and Rapid Retrieval
- Link: OpenReview
Rethinking Consistent Multi-Label Classification Under Inexact Supervision
- Link: OpenReview
SSD-GS: Scattering and Shadow Decomposition for Relightable 3D Gaussian Splatting
- Link: OpenReview
Open-Set Semantic Gaussian Splatting SLAM with Expandable Representation
- Link: OpenReview
Gradient-Direction-Aware Density Control for 3D Gaussian Splatting
- Link: OpenReview
DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning
- Link: OpenReview
Patch-as-Decodable-Token: Towards Unified Multi-Modal Vision Tasks in MLLMs
- Link: OpenReview
AVoCaDO: An Audiovisual Video Captioner Driven by Temporal Orchestration
- Link: OpenReview
SAM-Veteran: An MLLM-Based Human-like SAM Agent for Reasoning Segmentation
- Link: OpenReview
Thicker and Quicker: The Jumbo Token for Fast Plain Vision Transformers
- Link: OpenReview
ViMo: A Generative Visual GUI World Model for App Agents
- Link: OpenReview
Demystifying Supervision Data Generalization in Multimodal LMs
- Link: OpenReview
Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations
- Link: OpenReview
ERGO: Efficient High-Resolution Visual Understanding for Vision-Language Models
- Link: OpenReview
HierLoc: Hyperbolic Entity Embeddings for Hierarchical Visual Geolocation
- Link: OpenReview
ExpVid: A Benchmark for Experiment Video Understanding & Reasoning
- Link: OpenReview
Self-Aug: Query and Entropy Adaptive Decoding for Large Vision-Language Models
- Link: OpenReview
JUDO: A Juxtaposed Domain-Oriented Multimodal Reasoner for Industrial Anomaly QA
- Link: OpenReview
Part-X-MLLM: Part-aware 3D Multimodal Large Language Model
- Link: OpenReview
Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks
- Link: OpenReview
FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded Reasoning
- Link: OpenReview
Scenethesis: A Language and Vision Agentic Framework for 3D Scene Generation
- Link: OpenReview
H2OFlow: Grounding Human-Object Affordances with 3D Generative Models and Dense Diffused Flows
- Link: OpenReview
SpatialViz-Bench: A Cognitively-Grounded Benchmark for Diagnosing Spatial Visualization in MLLMs
- Link: OpenReview
WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models
- Link: OpenReview
FLARE: Fully Integration of Vision-Language Representations for Deep Cross-Modal Understanding
- Link: OpenReview
SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
- Link: OpenReview
Vision-SR1: Self-Rewarding Vision-Language Model via Reasoning Decomposition and Multi-Reward Policy Optimization
- Link: OpenReview
WebWatcher: Breaking New Frontiers of Vision-Language Deep Research Agent
- Link: OpenReview
AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions
- Link: OpenReview
Decomposed Attention Fusion in MLLMs for Training-free Video Reasoning Segmentation
- Link: OpenReview
TABLET: A Large-Scale Dataset for Robust Visual Table Understanding
- Link: OpenReview
SCoT: Teaching 3D-LLMs to Think Spatially with Million-scale CoT Annotations
- Link: OpenReview
VisionTrim: Unified Vision Token Compression for Training-Free MLLM Acceleration
- Link: OpenReview
IVC-Prune: Revealing the Implicit Visual Coordinates in LVLMs for Vision Token Pruning
- Link: OpenReview
CoT-RVS: Zero-Shot Chain-of-Thought Reasoning Segmentation for Videos
- Link: OpenReview
FOCUS: Efficient Keyframe Selection for Long Video Understanding
- Link: OpenReview
OVSeg3R: Learn Open-vocabulary Instance Segmentation from 2D via 3D Reconstruction
- Link: OpenReview
Let's Split Up: Zero-Shot Classifier Edits for Fine-Grained Video Understanding
- Link: OpenReview
Interpretable 3D Neural Object Volumes for Robust Conceptual Reasoning
- Link: OpenReview
QueryStream: Advancing Streaming Video Understanding with Query-Aware Pruning and Proactive Response
- Link: OpenReview
Grounding-IQA: Grounding Multimodal Language Model for Image Quality Assessment
- Link: OpenReview
UniRestorer: Universal Image Restoration via Adaptively Estimating Image Degradation at Proper Granularity
- Link: OpenReview
Parameterization-Based Dataset Distillation of 3D Point Clouds through Learnable Shape Morphing
- Link: OpenReview
CoEmoGen: Towards Semantically-Coherent and Scalable Emotional Image Content Generation
- Link: OpenReview
Anime-Ready: Controllable 3D Anime Character Generation with Body-Aligned Component-Wise Garment Modeling
- Link: OpenReview
LiFR-Seg: Anytime High-Frame-Rate Segmentation via Event-Guided Propagation
- Link: OpenReview
DETR-ViP: Detection Transformer with Robust Discriminative Visual Prompts
- Link: OpenReview
VSF: Simple, Efficient, and Effective Negative Guidance in Few-Step Image Generation Models By Value Sign Flip
- Link: OpenReview
WeTok: Powerful Discrete Tokenization for High-Fidelity Visual Reconstruction
- Link: OpenReview
Output Supervision Can Obfuscate the Chain of Thought
- Link: OpenReview
RelayFormer: A Unified Local-Global Attention Framework for Scalable Image and Video Manipulation Localization
- Link: OpenReview
Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs
- Link: OpenReview
VUDG: A Dataset for Video Understanding Domain Generalization
- Link: OpenReview
Block Recurrent Dynamics in Vision Transformers
- Link: OpenReview
Station2Radar: Query‑Conditioned Gaussian Splatting for Precipitation Field
- Link: OpenReview
DeepPrim: a Physics-Driven 3D Short-term Weather Forecaster via Primitive Equation Learning
- Link: OpenReview
PathChat-SegR1: Reasoning Segmentation in Pathology via SO-GRPO
- Link: OpenReview
P3D: Highly Scalable 3D Neural Surrogates for Physics Simulations with Global Context
- Link: OpenReview
OmniEVA: Embodied Versatile Planner via Task-Adaptive 3D-Grounded and Embodiment-aware Reasoning
- Link: OpenReview
FASTer: Toward Powerful and Efficient Autoregressive Vision–Language–Action Models with Learnable Action Tokenizer and Block-wise Decoding
- Link: OpenReview
From Spatial to Actions: Grounding Vision-Language-Action Model in Spatial Foundation Priors
- Link: OpenReview
Sparse Imagination for Efficient Visual World Model Planning
- Link: OpenReview
EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic Manipulation
- Link: OpenReview
CoNavBench: Collaborative Long-Horizon Vision-Language Navigation Benchmark
- Link: OpenReview
AsyncBEV: Cross-modal flow alignment in Asynchronous 3D Object Detection
- Link: OpenReview
3D-aware Disentangled Representation for Compositional Reinforcement Learning
- Link: OpenReview
Entropy Regularizing Activation: Boosting Continuous Control, Large Language Models, and Image Classification with Activation as Entropy Constraints
- Link: OpenReview
When would Vision-Proprioception Policies Fail in Robotic Manipulation?
- Link: OpenReview
PINFDiT: Energy-Based Physics-Informed Diffusion Transformers for General-purpose Time Series Tasks
- Link: OpenReview
Fracture-GS: Dynamic Fracture Simulation with Physics-Integrated Gaussian Splatting
- Link: OpenReview
Mobile-GS: Real-time Gaussian Splatting for Mobile Devices
- Link: OpenReview
Advancing Complex Video Object Segmentation via Progressive Concept Construction
- Link: OpenReview
PixelCraft: A Multi-Agent system for High-Fidelity Visual Reasoning on Structured Images
- Link: OpenReview
GoT-R1: Unleashing Reasoning Capability of Autoregressive Visual Generation with Reinforcement Learning
- Link: OpenReview
SigLIP-HD by Fine-to-Coarse Supervision
- Link: OpenReview
Zeros can be Informative: Masked Binary U-Net for Image Segmentation on Tensor Cores
- Link: OpenReview
Falcon: Fast Proximal Linearization of Normalized Cuts for Unsupervised Image Segmentation
- Link: OpenReview
TAVAE: A VAE with Adaptable Priors Explains Contextual Modulation in the Visual Cortex
- Link: OpenReview
Look Carefully: Adaptive Visual Reinforcements in Multimodal Large Language Models for Hallucination Mitigation
- Link: OpenReview
Seeing but Not Believing: Probing the Disconnect Between Visual Attention and Answer Correctness in VLMs
- Link: OpenReview
Local Success Does Not Compose: Benchmarking Large Language Models for Compositional Formal Verification
- Link: OpenReview
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale
- Link: OpenReview
Urban Socio-Semantic Segmentation with Vision-Language Reasoning
- Link: OpenReview
Towards Faithful Reasoning in Remote Sensing: A Perceptually-Grounded GeoSpatial Chain-of-Thought for Vision-Language Models
- Link: OpenReview
Visual Compositional Tuning
- Link: OpenReview
Visual Prompt-Agnostic Evolution
- Link: OpenReview
Point-MoE: Large-Scale Multi-Dataset Training with Mixture-of-Experts for 3D Semantic Segmentation
- Link: OpenReview
Error Notebook-Guided, Training-Free Part Retrieval in 3D CAD Assemblies via Vision-Language Models
- Link: OpenReview
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
- Link: OpenReview
Adaptive Domain Shift in Diffusion Models for Cross-Modality Image Translation
- Link: OpenReview
InfoScan: Information-Efficient Visual Scanning via Resource-Adaptive Walks
- Link: OpenReview
EA3D: Event-Augmented 3D Diffusion for Generalizable Novel View Synthesis
- Link: OpenReview
SpikeGen: Decoupled “Rods and Cones” Visual Representation Processing with Latent Generative Framework
- Link: OpenReview
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
- Link: OpenReview
IV-Bench: A Benchmark for Image-Grounded Video Perception and Reasoning in Multimodal LLMs
- Link: OpenReview
GPT4Scene: Understand 3D Scenes from Videos with Vision-Language Models
- Link: OpenReview
PICS: Pairwise Image Compositing with Spatial Interactions
- Link: OpenReview
DistillKac: Few-Step Image Generation via Damped Wave Equations
- Link: OpenReview
Constantly Improving Image Models Need Constantly Improving Benchmarks
- Link: OpenReview
Temporal Slowness in Central Vision Drives Semantic Object Learning
- Link: OpenReview
MVAR: Visual Autoregressive Modeling with Scale and Spatial Markovian Conditioning
- Link: OpenReview
W-EDIT: A Wavelet-Based Frequency-Aware Framework for Text-Driven Image Editing
- Link: OpenReview
VisualPrompter: Semantic-Aware Prompt Optimization with Visual Feedback for Text-to-Image Synthesis
- Link: OpenReview
IC-Custom: Diverse Image Customization via In-Context Learning
- Link: OpenReview
Visual Autoregressive Modeling for Instruction-Guided Image Editing
- Link: OpenReview
EditReward: A Human-Aligned Reward Model for Instruction-Guided Image Editing
- Link: OpenReview
PHyCLIP: -Product of Hyperbolic Factors Unifies Hierarchy and Compositionality in Vision-Language Representation Learning
- Link: OpenReview
VIRTUE: Visual-Interactive Text-Image Universal Embedder
- Link: OpenReview
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
- Link: OpenReview
Learning an Image Editing Model without Image Editing Pairs
- Link: OpenReview
RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement Learning
- Link: OpenReview
EditScore: Unlocking Online RL for Image Editing via High-Fidelity Reward Modeling
- Link: OpenReview
PICABench: How Far are We from Physical Realistic Image Editing?
- Link: OpenReview
Enhanced Continual Learning of Vision-Language Models with Model Fusion
- Link: OpenReview
Soft-Di[M]O: Improving One-Step Discrete Image Generation with Soft Embeddings
- Link: OpenReview
ReDDiT: Rehashing Noise for Discrete Visual Generation
- Link: OpenReview
Directional Textual Inversion for Personalized Text-to-Image Generation
- Link: OpenReview
GIR-Bench: Versatile Benchmark for Generating Images with Reasoning
- Link: OpenReview
FantasyWorld: Geometry-Consistent World Modeling via Unified Video and 3D Prediction
- Link: OpenReview
ToonComposer: Streamlining Cartoon Production with Generative Post-Keyframing
- Link: OpenReview
Generalized Compressed Sensing for Image Reconstruction with Diffusion Probabilistic Models
- Link: OpenReview
3760. In-Context Learning of Temporal Point Processes with Foundation Inference Models
- Topics: Efficiency & Compression
FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting
- Link: OpenReview
A Scene is Worth a Thousand Features: Feed-Forward Camera Localization from a Collection of Image Features
- Link: OpenReview
MEGS^2: Memory-Efficient Gaussian Splatting via Spherical Gaussians and Unified Pruning
- Link: OpenReview
StreamSplat: Towards Online Dynamic 3D Reconstruction from Uncalibrated Video Streams
- Link: OpenReview
DreamCS: Geometry-Aware Text-to-3D Generation with Unpaired 3D Reward Supervision
- Link: OpenReview
FastAvatar: Towards Unified and Fast 3D Avatar Reconstruction with Large Gaussian Reconstruction Transformers
- Link: OpenReview
Densemarks: Learning Canonical Embeddings for Human Heads Images via Point Tracks
- Link: OpenReview
Towards Physically Executable 3D Gaussian for Embodied Navigation
- Link: OpenReview
Distractor-free Generalizable 3D Gaussian Splatting
- Link: OpenReview
Neural Compression of 3D Meshes using Sparse Implicit Representation
- Link: OpenReview
ReSplat: Degradation-agnostic Feed-forward Gaussian Splatting via Self-guided Residual Diffusion
- Link: OpenReview
Color3D: Controllable and Consistent 3D Colorization with Personalized Colorizer
- Link: OpenReview
Teaching VLMs to Admit Uncertainty in OCR from Lossy Visual Inputs
- Link: OpenReview
Revisiting [CLS] and Patch Token Interaction in Vision Transformers
- Link: OpenReview
RAR: Reversing Visual Attention Re-Sinking for Unlocking Potential in Multimodal Large Language Models
- Link: OpenReview
ChainMPQ: Interleaved Text-Image Reasoning Chains for Mitigating Relation Hallucinations
- Link: OpenReview
VL-JEPA: Joint Embedding Predictive Architecture for Vision-language
- Link: OpenReview
Revisiting Multimodal Positional Encoding in Vision–Language Models
- Link: OpenReview
Towards Text-Mask Consistency in Medical Image Segmentation
- Link: OpenReview
Progressive Online Video Understanding with Evidence-Aligned Timing and Transparent Decisions
- Link: OpenReview
Divid: Disentangled Spatial-Temporal Modeling within LLMs for Temporally Grounded Video Understanding
- Link: OpenReview
VisJudge-Bench: Aesthetics and Quality Assessment of Visualizations
- Link: OpenReview
V2P-Bench: Evaluating Video-Language Understanding with Visual Prompts for Better Human-Model Interaction
- Link: OpenReview
Fine-R1: Make Multi-modal LLMs Excel in Fine-Grained Visual Recognition by Chain-of-Thought Reasoning
- Link: OpenReview
DAVE: A VLM Vision Encoder for Document Understanding and Web Agents
- Link: OpenReview
A Training-Free Framework for Long Video Understanding via Video-Query-Options Similarity
- Link: OpenReview
Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes
- Link: OpenReview
MetaCaptioner: Towards Generalist Visual Captioning with Open-source Suites
- Link: OpenReview
ProxyThinker: Test-Time Guidance through Small Visual Reasoners
- Link: OpenReview
RayI2P: Learning Rays for Image-to-Point Cloud Registration
- Link: OpenReview
OmniCVR: A Benchmark for Omni-Composed Video Retrieval with Vision, Audio, and Text
- Link: OpenReview
PTNET: A PROPOSAL-CENTRIC TRANSFORMER NET- WORK FOR 3D OBJECT DETECTION
- Link: OpenReview
Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via Speculation
- Link: OpenReview
CircuitSense: A Hierarchical MLLM Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process
- Link: OpenReview
EventFlash: Towards Efficient MLLMs for Event-Based Vision
- Link: OpenReview
PAGE-4D: Disentangled Pose and Geometry Estimation for VGGT-4D Perception
- Link: OpenReview
Human Uncertainty-Aware Data Selection and Automatic Labeling in Visual Question Answering
- Link: OpenReview
Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
- Link: OpenReview
VTool-R1: VLMs Learn to Think with Images via Reinforcement Learning on Multimodal Tool Use
- Link: OpenReview
3D Aware Region Prompted Vision Language Model
- Link: OpenReview
Foundation Visual Encoders Are Secretly Few-Shot Anomaly Detectors
- Link: OpenReview
NePTune: A Neuro-Pythonic Framework for Tunable Compositional Reasoning on Vision-Language
- Link: OpenReview
Plug, Play, and Fortify: A Low-Cost Module for Robust Multimodal Image Understanding Models
- Link: OpenReview
Language-guided Open-world Video Anomaly Detection under Weak Supervision
- Link: OpenReview
RAVENEA: A Benchmark for Multimodal Retrieval-Augmented Visual Culture Understanding
- Link: OpenReview
Agentic Jigsaw Interaction Learning for Enhancing Visual Perception and Reasoning in Vision-Language Models
- Link: OpenReview
VideoAnchor: Reinforcing Subspace-Structured Visual Cues for Coherent Visual-Spatial Reasoning
- Link: OpenReview
PartSAM: A Scalable Promptable Part Segmentation Model Trained on Native 3D Data
- Link: OpenReview
Object-Centric Refinement for Enhanced Zero-Shot Segmentation
- Link: OpenReview
Adaptive Augmentation-Aware Latent Learning for Robust LiDAR Semantic Segmentation
- Link: OpenReview
Rethinking Model Calibration through Spectral Entropy Regularization in Medical Image Segmentation
- Link: OpenReview
GOOD: Geometry-guided Out-of-Distribution Modeling for Open-set Test-time Adaptation in Point Cloud Semantic Segmentation
- Link: OpenReview
Text-Aware Image Restoration with Diffusion Models
- Link: OpenReview
Pixel3DMM: Versatile Screen-Space Priors for Single-Image 3D Face Reconstruction
- Link: OpenReview
InT: Self-Proposed Interventions Enable Credit Assignment in LLM Reasoning
- Link: OpenReview
Learning Domain-Aware Task Prompt Representations for Multi-Domain All-in-One Image Restoration
- Link: OpenReview
Energy-oriented Diffusion Bridge for Image Restoration with Foundational Diffusion Models
- Link: OpenReview
DNOD: Deformable Neural Operators for Object Detection in SAR Images
- Link: OpenReview
3989. Reasoning Scaffolding: Distilling the Flow of Thought from LLMs
- Topics: LLMs & Foundation Models
Autoregressive-based Progressive Coding for Ultra-Low Bitrate Image Compression
- Link: OpenReview
Transferable and Stealthy Adversarial Attacks on Large Vision-Language Models
- Link: OpenReview
Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language Models
- Link: OpenReview
Uncovering Conceptual Blindspots in Generative Image Models Using Sparse Autoencoders
- Link: OpenReview
Pay Less Attention to Function Words for Free Robustness of Vision-Language Models
- Link: OpenReview
Fed-Duet: Dual Expert-Orchestrated Framework for Continual Federated Vision-Language Learning
- Link: OpenReview
Composer: A Search Framework for Hybrid Neural Architecture Design
- Link: OpenReview
Naming to Learn: Class Incremental Learning for Vision-Language Model with Unlabeled Data
- Link: OpenReview
DRIFT: Decompose, Retrieve, Illustrate, then Formalize Theorems
- Link: OpenReview
Feature segregation by signed weights in artificial vision systems and biological models
- Link: OpenReview
PoseX: AI Defeats Physics-based Methods on Protein Ligand Cross-Docking
- Link: OpenReview
Training Dynamics of Learning 3D-Rotational Equivariance
- Link: OpenReview
4171. Explaining Grokking and Information Bottleneck through Neural Collapse Emergence
- Topics: Interpretability & Mechanistic Interpretability, Theory & Deep Learning Theory
RIDER: 3D RNA Inverse Design with Reinforcement Learning-Guided Diffusion
- Link: OpenReview
CryoLVM: Self-supervised Learning from Cryo-EM Density Maps with Large Vision Models
- Link: OpenReview
Random Anchors with Low-rank Decorrelated Learning: A Minimalist Pipeline for Class-Incremental Medical Image Classification
- Link: OpenReview
U2-BENCH: Benchmarking Large Vision-Language Models on Ultrasound Understanding
- Link: OpenReview
MedVR: Annotation-Free Medical Visual Reasoning via Agentic Reinforcement Learning
- Link: OpenReview
MedLesionVQA: A Multimodal Benchmark Emulating Clinical Visual Diagnosis for Body Surface Health
- Link: OpenReview
Search Self-Play: Pushing the Frontier of Agent Capability without Supervision
- Link: OpenReview
Robust Fine-tuning of Vision-Language-Action Robot Policies via Parameter Merging
- Link: OpenReview
All-day Multi-scenes Lifelong Vision-and-Language Navigation with Tucker Adaptation
- Link: OpenReview
ME: Continual Vision-and-Language Navigation via Mixture of Macro and Micro Experts
- Link: OpenReview
VLMgineer: Vision-Language Models as Robotic Toolsmiths
- Link: OpenReview
Self-Improving Vision-Language-Action Models with Data Generation via Residual RL
- Link: OpenReview
VER: Vision Expert Transformer for Robot Learning via Foundation Distillation and Dynamic Routing
- Link: OpenReview
Align-Then-stEer: Adapting the Vision-Language Action Models through Unified Latent Guidance
- Link: OpenReview
Uncertainty-Aware Gaussian Map for Vision-Language Navigation
- Link: OpenReview
Ground Slow, Move Fast: A Dual-System Foundation Model for Generalizable Vision-Language Navigation
- Link: OpenReview
Capturing Visual Environment Structure Correlates with Control Performance
- Link: OpenReview
Compositional Visual Planning via Inference-Time Diffusion Scaling
- Link: OpenReview
Scaling Large Vision-Language Model RL Training via Efficient Load Balancing
- Link: OpenReview
MMR-Life: Piecing Together Real-life Scenes for Multimodal Multi-image Reasoning
- Link: OpenReview
Efficient Degradation-agnostic Image Restoration via Channel-Wise Functional Decomposition and Manifold Regularization
- Link: OpenReview
How Do Medical MLLMs Fail? A Study on Visual Grounding in Medical Images
- Link: OpenReview
SpatialHand: Generative Object Manipulation from 3D Prespective
- Link: OpenReview
The Human Brain as a Dynamic Mixture of Expert Models in Video Understanding
- Link: OpenReview
Towards Interpretable Visual Decoding with Attention to Brain Representations
- Link: OpenReview
CerebraGloss: Instruction-Tuning a Large Vision-Language Model for Fine-Grained Clinical EEG Interpretation
- Link: OpenReview
InclusiveVidPose: Bridging the Pose Estimation Gap for Individuals with Limb Deficiencies in Video-Based Motion
- Link: OpenReview
The Less You Depend, The More You Learn: Synthesizing Novel Views from Sparse, Unposed Images without Any 3D Knowledge
- Link: OpenReview
A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual Decoding
- Link: OpenReview
Rethinking Causal Mask Attention for Vision-Language Inference
- Link: OpenReview
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
- Link: OpenReview
CoDA: Agentic Systems for Collaborative Data Visualization
- Link: OpenReview
SemHiTok: A Unified Image Tokenizer via Semantic-Guided Hierarchical Codebook for Multimodal Understanding and Generation
- Link: OpenReview
Deforming Videos to Masks: Flow Matching for Referring Video Segmentation
- Link: OpenReview
SynCoGen: Synthesizable 3D Molecule Generation via Joint Reaction and Coordinate Modeling
- Link: OpenReview
Reconciling Visual Perception and Generation in Diffusion Models
- Link: OpenReview
Disco: Densely-overlapping Cell Instance Segmentation via Adjacency-aware Collaborative Coloring
- Link: OpenReview
Visual Jigsaw Post-Training Improves MLLMs
- Link: OpenReview
Pyramid Patchification Flow for Visual Generation
- Link: OpenReview
On Robustness of Vision-Language-Action Model against Multi-Modal Perturbations
- Link: OpenReview
Cross-Modal Redundancy and the Geometry of Vision–Language Embeddings
- Link: OpenReview
AttTok: Marrying Attribute Tokens with Generative Pre-trained Vision-Language Models towards Medical Image Understanding
- Link: OpenReview
SketchEvo: Leveraging Drawing Dynamics for Enhanced Image Synthesis
- Link: OpenReview
Consis-GCPO: Consistency-Preserving Group Causal Preference Optimization for Vision Customization
- Link: OpenReview
From EduVisBench to EduVisAgent: A Benchmark and Multi-Agent Framework for Reasoning-Driven Pedagogical Visualization
- Link: OpenReview
CoDA: From Text-to-Image Diffusion Models to Training-Free Dataset Distillation
- Link: OpenReview
Cartridges: Lightweight and general-purpose long context representations via self-study
- Link: OpenReview
Group Critical-token Policy Optimization for Autoregressive Image Generation
- Link: OpenReview
Continuous Space-Time Video Super-Resolution with 3D Fourier Fields
- Link: OpenReview
Arbitrary-Shaped Image Generation via Spherical Neural Field Diffusion
- Link: OpenReview
Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World Modeling
- Link: OpenReview
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation
- Link: OpenReview
Unified 3D Scene Understanding Through Physical World Modeling
- Link: OpenReview
From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image Generation
- Link: OpenReview
PI-Light: Physics-Inspired Diffusion for Full-Image Relighting
- Link: OpenReview
TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment
- Link: OpenReview
Beyond Static Vision: Scene Dynamic Field Unlocks Intuitive Physics Understanding in Multi-modal Large Language Models
- Link: OpenReview
Factuality Matters: When Image Generation and Editing Meet Structured Visuals
- Link: OpenReview
RegionE: Adaptive Region-Aware Generation for Efficient Image Editing
- Link: OpenReview
SoftCFG: Uncertainty-guided Stable Guidance for Visual Autoregressive Model
- Link: OpenReview
FastGHA: Generalized Few-Shot 3D Gaussian Head Avatars with Real-Time Animation
- Link: OpenReview
CoDi: Subject-Consistent and Pose-Diverse Text-to-Image Generation
- Link: OpenReview
Latent Wavelet Diffusion For Ultra High-Resolution Image Synthesis
- Link: OpenReview
One-Step Flow for Image Super-Resolution with Tunable Fidelity-Realism Trade-offs
- Link: OpenReview
Disentangled Hierarchical VAE for 3D Human-Human Interaction Generation
- Link: OpenReview
Inverse Virtual Try-On: Generating Multi-Category Product-Style Images from Clothed Individuals
- Link: OpenReview
Revisit Visual Prompt Tuning: The Expressiveness of Prompt Experts
- Link: OpenReview
Charts Are Not Images: On the Challenges of Scientific Chart Editing
- Link: OpenReview
Geometric Image Editing via Effects-Sensitive In-Context Inpainting with Diffusion Transformers
- Link: OpenReview
Retain and Adapt: Auto-Balanced Model Editing for Open-Vocabulary Object Detection under Domain Shifts
- Link: OpenReview
Aligned Novel View Image and Geometry Synthesis via Cross-modal Attention Instillation
- Link: OpenReview
CONSIGN: Conformal Segmentation Informed by Spatial Groupings via Decomposition
- Link: OpenReview
STream3R: Scalable Sequential 3D Reconstruction with Causal Transformer
- Link: OpenReview
From Tokens to Nodes: Semantic-Guided Motion Control for Dynamic 3D Gaussian Splatting
- Link: OpenReview
Quantized Visual Geometry Grounded Transformer
- Link: OpenReview
ODE-GS: Latent ODEs for Dynamic Scene Extrapolation with 3D Gaussian Splatting
- Link: OpenReview
A^2TG: Adaptive Anisotropic Textured Gaussians for Efficient 3D Scene Representation
- Link: OpenReview
Sat3DGen: Comprehensive Street-Level 3D Scene Generation from Single Satellite Image
- Link: OpenReview
IncVGGT: Incremental VGGT for Memory-Bounded Long-Range 3D Reconstruction
- Link: OpenReview
3DGEER: 3D Gaussian Rendering Made Exact and Efficient for Generic Cameras
- Link: OpenReview
Streaming Visual Geometry Transformer
- Link: OpenReview
pySpatial: Generating 3D Visual Programs for Zero-Shot Spatial Reasoning
- Link: OpenReview
ReWatch-R1: Boosting Complex Video Reasoning in Large Vision-Language Models through Agentic Data Synthesis
- Link: OpenReview
ViTSP: A Vision Language Models Guided Framework for Solving Large-Scale Traveling Salesman Problems
- Link: OpenReview
FLoC: Facility Location-Based Efficient Visual Token Compression for Long Video Understanding
- Link: OpenReview
iLLaVA: An Image is Worth Fewer Than 1/3 Input Tokens in Large Multimodal Models
- Link: OpenReview
OmniActor: A Generalist GUI and Embodied Agent for 2D&3D Worlds
- Link: OpenReview
VGR: Visual Grounded Reasoning
- Link: OpenReview
Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression Tasks
- Link: OpenReview
Benchmarking Large Vision-Language Models on Fine-Grained Image Tasks: A Comprehensive Evaluation
- Link: OpenReview
Investigating Redundancy in Multimodal Large Language Models with Multiple Vision Encoders
- Link: OpenReview
Dynamic Multimodal Activation Steering for Hallucination Mitigation in Large Vision-Language Models
- Link: OpenReview
More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language Models
- Link: OpenReview
UniUGG: Unified 3D Understanding and Generation via Geometric-Semantic Encoding
- Link: OpenReview
WIMFRIS: WIndow Mamba Fusion and Parameter Efficient Tuning for Referring Image Segmentation
- Link: OpenReview
PoSh: Using Scene Graphs to Guide LLMs-as-a-Judge for Detailed Image Descriptions
- Link: OpenReview
Faster Vision Transformers with Adaptive Patches
- Link: OpenReview
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
- Link: OpenReview
Flatness Guided Test-Time Adaptation for Vision-Language Models
- Link: OpenReview
FRIEDA: Benchmarking Multi-Step Cartographic Reasoning in Vision-Language Models
- Link: OpenReview
InternSpatial: A Comprehensive Dataset for Spatial Reasoning in Vision-Language Models
- Link: OpenReview
RegionReasoner: Region-Grounded Multi-Round Visual Reasoning
- Link: OpenReview
Knowledge Exchange with Confidence: Cost-Effective LLM Integration for Reliable and Efficient Visual Question Answering
- Link: OpenReview
VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks
- Link: OpenReview
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
- Link: OpenReview
Efficient-SAM2: Accelerating SAM2 with Object-Aware Visual Encoding and Memory Retrieval
- Link: OpenReview
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
- Link: OpenReview
LearnPruner: Rethinking Attention-based Token Pruning in Vision Language Models
- Link: OpenReview
LINK: Learning Instance-level Knowledge from Vision-Language Models for Human-Object Interaction Detection
- Link: OpenReview
ScaleLong: A Multi-Timescale Benchmark for Long Video Understanding
- Link: OpenReview
Mixture-of-Visual-Thoughts: Exploring Context-Adaptive Reasoning Mode Selection for General Visual Reasoning
- Link: OpenReview
VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?
- Link: OpenReview
Sequential Information Bottleneck Fusion: Towards Robust and Generalizable Multi-Modal Brain Tumor Segmentation
- Link: OpenReview
GeoPurify: A Data-Efficient Geometric Distillation Framework for Open-Vocabulary 3D Segmentation
- Link: OpenReview
Johnson-Lindenstrauss Lemma Guided Network for Efficient 3D Medical Segmentation
- Link: OpenReview
All Patches Matter, More Patches Better: Enhance AI-Generated Image Detection via Panoptic Patch Learning
- Link: OpenReview
HSIC Bottleneck for Cross-Generator and Domain-Incremental Synthetic Image Detection
- Link: OpenReview
Part-level Semantic-guided Contrastive Learning for Fine-grained Visual Classification
- Link: OpenReview
No Pixel Left Behind: A Detail-Preserving Architecture for Robust High-Resolution AI-Generated Image Detection
- Link: OpenReview
RestoreVAR: Visual Autoregressive Generation for All-in-One Image Restoration
- Link: OpenReview
LucidFlux: Caption-Free Universal Image Restoration via a Large-Scale Diffusion Transformer
- Link: OpenReview
FARTrack: Fast Autoregressive Visual Tracking with High Performance
- Link: OpenReview
QuaMo: Quaternion Motions for Vision-based 3D Human Kinematics Capture
- Link: OpenReview
Fore-Mamba3D: Mamba-based Foreground-Enhanced Encoding for 3D Object Detection
- Link: OpenReview
Video Scene Segmentation with Genre and Duration Signals
- Link: OpenReview
RobustSpring: Benchmarking Robustness to Image Corruptions for Optical Flow, Scene Flow and Stereo
- Link: OpenReview
Self-Guided Low Light Object Detection Framework
- Link: OpenReview
Seeing Through the PRISM: Compound & Controllable Restoration of Scientific Images
- Link: OpenReview
Training-Free Reward-Guided Image Editing via Trajectory Optimal Control
- Link: OpenReview
CardioComposer: Leveraging Differentiable Geometry for Compositional Control of Anatomical Diffusion Models
- Link: OpenReview
Towards Scalable Oversight via Partitioned Human Supervision
- Link: OpenReview
SERUM: Simple, Efficient, Robust, and Unifying Marking for Diffusion-based Image Generation
- Link: OpenReview
Physics-Inspired All-Pair Interaction Learning for 3D Dynamics Modeling
- Link: OpenReview
VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language Models
- Link: OpenReview
Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models
- Link: OpenReview
Do We Need All the Synthetic Data? Targeted Image Augmentation via Diffusion Models
- Link: OpenReview
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
- Link: OpenReview
A-TPT: Angular Diversity Calibration Properties for Test-Time Prompt Tuning of Vision-Language Models
- Link: OpenReview
CHAMMI-75: Pre-training multi-channel models with heterogeneous microscopy images
- Link: OpenReview
Exploiting Low-Dimensional Manifold of Features for Few-Shot Whole Slide Image Classification
- Link: OpenReview
PA3FF:Learning Part-Aware Dense 3D Feature Field For Generalizable Articulated Object Manipulation
- Link: OpenReview
Action-aware Dynamic Pruning for Efficient Vision-Language-Action Manipulation
- Link: OpenReview
Generalizable Coarse-to-Fine Robot Manipulation via Language-Aligned 3D Keypoints
- Link: OpenReview
SAGE: Spatial-visual Adaptive Graph Exploration for Efficient Visual Place Recognition
- Link: OpenReview
DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
- Link: OpenReview
TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models
- Link: OpenReview
LogicReward: Incentivizing LLM Reasoning via Step-Wise Logical Supervision
- Link: OpenReview
IGGT: Instance-Grounded Geometry Transformer for Semantic 3D Reconstruction
- Link: OpenReview
Vision-Zero: Scalable VLM Self-Evolution via Multi-Agent Self-Play
- Link: OpenReview
Language-Instructed Vision Embeddings for Controllable and Generalizable Perception
- Link: OpenReview
PreferThinker: Reasoning-based Personalized Image Preference Assessment
- Link: OpenReview
Repurposing Synthetic Data for Fine-grained Search Agent Supervision
- Link: OpenReview
StepORLM: A Self-Evolving Framework With Generative Process Supervision For Operations Research Language Models
- Link: OpenReview
A tale of two tails: Preferred and anti-preferred natural stimuli in visual cortex
- Link: OpenReview
GIQ: Benchmarking 3D Geometric Reasoning of Vision Foundation Models with Simulated and Real Polyhedra
- Link: OpenReview
Guidance Matters: Rethinking the Evaluation Pitfall for Text-to-Image Generation
- Link: OpenReview
Brain-IT: Image Reconstruction from fMRI via Brain-Interaction Transformer
- Link: OpenReview
MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided Diffusion
- Link: OpenReview
Mixing Importance with Diversity: Joint Optimization for KV Cache Compression in Large Vision-Language Models
- Link: OpenReview
Zero-shot Human Pose Estimation using Diffusion-based Inverse solvers
- Link: OpenReview
Why Keep Your Doubts to Yourself? Trading Visual Uncertainties among Vision-Language Models
- Link: OpenReview
MICLIP: Learning to Interpret Representation in Vision Models
- Link: OpenReview
VLM4VLA: Revisiting Vision-Language-Models in Vision-Language-Action Models
- Link: OpenReview
Mango-GS: Enhancing Spatio-Temporal Consistency in Dynamic Scenes Reconstruction using Multi-Frame Node-Guided 4D Gaussian Splatting
- Link: OpenReview
To Sink or Not to Sink: Visual Information Pathways in Large Vision-Language Models
- Link: OpenReview
Test-Time Optimization of 3D Point Cloud LLM via Manifold-Aware In-Context Guidance and Refinement
- Link: OpenReview
A High Quality Dataset and Reliable Evaluation for Interleaved Image-Text Generation
- Link: OpenReview
SpaceControl: Introducing Test-Time Spatial Control to 3D Generative Modeling
- Link: OpenReview
Interleaving Reasoning for Better Text-to-Image Generation
- Link: OpenReview
Modeling the Density of Pixel-level Self-supervised Embeddings for Unsupervised Pathology Segmentation in Medical CT
- Link: OpenReview
EVLP: Learning Unified Embodied Vision-Language Planner with Reinforced Supervised Fine-Tuning
- Link: OpenReview
VARestorer: One-Step VAR Distillation for Real-World Image Super-Resolution
- Link: OpenReview
Can Vision–Language Models Assess Graphic Design Aesthetics? A Benchmark, Evaluation, and Dataset Perspective.
- Link: OpenReview
JanusCoder: Towards a Foundational Visual-Programmatic Interface for Code Intelligence
- Link: OpenReview
OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- Link: OpenReview