- Published on
ICLR 2026 — Trust & Safety
Trust & Safety
442 papers (0 oral)
Benchmarking Empirical Privacy Protection for Adaptations of Large Language Models
- Link: OpenReview
Invisible Safety Threat: Malicious Finetuning for LLM via Steganography
- Link: OpenReview
The Shape of Adversarial Influence: Characterizing LLM Latent Spaces with Persistent Homology
- Link: OpenReview
Watch your steps: Dormant Adversarial Behaviors that Activate upon LLM Finetuning
- Link: OpenReview
LLM Fingerprinting via Semantically Conditioned Watermarks
- Link: OpenReview
Modality-free Graph In-context Alignment
- Link: OpenReview
Learning with Dual-level Noisy Correspondence for Multi-modal Entity Alignment
- Link: OpenReview
GLASS Flows: Efficient Inference for Reward Alignment of Flow and Diffusion Models
- Link: OpenReview
EigenBench: A Comparative Behavioral Measure of Value Alignment
- Link: OpenReview
SWINGARENA: Adversarial Programming Arena for Long-context GitHub Issue Solving
- Link: OpenReview
SafeDPO: A Simple Approach to Direct Preference Optimization with Enhanced Safety
- Link: OpenReview
Train-before-Test Harmonizes Language Model Rankings
- Link: OpenReview
SAFETY-GUIDED FLOW (SGF): A UNIFIED FRAMEWORK FOR NEGATIVE GUIDANCE IN SAFE GENERATION
- Link: OpenReview
RedTeamCUA: Realistic Adversarial Testing of Computer-Use Agents in Hybrid Web-OS Environments
- Link: OpenReview
CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering
- Link: OpenReview
Spherical Watermark: Encryption-Free, Lossless Watermarking for Diffusion Models
- Link: OpenReview
Beware Untrusted Simulators -- Reward-Free Backdoor Attacks in Reinforcement Learning
- Link: OpenReview
6. Enhancing Sparse Event Detection in Healthcare Time-Series via Adaptive Gate of Context–Detail Interaction
- Topics: Medical & Healthcare
SafeMoE: Safe Fine-Tuning for MoE LLMs by Aligning Harmful Input Routing
- Link: OpenReview
12. FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents
- Topics: LLMs & Foundation Models, Agents & Tool Use, Data-centric & Curation
BA-LoRA: Bias-Alleviating Low-Rank Adaptation to Mitigate Catastrophic Inheritance in Large Language Models
- Link: OpenReview
16. GeoBench: Rethinking Multimodal Geometric Problem-Solving via Hierarchical Evaluation
- Topics: Multi-modal & Vision-Language
Teach to Reason Safely: Policy-Guided Safety Tuning for MLRMs
- Link: OpenReview
32. Learning to Interpret Weight Differences in Language Models
- Topics: LLMs & Foundation Models
AI-for-Science Low-code Platform with Bayesian Adversarial Multi-Agent Framework
- Link: OpenReview
SesaHand: Enhancing 3D Hand Reconstruction via Controllable Generation with Semantic and Structural Alignment
- Link: OpenReview
Align-SAM: Seeking Flatter Minima for Better Cross-Subset Alignment
- Link: OpenReview
Dynamic Reflections: Probing Video Representations with Text Alignment
- Link: OpenReview
KBVQ-MoE: KLT-guided SVD with Bias-Corrected Vector Quantization for MoE Large Language Models
- Link: OpenReview
Alignment-Enhanced Integration of Connectivity and Spectral Sparsity in Dynamic Sparse Training of LLM
- Link: OpenReview
Learning Dynamics of Logits Debiasing for Long-Tailed Semi-Supervised Learning
- Link: OpenReview
Bridging Draft Policy Misalignment: Group Tree Optimization for Speculative Decoding
- Link: OpenReview
PromptHub: Enhancing Multi-Prompt Visual In-Context Learning with Locality-Aware Fusion, Concentration and Alignment
- Link: OpenReview
Video-LevelGauge: Investigating Contextual Positional Bias in Video Language Models.
- Link: OpenReview
Token Alignment Heads: Unveiling Attention's Role in LLM Multilingual Translation
- Link: OpenReview
DVLA-RL: Dual-Level Vision–Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
- Link: OpenReview
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
- Link: OpenReview
Privacy-Protected Causal Survival Analysis Under Distribution Shift
- Link: OpenReview
Reducing information dependency does not cause training data privacy. Adversarially non-robust features do.
- Link: OpenReview
Reliable Poisoned Sample Detection against Backdoor Attacks Enhanced by Sharpness Aware Minimization
- Link: OpenReview
GAVEL: Towards Rule-Based Safety through Activation Monitoring
- Link: OpenReview
Concept-Aware Privacy Mechanisms for Defending Embedding Inversion Attacks
- Link: OpenReview
TAO-Attack: Toward Advanced Optimization-Based Jailbreak Attacks for Large Language Models
- Link: OpenReview
Antibody: Strengthening Defense Against Harmful Fine-Tuning for Large Language Models via Attenuating Harmful Gradient Influence
- Link: OpenReview
Learning to Lie: Adversarial Attacks on Human-AI Teams and LLMs
- Link: OpenReview
Diffusion Blend: Inference-Time Multi-Preference Alignment for Diffusion Models
- Link: OpenReview
Scalable Energy-Based Models via Adversarial Training: Unifying Discrimination and Generation
- Link: OpenReview
ASGuard: Activation-Scaling Guard to Mitigate Targeted Jailbreaking Attack
- Link: OpenReview
A Statistical Learning Perspective on Semi-dual Adversarial Neural Optimal Transport Solvers
- Link: OpenReview
Time-To-Inconsistency: A Survival Analysis of Large Language Model Robustness to Adversarial Attacks
- Link: OpenReview
When Thinking Backfires: Mechanistic Insights into Reason-induced Misalignment
- Link: OpenReview
Vision Language Models are Biased
- Link: OpenReview
Watermarking Diffusion Language Models
- Link: OpenReview
SocialHarmBench: Revealing LLM Vulnerabilities to Socially Harmful Requests
- Link: OpenReview
MIDAS: Multi-Image Dispersion and Semantic Reconstruction for Jailbreaking MLLMs
- Link: OpenReview
Assessing Robustness via Score-Based Adversarial Image Generation
- Link: OpenReview
414. Towards Anomaly-Aware Pre-Training and Fine-Tuning for Graph Anomaly Detection
- Topics: LLMs & Foundation Models, Graph Neural Networks, Graphs & Combinatorial
FARI: Robust One-Step Inversion for Watermarking in Diffusion Models
- Link: OpenReview
From Evaluation to Defense: Advancing Safety in Video Large Language Models
- Link: OpenReview
STEDiff: Revealing the Spatial and Temporal Redundancy of Backdoor Attacks in Text-to-Image Diffusion Models
- Link: OpenReview
Safety Instincts: LLMs Learn to Trust Their Internal Compass for Self-Defense
- Link: OpenReview
Robust Adversarial Quantification via Conflict-Aware Evidential Deep Learning
- Link: OpenReview
Benchmarking Stochastic Approximation Algorithms for Fairness-Constrained Training of Deep Neural Networks
- Link: OpenReview
Mitigating the Safety Alignment Tax with Null-Space Constrained Policy Optimization
- Link: OpenReview
SlotGCG: Exploiting the Positional Vulnerability in LLMs for Jailbreak Attacks
- Link: OpenReview
BiasBusters: Uncovering and Mitigating Tool Selection Bias in Large Language Models
- Link: OpenReview
Fine-Grained Privacy Extraction from Retrieval-Augmented Generation Systems by Exploiting Knowledge Asymmetry
- Link: OpenReview
PluriHarms: Benchmarking the Full Spectrum of Human Judgments on AI Harm
- Link: OpenReview
WaterDrum: Watermark-based Data-centric Unlearning Metric
- Link: OpenReview
Efficient Adversarial Attacks on High-dimensional Offline Bandits
- Link: OpenReview
Learnability and Privacy Vulnerability are Entangled in a Few Critical Weights
- Link: OpenReview
Enabling Fine-Tuning of Direct Feedback Alignment via Feedback-Weight Matching
- Link: OpenReview
GradPCA: Leveraging NTK Alignment for Reliable Out-of-Distribution Detection
- Link: OpenReview
PALC: Preference Alignment via Logit Calibration
- Link: OpenReview
Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-Rankness
- Link: OpenReview
Tackling Heavy-Tailed Q-Value Bias in Offline-to-Online Reinforcement Learning with Laplace-Robust Modeling
- Link: OpenReview
Steerable Adversarial Scenario Generation through Test-Time Preference Alignment
- Link: OpenReview
Aligned Agents, Biased Swarm: Measuring Bias Amplification in Multi-Agent Systems
- Link: OpenReview
Bridging Past and Future: Distribution-Aware Alignment for Time Series Forecasting
- Link: OpenReview
Towards Multimodal Time Series Anomaly Detection with Semantic Alignment and Condensed Interaction
- Link: OpenReview
On-Policy RL Meets Off-Policy Experts: Harmonizing Supervised Fine-Tuning and Reinforcement Learning via Dynamic Weighting
- Link: OpenReview
From Natural Alignment to Conditional Controllability in Multimodal Dialogue
- Link: OpenReview
SELF-HARMONY: LEARNING TO HARMONIZE SELF-SUPERVISION AND SELF-PLAY IN TEST-TIME REINFORCEMENT LEARNING
- Link: OpenReview
Towards Understanding Valuable Preference Data for Large Language Model Alignment
- Link: OpenReview
Enforcing Axioms for AI Alignment under Loss-Based Rules
- Link: OpenReview
Representational Alignment Across Model Layers and Brain Regions with Multi-Level Optimal Transport
- Link: OpenReview
Enhancing Vision-Language Model with Unmasked Token Alignment
- Link: OpenReview
821. AB-UPT: Scaling Neural CFD Surrogates for High- Fidelity Automotive Aerodynamics Simulations via Anchored- Branched Universal Physics Transformers
- Topics: LLMs & Foundation Models
822. Seek-CAD: A Self-refined Generative Modeling for 3D Parametric CAD Using Local Inference via DeepSeek
- Topics: Diffusion Models & Generative AI, Computer Vision, Efficiency & Compression
DenseGRPO: From Sparse to Dense Reward for Flow Matching Model Alignment
- Link: OpenReview
Robust Preference Alignment via Directional Neighborhood Consensus
- Link: OpenReview
Adaptive Methods Are Preferable in High Privacy Settings: An SDE Perspective
- Link: OpenReview
How does the optimizer implicitly bias the model merging loss landscape?
- Link: OpenReview
Analyzing and Evaluating Unbiased Language Model Watermark
- Link: OpenReview
-DPO: Robust Preference Alignment for Diffusion Models via Divergence
- Link: OpenReview
NoisePrints: Distortion-Free Watermarks for Authorship in Private Diffusion Models
- Link: OpenReview
Adaptive Debiasing Tsallis Entropy for Test-Time Adaptation
- Link: OpenReview
Improving Semantic Proximity in Information Retrieval through Cross-Lingual Alignment
- Link: OpenReview
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement
- Link: OpenReview
NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting
- Link: OpenReview
SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense
- Link: OpenReview
Divide, Harmonize, Then Conquer It: Shooting Multi-Commodity Flow Problems with Multimodal Language Models
- Link: OpenReview
Harmonized Cone for Feasible and Non-conflict Directions in Training Physics-Informed Neural Networks
- Link: OpenReview
Hyperbolic Aware Minimization: Implicit Bias for Sparsity
- Link: OpenReview
Improving the Trade-off Between Watermark Strength and Speculative Sampling Efficiency for Language Models
- Link: OpenReview
Rethinking Bottlenecks in Safety Fine-Tuning of Vision Language Models
- Link: OpenReview
Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
- Link: OpenReview
Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks
- Link: OpenReview
ScaleCap: Scalable Image Captioning via Dual-Modality Debiasing
- Link: OpenReview
QPrompt-R1: Real-Time Reasoning for Domain-Generalized Semantic Segmentation via Group-Relative Query Alignment
- Link: OpenReview
Searching for Privacy Risks in LLM Agents via Simulation
- Link: OpenReview
SMOTE and Mirrors: Exposing Privacy Leakage from Synthetic Minority Oversampling
- Link: OpenReview
Towards Privacy-Guaranteed Label Unlearning in Vertical Federated Learning: Few-Shot Forgetting Without Disclosure
- Link: OpenReview
DeRaDiff: Denoising Time Realignment of Diffusion Models
- Link: OpenReview
Prediction with Expert Advice under Local Differential Privacy
- Link: OpenReview
Fairness via Independence: A General Regularization Framework for Machine Learning
- Link: OpenReview
PropensityBench: Evaluating Latent Safety Risks in Large Language Models via an Agentic Approach
- Link: OpenReview
GAS: Improving Discretization of Diffusion ODEs via Generalized Adversarial Solver
- Link: OpenReview
Annotation-Efficient Honesty Alignment via Confidence Elicitation and Calibration
- Link: OpenReview
Fairness-Aware Multi-view Evidential Learning with Adaptive Prior
- Link: OpenReview
Test-Time Poisoned Sample Detection by Exploiting Shallow Malicious Matching in Backdoored CLIP
- Link: OpenReview
On the Impossibility of Separating Intelligence from Judgment: The Computational Intractability of Filtering for AI Alignment
- Link: OpenReview
Towards Sequence Modeling Alignment between Tokenizer and Autoregressive Model
- Link: OpenReview
AEGIS: Adversarial Target-Guided Retention-Data-Free Robust Concept Erasure from Diffusion Models
- Link: OpenReview
Self-Jailbreaking: Language Models Can Reason Themselves Out of Safety Alignment After Benign Reasoning Training
- Link: OpenReview
CodeGenGuard: A Watermark for Code Generation Models
- Link: OpenReview
From Sure" to Sorry": Detecting Jailbreak in Large Vision Language Model via JailNeurons
- Link: OpenReview
On the Interaction of Compressibility and Adversarial Robustness
- Link: OpenReview
Automatic Dialectic Jailbreak: A Framework for Generating Effective Jailbreak Strategies
- Link: OpenReview
When Flatness Does (Not) Guarantee Adversarial Robustness
- Link: OpenReview
Beyond Match Maximization and Fairness: Retention-Optimized Two-Sided Matching
- Link: OpenReview
MCP-SafetyBench: A Benchmark for Safety Evaluation of Large Language Models with Real-World MCP Servers
- Link: OpenReview
STAR: Strategy-driven Automatic Jailbreak Red-teaming For Large Language Model
- Link: OpenReview
Fair Graph Machine Learning under Adversarial Missingness Processes
- Link: OpenReview
BiasScope: Towards Automated Detection of Bias in LLM-as-a-Judge Evaluation
- Link: OpenReview
Pretrain–Test Task Alignment Governs Generalization in In-Context Learning
- Link: OpenReview
BiasFreeBench: a Benchmark for Mitigating Bias in Large Language Model Responses
- Link: OpenReview
LLMS ON TRIAL: Evaluating Judicial Fairness For Large Language Models
- Link: OpenReview
Superficial Safety Alignment Hypothesis
- Link: OpenReview
The Devil behind the mask: An emergent safety vulnerability of Diffusion LLMs
- Link: OpenReview
Learning Adaptive Distribution Alignment with Neural Characteristic Function for Graph Domain Adaptation
- Link: OpenReview
Architecture-Agnostic Test-Time Adaptation via Backprop-Free Embedding Alignment
- Link: OpenReview
GRADIEND: Feature Learning within Neural Networks Exemplified through Biases
- Link: OpenReview
Spatial Forcing: Implicit Spatial Representation Alignment for Vision-language-action Model
- Link: OpenReview
Aligner, Diagnose Thyself: A Meta-Learning Paradigm for Fusing Intrinsic Feedback in Preference Alignment
- Link: OpenReview
Can SAEs reveal and mitigate racial biases of LLMs in healthcare?
- Link: OpenReview
Towards Cognitively-Faithful Decision-Making Models to Improve AI Alignment
- Link: OpenReview
When Greedy Wins: Emergent Exploitation Bias in Meta-Bandit LLM Training
- Link: OpenReview
Multi-objective Large Language Model Alignment with Hierarchical Experts
- Link: OpenReview
PAMDP: Interact to Persona Alignment via a Partially Observable Markov Decision Process
- Link: OpenReview
Breaking Safety Paradox with Feasible Dual Policy Iteration
- Link: OpenReview
PLANETALIGN: A Comprehensive Python Library for Benchmarking Network Alignment
- Link: OpenReview
Robust Spiking Neural Networks Against Adversarial Attacks
- Link: OpenReview
When Weak LLMs Speak with Confidence, Preference Alignment Gets Stronger
- Link: OpenReview
DistDF: Time-series Forecasting Needs Joint-distribution Wasserstein Alignment
- Link: OpenReview
ProSafePrune: Projected Safety Pruning for Mitigating Over-Refusal in LLMs
- Link: OpenReview
Auto-RT: Automatic Jailbreak Strategy Exploration for Red-Teaming Large Language Models
- Link: OpenReview
Meta-Learning Theory-Informed Inductive Biases using Deep Kernel Gaussian Processes
- Link: OpenReview
MindMix: A Multimodal Foundation Model for Auditory Perception Decoding via Deep Neural-Acoustic Alignment
- Link: OpenReview
Safety at One Shot: Patching Fine-Tuned LLMs with A Single Instance
- Link: OpenReview
Reasoned Safety Alignment: Ensuring Jailbreak Defense via Answer-Then-Check
- Link: OpenReview
Unified Vision–Language Modeling via Concept Space Alignment
- Link: OpenReview
Closing the Safety Gap: Surgical Concept Erasure in Visual Autoregressive Models
- Link: OpenReview
AMLRIS: Alignment-aware Masked Learning for Referring Image Segmentation
- Link: OpenReview
Transformers with Endogenous In-Context Learning: Bias Characterization and Mitigation
- Link: OpenReview
DeLiVR: Differential Spatiotemporal Lie Bias for Efficient Video Deraining
- Link: OpenReview
No, of Course I Can! Deeper Fine-Tuning Attacks That Bypass Token-Level Safety Mechanisms
- Link: OpenReview
Missingness Bias Calibration in Feature Attribution Explanations
- Link: OpenReview
Information Theoretic Guarantees For Policy Alignment In Large Language Models
- Link: OpenReview
1814. Adjusting Prediction Model Through Wasserstein Geodesic for Causal Inference
- Topics: Efficiency & Compression
Decoupling Primitive with Experts: Dynamic Feature Alignment for Compositional Zero-Shot Learning
- Link: OpenReview
ORION: Decoupling and Alignment for Unified Autoregressive Understanding and Generation
- Link: OpenReview
GAPrune: Gradient-Alignment Pruning for Domain-Aware Embeddings
- Link: OpenReview
Path Matters: Unveiling Geometric Implicit Bias via Curvature-Aware Sparse View Optimization
- Link: OpenReview
Two-Way Is Better Than One: Bidirectional Alignment with Cycle Consistency for Exemplar-Free Class-Incremental Learning
- Link: OpenReview
Reversible Primitive–Composition Alignment for Continual Vision–Language Learning
- Link: OpenReview
Distributional Vision-Language Alignment by Cauchy-Schwarz Divergence
- Link: OpenReview
Detecting Temporal Misalignment Attacks in Multimodal Fusion for Autonomous Driving
- Link: OpenReview
The Unseen Bias: How Norm Discrepancy in Pre-Norm MLLMs Leads to Visual Information Loss
- Link: OpenReview
Exponential-Wrapped Mechanisms: Differential Privacy on Hadamard Manifolds Made Practical
- Link: OpenReview
Membership Privacy Risks of Sharpness Aware Minimization
- Link: OpenReview
Unified Privacy Guarantees for Decentralized Learning via Matrix Factorization
- Link: OpenReview
PMark: Towards Robust and Distortion-free Semantic-level Watermarking with Channel Constraints
- Link: OpenReview
Beyond Membership: Limitations of Add/Remove Adjacency in Differential Privacy
- Link: OpenReview
AdvChain: Adversarial Chain-of-Thought Tuning for Robust Safety Alignment of Large Reasoning Models
- Link: OpenReview
Understanding Sensitivity of Differential Attention through the Lens of Adversarial Robustness
- Link: OpenReview
Defending against Backdoor Attacks via Module Switching
- Link: OpenReview
SecP-Tuning: Efficient Privacy-Preserving Prompt Tuning for Large Language Models via MPC
- Link: OpenReview
Structurally Human, Semantically Biased: Detecting LLM-Generated References with Embeddings and GNNs
- Link: OpenReview
RESFL: An Uncertainty-Aware Framework for Responsible Federated Learning by Balancing Privacy, Fairness and Utility
- Link: OpenReview
Online Conformal Prediction with Adversarial Semi-bandit Feedback via Regret Minimization
- Link: OpenReview
MUSE: Model-Agnostic Tabular Watermarking via Multi-Sample Selection
- Link: OpenReview
Safety Mirage: How Spurious Correlations Undermine VLM Safety Fine-Tuning and Can Be Mitigated by Machine Unlearning
- Link: OpenReview
DRIFT: Divergent Response in Filtered Transformations for Robust Adversarial Defense
- Link: OpenReview
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
- Link: OpenReview
Priors in time: Missing inductive biases for language model interpretability
- Link: OpenReview
OpenAgentSafety: A Comprehensive Framework For Evaluating Real-World AI Agent Safety
- Link: OpenReview
Bridging Fairness and Explainability: Can Input-Based Explanations Promote Fairness in Hate Speech Detection?
- Link: OpenReview
When Style Breaks Safety: Defending LLMs Against Superficial Style Alignment
- Link: OpenReview
VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety
- Link: OpenReview
CompMarkGS: Robust Watermarking for Compressed 3D Gaussian Splatting
- Link: OpenReview
Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs
- Link: OpenReview
ManagerBench: Evaluating the Safety-Pragmatism Trade-off in Autonomous LLMs
- Link: OpenReview
SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- Link: OpenReview
Robust Deep Reinforcement Learning against Adversarial Behavior Manipulation
- Link: OpenReview
SEMA: Simple yet Effective Learning for Multi-Turn Jailbreak Attacks
- Link: OpenReview
SoSBench: Benchmarking Safety Alignment on Six Scientific Domains
- Link: OpenReview
Closed-form norm scaling with data for overparameterized linear regression and diagonal linear networks under bias
- Link: OpenReview
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
- Link: OpenReview
Gradient Intrinsic Dimensionality Alignment:Narrowing The Gap Between Low-Rank Adaptation and Full Fine-Tuning
- Link: OpenReview
Unbiased Gradient Estimation for Event Binning via Functional Backpropagation
- Link: OpenReview
Adversarially Pretrained Transformers May Be Universally Robust In-Context Learners
- Link: OpenReview
Adaptive Data-Knowledge Alignment in Genetic Perturbation Prediction
- Link: OpenReview
Primal-Dual Policy Optimization for Linear CMDPs with Adversarial Losses
- Link: OpenReview
Reward Model Routing in Alignment
- Link: OpenReview
Improving Human-AI Coordination through Online Adversarial Training and Generative Models
- Link: OpenReview
Latent Wasserstein Adversarial Imitation Learning
- Link: OpenReview
Model Predictive Adversarial Imitation Learning for Planning from Observation
- Link: OpenReview
Watermark-based Attribution of AI-Generated Content
- Link: OpenReview
Measuring and Mitigating Rapport Bias of Large Language Models under Multi-Agent Social Interactions
- Link: OpenReview
Bures-Isotropy Alignment: Manifold Learning of Generalized Category Discovery
- Link: OpenReview
HarmonyGNNs: Harmonizing Heterophily and Homophily in GNNs via Self-Supervised Node Encoding
- Link: OpenReview
Learning to Summarize by Learning to Quiz: Adversarial Agentic Collaboration for Long Document Summarization
- Link: OpenReview
TwinFlow: Realizing One-step Generation on Large Models with Self-adversarial Flows
- Link: OpenReview
Representation Alignment for Diffusion Transformers without External Components
- Link: OpenReview
PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation Models
- Link: OpenReview
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
- Link: OpenReview
Tug-of-War No More: Harmonizing Accuracy and Robustness in Vision-Language Models via Stability-Aware Task Vector Merging
- Link: OpenReview
Inverse Reinforcement Learning with Dynamic Reward Scaling for LLM Alignment
- Link: OpenReview
Only Brains Align with Brains: Cross-Region Alignment Patterns Expose Limits of Normative Models
- Link: OpenReview
Low-Pass Filtering Improves Behavioral Alignment of Vision Models
- Link: OpenReview
VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
- Link: OpenReview
Debiased and Denoised Representation Learning for Incomplete Multi-view Clustering
- Link: OpenReview
NAIPv2: Debiased Pairwise Learning for Efficient Paper Quality Estimation
- Link: OpenReview
Routing Manifold Alignment Improves Generalization of Mixture-of-Experts LLMs
- Link: OpenReview
The Alignment Waltz: Jointly Training Agents to Collaborate for Safety
- Link: OpenReview
Generative Adversarial Post-Training Mitigates Reward Hacking in Live Human-AI Music Interaction
- Link: OpenReview
An Ensemble Framework for Unbiased Language Model Watermarking
- Link: OpenReview
Distilling the Thought, Watermarking the Answer: A Principle Semantic Guided Watermark for Reasoning Large Language Models
- Link: OpenReview
Debiased Front-Door Learners for Heterogeneous Effects
- Link: OpenReview
SeedVR2: One-Step Video Restoration via Diffusion Adversarial Post-Training
- Link: OpenReview
SIGMark: Scalable In-Generation Watermark with Blind Extraction for Video Diffusion
- Link: OpenReview
Reconstruction Alignment Improves Unified Multimodal Models
- Link: OpenReview
Relationship Alignment for View-aware Multi-view Clustering
- Link: OpenReview
UniCon: Unified Framework for Efficient Contrastive Alignment via Kernels
- Link: OpenReview
LCA: Local Classifier Alignment for Continual Learning
- Link: OpenReview
Towards One-step Causal Video Generation via Adversarial Self-Distillation
- Link: OpenReview
HiTeA: Hierarchical Temporal Alignment for Training-Free Long-Video Temporal Grounding
- Link: OpenReview
Shuffling the Data, Extrapolating the Step: Sharper Bias In Constant Step-Size SGD
- Link: OpenReview
Inconsistency Biases in Dynamic Data Pruning
- Link: OpenReview
IA2: Alignment with ICL Activations improves Supervised Fine-Tuning
- Link: OpenReview
Semantic-aware Wasserstein Policy Regularization for Large Language Model Alignment
- Link: OpenReview
Operationalizing Data Minimization for Privacy-Preserving LLM Prompting
- Link: OpenReview
Natural Identifiers for Privacy and Data Audits in Large Language Models
- Link: OpenReview
Mitigating Privacy Risk via Forget Set-Free Unlearning
- Link: OpenReview
Winter Soldier: Backdooring Language Models at Pre-Training with Indirect Data Poisoning
- Link: OpenReview
INO-SGD: Addressing Utility Imbalance under Individualized Differential Privacy
- Link: OpenReview
Convergent Differential Privacy Analysis for General Federated Learning
- Link: OpenReview
Don't Shift the Trigger: Robust Gradient Ascent for Backdoor Unlearning
- Link: OpenReview
Adversarial Déjà Vu: Jailbreak Dictionary Learning for Stronger Generalization to Unseen Attacks
- Link: OpenReview
Inference-Time Personalized Safety Control via Paired Difference-in-Means Intervention
- Link: OpenReview
Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean Reference
- Link: OpenReview
JULI: Jailbreak Large Language Models by Self-Introspection
- Link: OpenReview
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
- Link: OpenReview
Diffusion & Adversarial Schrödinger Bridges via Iterative Proportional Markovian Fitting
- Link: OpenReview
Doxing via the Lens: Revealing Location-related Privacy Leakage on Multi-modal Large Reasoning Models
- Link: OpenReview
On the trade-off between expressivity and privacy in graph representation learning
- Link: OpenReview
Jailbreaking on Text-to-Video Models via Scene Splitting Strategy
- Link: OpenReview
GuidedBench: Measuring and Mitigating the Evaluation Discrepancies of In-the-wild LLM Jailbreak Methods
- Link: OpenReview
Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMs
- Link: OpenReview
GraphShield: Graph-Theoretic Modeling of Network-Level Dynamics for Robust Jailbreak Detection
- Link: OpenReview
Jailbreaking the Matrix: Nullspace Steering for Controlled Model Subversion
- Link: OpenReview
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
- Link: OpenReview
Robust Adversarial Attacks Against Unknown Disturbance via Inverse Gradient Sample
- Link: OpenReview
Fine-Grained Class-Conditional Distribution Balancing for Debiased Learning
- Link: OpenReview
Doubly-Regressing Approach for Subgroup Fairness
- Link: OpenReview
Zero-Sacrifice Persistent-Robustness Adversarial Defense for Pre-Trained Encoders
- Link: OpenReview
Any-Depth Alignment: Unlocking Innate Safety Alignment of LLMs to Any-Depth
- Link: OpenReview
Safety Subspaces are Not Linearly Distinct: A Fine-Tuning Case Study
- Link: OpenReview
Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle Escape
- Link: OpenReview
Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network Architectures
- Link: OpenReview
When Bias Meets Trainability: Connecting Theories of Initialization
- Link: OpenReview
Property-Driven Protein Inverse Folding with Multi-Objective Preference Alignment
- Link: OpenReview
Fast and Interpretable Protein Substructure Alignment via Optimal Transport
- Link: OpenReview
RLAC: Reinforcement Learning with Adversarial Critic for Free-Form Generation Tasks
- Link: OpenReview
AsyncBEV: Cross-modal flow alignment in Asynchronous 3D Object Detection
- Link: OpenReview
Mirage or Method? How Model–Task Alignment Induces Divergent RL Conclusions
- Link: OpenReview
ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
- Link: OpenReview
Understanding the Implicit Biases of Design Choices for Time Series Foundation Models
- Link: OpenReview
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving
- Link: OpenReview
Defending Against Unknown Corrupted Agents: Reinforcement Learning of Adversarially Robust Nash Equilibria
- Link: OpenReview
3399. SciTS: Scientific Time Series Understanding and Generation with LLMs
- Topics: LLMs & Foundation Models, Graphs & Structured Data
Verification and Co-Alignment via Heterogeneous Consistency for Preference-Aligned LLM Annotations
- Link: OpenReview
RE-PO: Robust Enhanced Policy Optimization as a General Framework for LLM Alignment
- Link: OpenReview
Generative Adversarial Reasoner: Enhancing LLM Reasoning with Adversarial Reinforcement Learning
- Link: OpenReview
What matters for Representation Alignment: Global Information or Spatial Structure?
- Link: OpenReview
MATRIX: Mask Track Alignment for Interaction-aware Video Generation
- Link: OpenReview
Discrete Diffusion Trajectory Alignment via Stepwise Decomposition
- Link: OpenReview
The Matthew Effect of AI Programming Assistants: A Hidden Bias in Software Evolution
- Link: OpenReview
When Language Models Lose Their Mind: The Consequences of Brain Misalignment
- Link: OpenReview
Alignment through Meta-Weighted Online Sampling: Bridging the Gap between Data Generation and Preference Optimization
- Link: OpenReview
JailNewsBench: Multi-Lingual and Regional Benchmark for Fake News Generation under Jailbreak Attacks
- Link: OpenReview
AdPO: Enhancing the Adversarial Robustness of Large Vision-Language Models with Preference Optimization
- Link: OpenReview
SAGA: Structural Aggregation Guided Alignment with Dynamic View and Neighborhood Order Selection for Multiview Graph Domain Adaptation
- Link: OpenReview
Humanline: Online Alignment as Perceptual Loss
- Link: OpenReview
DecAlign: Hierarchical Cross-Modal Alignment for Decoupled Multimodal Representation Learning
- Link: OpenReview
Understanding and Improving Continuous LLM Adversarial Training via In-context Learning Theory
- Link: OpenReview
Dual-Branch Representations with Dynamic Gated Fusion and Triple-Granularity Alignment for Deep Multi-View Clustering
- Link: OpenReview
Beyond Instance-Level Alignment: Dual-Level Optimal Transport for Audio-Text Retrieval
- Link: OpenReview
On the Alignment Between Supervised and Self-Supervised Contrastive Learning
- Link: OpenReview
Unbiased Object Detection Beyond Frequency with Visually Prompted Image Synthesis
- Link: OpenReview
MoAlign: Motion-Centric Representation Alignment for Video Diffusion Models
- Link: OpenReview
Verification of the Implicit World Model in a Generative Model via Adversarial Sequences
- Link: OpenReview
Expert Merging: Model Merging with Unsupervised Expert Alignment and Importance-Guided Layer Chunking
- Link: OpenReview
CHROMA: Consistent Harmonization of Multi-View Appearance via Bilateral Grid Prediction
- Link: OpenReview
Scaling Direct Feedback Learning with Jacobian Alignment Guarantees
- Link: OpenReview
Deep Latent Variable Model based Vertical Federated Learning with Flexible Alignment and Labeling Scenarios
- Link: OpenReview
Disentangling Length Bias in Preference Learning via Response-Conditioned Modeling
- Link: OpenReview
NatADiff: Adversarial Boundary Guidance for Natural Adversarial Diffusion
- Link: OpenReview
Noisy-Pair Robust Representation Alignment for Positive-Unlabeled Learning
- Link: OpenReview
EAMET: ROBUST MASSIVE MODEL EDITING VIA EMBEDDING ALIGNMENT OPTIMIZATION
- Link: OpenReview
Federated Learning of Quantile Inference under Local Differential Privacy
- Link: OpenReview
Flattery, Fluff, and Fog: Diagnosing and Mitigating Idiosyncratic Biases in Preference Models
- Link: OpenReview
Dual-Path Condition Alignment for Diffusion Transformers
- Link: OpenReview
Transferable and Stealthy Adversarial Attacks on Large Vision-Language Models
- Link: OpenReview
Identifying Robust Neural Pathways: Few-Shot Adversarial Mask Tuning for Vision-Language Models
- Link: OpenReview
TriQDef: Disrupting Semantic and Gradient Alignment to Prevent Adversarial Patch Transferability in Quantized Neural Networks
- Link: OpenReview
Measuring Physical-World Privacy Awareness of Large Language Models: An Evaluation Benchmark
- Link: OpenReview
JALMBench: Benchmarking Jailbreak Vulnerabilities in Audio Language Models
- Link: OpenReview
On Fairness of Task Arithmetic: The Role of Task Vectors
- Link: OpenReview
CheckMate! Watermarking Graph Diffusion Models in Polynomial Time
- Link: OpenReview
Disrupting Hierarchical Reasoning: Adversarial Protection for Geographic Privacy in Multimodal Reasoning Models
- Link: OpenReview
PLAGUE: Plug-and-play framework for Lifelong Adaptive Generation of mUlti-turn jailbrEaks
- Link: OpenReview
Adversarial Robustness of Graph Transformers
- Link: OpenReview
4040. Explainable Mixture Models through Differentiable Rule Learning
- Topics: Interpretability & Mechanistic Interpretability
Sampling-aware Adversarial Attacks Against Large Language Models
- Link: OpenReview
WRING Out The Bias: A Rotation-Based Alternative To Projection Debiasing
- Link: OpenReview
CERTIFIED VS. EMPIRICAL ADVERSARIAL ROBUSTNESS VIA HYBRID CONVOLUTIONS WITH ATTENTION STOCHASTICITY
- Link: OpenReview
Reward Models Inherit Value Biases from Pretraining
- Link: OpenReview
Why Adversarially Train Diffusion Models?
- Link: OpenReview
FERD: Fairness-Enhanced Data-Free Adversarial Robustness Distillation
- Link: OpenReview
Are Deep Speech Denoising Models Robust to Adversarial Noise?
- Link: OpenReview
Concept-based Adversarial Attack: a Probabilistic Perspective
- Link: OpenReview
Fine-Grained Iterative Adversarial Attacks with Limited Computation Budget
- Link: OpenReview
THE SELF-RE-WATERMARKING TRAP: FROM EXPLOIT TO RESILIENCE
- Link: OpenReview
Implicit bias produces neural scaling laws in learning curves, from perceptrons to deep networks
- Link: OpenReview
Align to Misalign: Automatic LLM Jailbreak with Meta-Optimized LLM Judges
- Link: OpenReview
Regulating Internal Alignment Flows for Robust Learning Under Spurious Correlations
- Link: OpenReview
Risk Phase Transitions in Spiked Regression: Alignment Driven Benign and Catastrophic Overfitting
- Link: OpenReview
JailbreakLoRA: Your Downloaded LoRA from Sharing Platforms might be Unsafe
- Link: OpenReview
Reasoning Boosts Opinion Alignment in LLMs
- Link: OpenReview
Bandit Learning in Matching Markets Robust to Adversarial Corruptions
- Link: OpenReview
Implicit Bias of Per-sample Adam on Separable Data: Departure from the Full-batch Regime
- Link: OpenReview
Minor First, Major Last: A Depth-Induced Implicit Bias of Sharpness-Aware Minimization
- Link: OpenReview
GeneBreaker: Jailbreak Attacks against DNA Language Models with Pathogenicity Guidance
- Link: OpenReview
Moving Beyond Medical Exams: A Clinician-Annotated Fairness Dataset of Real-World Tasks and Ambiguity in Mental Healthcare
- Link: OpenReview
Minimax Optimal Adversarial Reinforcement Learning
- Link: OpenReview
sleep2vec: Unified Cross-Modal Alignment for Heterogeneous Nocturnal Biosignals
- Link: OpenReview
Keep the Best, Forget the Rest: Reliable Alignment with Order-Aware Preference Optimization
- Link: OpenReview
ComPhy: Composing Physical Models with end-to-end Alignment
- Link: OpenReview
Test-Time Alignment for Large Language Models via Textual Model Predictive Control
- Link: OpenReview
Data Selection for LLM Alignment Using Fine-Grained Preferences
- Link: OpenReview
On the Tension Between Optimality and Adversarial Robustness in Policy Optimization
- Link: OpenReview
Learning From Dictionary: Enhancing Robustness of Machine-Generated Text Detection in Zero-Shot Language via Adversarial Training
- Link: OpenReview
Align Once, Benefit Multilingually: Enforcing Multilingual Consistency for LLM Safety Alignment
- Link: OpenReview
Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
- Link: OpenReview
Beyond Markovian Drifts: Action-Biased Geometric Walks with Memory for Personalized Summarization
- Link: OpenReview
In-Context Watermarks for Large Language Models
- Link: OpenReview
Jailbreak Transferability Emerges from Shared Representations
- Link: OpenReview
Improved Adversarial Diffusion Compression for Real-World Video Super-Resolution
- Link: OpenReview
Near-Optimal Second-Order Guarantees for Model-Based Adversarial Imitation Learning
- Link: OpenReview
Discovering alternative solutions beyond the simplicity bias in recurrent neural networks
- Link: OpenReview
References Improve LLM Alignment in Non-Verifiable Domains
- Link: OpenReview
Entropy-Guided Dynamic Tokens for Graph-LLM Alignment in Molecular Understanding
- Link: OpenReview
BideDPO: Conditional Image Generation with Simultaneous Text and Condition Alignment
- Link: OpenReview
Persona Features Control Emergent Misalignment
- Link: OpenReview
SAIL: Self-Amplified Iterative Learning for Diffusion Model Alignment with Minimal Human Feedback
- Link: OpenReview
Nasty Adversarial Training: A Probability Sparsity Perspective for Robustness Enhancement
- Link: OpenReview
Obscure but Effective: Classical Chinese Jailbreak Prompt Optimization via Bio-Inspired Search
- Link: OpenReview
The Hot Mess of AI: How Does Misalignment Scale With Model Intelligence and Task Complexity?
- Link: OpenReview
Modality Alignment across Trees on Heterogeneous Hyperbolic Manifolds
- Link: OpenReview
TIGaussian: Disentangle Gaussians for Spatial-Awared Text-Image-3D Alignment
- Link: OpenReview
Adversarial Encoding Perturbation and Synthesis for Set Representation Auxiliary Learning
- Link: OpenReview
DR-Submodular Maximization with Stochastic Biased Gradients: Classical and Quantum Gradient Algorithms
- Link: OpenReview
FedMuon: Federated Learning with Bias-corrected LMO-based Optimization
- Link: OpenReview
SCAD: Super-Class-Aware Debiasing for Long-Tailed Semi-Supervised Learning
- Link: OpenReview
Diffusion Alignment as Variational Expectation-Maximization
- Link: OpenReview
Rethinking LoRA for Privacy-Preserving Federated Learning in Large Models
- Link: OpenReview
Improved Object-Centric Diffusion Learning with Registers and Contrastive Alignment
- Link: OpenReview
Traceable Black-Box Watermarks For Federated Learning
- Link: OpenReview
SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt Learning
- Link: OpenReview
Alignment-Weighted DPO: A principled reasoning approach to improve safety alignment
- Link: OpenReview
A Guardrail for Safety Preservation: When Safety-Sensitive Subspace Meets Harmful-Resistant Null-Space
- Link: OpenReview
Guidance Watermarking for Diffusion Models
- Link: OpenReview
Dissecting Representation Misalignment in Contrastive Learning via Influence Function
- Link: OpenReview
Discrete Latent Features Ablate Adversarial Attack: A Robust Prompt Tuning Framework for VLMs
- Link: OpenReview
Constitutional Classifiers++: Efficient Production-Grade Defenses against Universal Jailbreaks
- Link: OpenReview
DiffuGuard: How Intrinsic Safety is Lost and Found in Diffusion Large Language Models
- Link: OpenReview
Eliciting Harmful Capabilities by Fine-Tuning on Safeguarded Outputs
- Link: OpenReview
Adaptive Logit Adjustment for Debiasing Multimodal Language Models
- Link: OpenReview
Adversarial Attacks Already Tell the Answer: Directional Bias-Guided Test-time Defense for Vision-Language Models
- Link: OpenReview
FLoRG: Federated Fine-tuning with Low-rank Gram Matrices and Procrustes Alignment
- Link: OpenReview
A2D: Any-Order, Any-Step Safety Alignment for Diffusion Language Models
- Link: OpenReview
Towards Understanding Subliminal Learning: When and How Hidden Biases Transfer
- Link: OpenReview
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
- Link: OpenReview
Misaligned Roles, Misplaced Images: Structural Input Perturbations Expose Multimodal Alignment Blind Spots
- Link: OpenReview
Harnessing Hyperbolic Geometry for Harmful Prompt Detection and Sanitization
- Link: OpenReview
Bias Similarity Measurement: A Black-Box Audit of Fairness Across LLMs
- Link: OpenReview
Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- Link: OpenReview
The First Impression Problem: Internal Bias Triggers Overthinking in Reasoning Models
- Link: OpenReview
Beyond Pairwise: Empowering LLM Alignment With (Ranked) Choice Modeling
- Link: OpenReview
DecompGAIL: Learning Realistic Traffic Behaviors with Decomposed Multi-Agent Generative Adversarial Imitation Learning
- Link: OpenReview
Remotely Detectable Robot Policy Watermarking
- Link: OpenReview
Unlocking the Value of Text: Event-Driven Reasoning and Multi-Level Alignment for Time Series Forecasting
- Link: OpenReview
ATGen: Adversarial Reinforcement Learning for Test Case Generation
- Link: OpenReview
IDEAL: Data Equilibrium Adaptation for Multi-Capability Language Model Alignment
- Link: OpenReview
Fluent Alignment with Disfluent Judges: Post-training for lower-resource languages
- Link: OpenReview
Is On-Policy Data always the Best Choice for Direct Preference Optimization-Based LM Alignment?
- Link: OpenReview
Learning to Generate Unit Test via Adversarial Reinforcement Learning
- Link: OpenReview
OrthAlign: Orthogonal Subspace Decomposition for Non-Interfering Multi-Objective Alignment
- Link: OpenReview
Emergent Misalignment is Easy, Narrow Misalignment is Hard
- Link: OpenReview
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
- Link: OpenReview
GuardAlign: Test-time Safety Alignment in Multimodal Large Language Models
- Link: OpenReview
Evaluating and Improving Cultural Awareness of Reward Models for LLM Alignment
- Link: OpenReview
The Mind's Transformer: Computational Neuroanatomy of LLM-Brain Alignment
- Link: OpenReview
Let's Think in Two Steps: Mitigating Agreement Bias in MLLMs with Self-Grounded Verification
- Link: OpenReview
Multi-Marginal Flow Matching with Adversarially Learnt Interpolants
- Link: OpenReview
Eliminating Inductive Bias in Reward Models with Information-Theoretic Guidance
- Link: OpenReview
Optimizing Canaries for Privacy Auditing with Metagradient Descent
- Link: OpenReview
AlphaAlign: Incentivizing Safety Alignment with Extremely Simplified Reinforcement Learning
- Link: OpenReview
Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to Models
- Link: OpenReview
An Improved Model-free Decision-estimation Coefficient with Applications in Adversarial MDPs
- Link: OpenReview
It's All Connected: A Journey Through Test-Time Memorization, Attentional Bias, Retention, and Online Optimization
- Link: OpenReview
Feature compression is the root cause of adversarial fragility in neural networks
- Link: OpenReview
Deconstructing Positional Information: From Attention Logits to Training Biases
- Link: OpenReview
From Gradient Volume to Shapley Fairness: Towards Fair Multi-Task Learning
- Link: OpenReview