R
Published on

ACL 2026 — Information Retrieval & Knowledge

Information Retrieval & Knowledge

385 papers Links not yet available — ACL proceedings pending on ACL Anthology.

  • EASE: Entity-Aware Sub-table Generation for Real-world Multi-table QA
  • Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
  • CLEAR: Cross-Lingual Enhancement in Retrieval via Reverse-training
  • Aligning Language Models with Real-time Knowledge Editing
  • MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail Knowledge
  • IntrAgent: An LLM Agent for Content-Grounded Information Retrieval through Literature Review
  • AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering
  • FastV-RAG: Towards Fast and Fine-Grained Video QA with Retrieval-Augmented Generation
  • MemSearch-o1: Empowering Large Language Models with Reasoning-Aligned Memory Growth in Agentic Search
  • ReasonEmbed: Enhanced Text Embeddings for Reasoning-Intensive Document Retrieval
  • PrefixNLI: Detecting Factual Inconsistencies as Soon as They Arise
  • RAM-SD: Retrieval-Augmented Multi-agent framework for Sarcasm Detection
  • SAGE: A Search-AuGmented Evaluation of Large Language Models on Free‑Form QA
  • Grammar Search for Multi-Agent Systems
  • ZoomR: Memory Efficient Reasoning through Multi-Granularity Key Value Retrieval
  • Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering
  • CHIMERA: A Knowledge Base of Scientific Idea Recombinations for Research Analysis and Ideation
  • Erasing Without Remembering: Implicit Knowledge Forgetting in Large Language Models
  • AraVQA: Building a New Arabic Factoid Visual Question Answering Dataset from Wikipedia
  • When Misinformation Speaks and Converses: Rethinking Fact-Checking in Audio Platforms
  • Does Self-Consistency Improve the Recall of Encyclopedic Knowledge?
  • EquivPruner: Boosting Efficiency and Quality in LLM-Based Search via Action Pruning
  • ChatR1: Reinforcement Learning for Conversational Reasoning and Retrieval Augmented Question Answering
  • MMSearch-R1: Incentivizing LMMs to Search
  • Thinking beyond the anthropomorphic paradigm benefits LLM research
  • Coordinating Search-Informed Reasoning and Reasoning-Guided Search in Claim Verification
  • EMCEE: Improving Multilingual Capability of LLMs via Bridging Knowledge and Reasoning with Extracted Synthetic Multilingual Context
  • Situated Embedding Models for Context-Aware Dense Retrieval
  • CiPO: Counterfactual Unlearning for Large Reasoning Models through Iterative Preference Optimization
  • Towards Scalable Lifelong Knowledge Editing with Selective Knowledge Suppression
  • Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders
  • CRAFT: Training-Free Cascaded Retrieval for Tabular QA
  • KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation
  • Mitigating Context Interference for Reliable and Efficient Search Agents
  • Long Context Modeling with Ranked Memory-Augmented Retrieval
  • Collaborative Chain-of-Agents for Parametric-Retrieved Knowledge Synergy
  • Bringing Real-World Relations into Video Generation with Graph-Structured Knowledge
  • Mask-to-Correct⁺: Leveraging Retriever Diversity for Masking-guided Faithful Fact Correction
  • From Factuality to Meta-Factivity: A Cognitive Blueprint for Trustworthy LLMs
  • UniversalRAG: Retrieval-Augmented Generation over Corpora of Diverse Modalities and Granularities
  • Disco-RAG: Discourse-Aware Retrieval-Augmented Generation
  • Guaranteeing Knowledge Integration with Joint Decoding for Retrieval-Augmented Generation
  • Retrieval as Generation: A Unified Framework with Self-Triggered Information Planning
  • Deliberative Searcher: Improving LLM Reliability via Reinforcement Learning with Constraints
  • REFLEX: Self-Refining Explainable Fact-Checking via Verdict-Anchored Style Control
  • Scaling Beyond Context: A Survey of Multimodal Retrieval-Augmented Generation for Document Understanding
  • Open Schrödinger’s Closed Box: Identifying Retrieval Augmented Generation in API-Accessible Large Language Model Services
  • D²Plan: Dual-Agent Dynamic Global Planning for Complex Retrieval-Augmented Reasoning
  • Enhancing Lexical Relation Mining with Structured Sememe Knowledge
  • Retrievals Can Be Detrimental: Unveiling the Backdoor Vulnerability of Retrieval-Augmented Diffusion Models
  • DPDV: Dual-Pathway and Dual-View Representation Learning for Bridging Information Asymmetry in Text-Video Retrieval
  • Representation Interventions Enable Lifelong Knowledge Memory Control in LLMs
  • LogicPoison: Logical Attacks on Graph Retrieval-Augmented Generation
  • Incentivizing Parametric Knowledge via Reinforcement Learning with Verifiable Rewards for Cross-Cultural Entity Translation
  • Compete to Complete: Co-opetition Adversarial Learning for Retrieval-Augmented Generation
  • FinSight: Towards Real-World Financial Deep Research
  • CiteGuard: Faithful Citation Attribution for LLMs via Retrieval-Augmented Validation
  • UNIKIE-BENCH: Benchmarking Large Multimodal Models for Key Information Extraction in Visual Documents
  • FS-Researcher: Test-Time Scaling for Long-Horizon Research Tasks with File-System-Based Agents
  • StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt Generation
  • Efficient Multi-Agent System Training with Data Influence-Oriented Tree Search
  • DeReA: Improving Idiom Translation with Detect-Retrieve-Arbitrate Reasoning
  • Don’t Be Misled by Style: A Style-Adaptive Reranker for Capturing Effective Knowledge in Retrieval-Augmented Generation
  • Text2Tabular – Reconstructing Tabular Research Data from Scientific Publications
  • In-depth Research Impact Summarization through Fine-Grained Temporal Citation Analysis
  • LAFaCT: Attribution-based Localization and Focused Sequential Analysis of Fact-Critical Tokens for Hallucination Detection
  • STEM: Structure-Tracing Evidence Mining for Knowledge Graphs-Driven Retrieval-Augmented Generation
  • LePREC: Reasoning as Classification over Structured Factors for Assessing Relevance of Legal Issues
  • AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering
  • From Completion to Editing: Unlocking Context-Aware Code Infilling via Search-and-Replace Instruction Tuning
  • ARK: Answer-Centric Retriever Tuning via KG-augmented Curriculum Learning
  • A Survey of Large Language Model-Based Search Agents
  • InferenceDynamics: Adaptive LLM Routing through Structured Capability and Knowledge Profiling
  • ReportLogic: Evaluating Logical Quality in Deep Research Reports
  • MulVul: Retrieval-augmented Multi-Agent Code Vulnerability Detection via Cross-Model Prompt Evolution
  • Visual Attention Reasoning via Hierarchical Search and Self-Verification
  • KCVR: Knowledge-Centric Video Reconstruction for Structured Pedagogical Summarization via Dynamic Graph Planning
  • TRACE: Traversal Retrieval-Augmented Chain of Evidence for Document Understanding
  • FinKario: Event-Enhanced Automated Construction of Financial Knowledge Graph
  • DREAM: Deep Research Evaluation with Agentic Metrics
  • GKnow: Measuring the Entanglement of Gender Bias and Factual Gender
  • Reinforcing Agentic Search Via Reward Density Optimization
  • Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration
  • OCR-Memory: Optical Context Retrieval for Long-Horizon Agent Memory
  • Masked by Consensus: Disentangling Privileged Knowledge in LLM Correctness
  • Reasoning with Ontology Graph: Toward Type-Constrained Knowledge Graph Question Answering
  • All That Glisters Is Not Gold: A Benchmark for Reference-Free Counterfactual Financial Misinformation Detection
  • ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification
  • Learning from Contrasts: Synthesizing Reasoning Paths from Diverse Search Trajectories
  • Improving Retrieval-Augmented Generation without Taxonomy-based Error Categorization
  • CoQuIR: A Comprehensive Benchmark for Code Quality-Aware Information Retrieval
  • Lifting Optimized Binaries to Canonical Compiler IR via Structure-Aware Retrieval and Iterative Verification
  • Adaptive Prompt Structure Factorization: A Framework for Self-Discovering and Optimizing Compositional Prompt Programs
  • Exploring Layer Activation Dynamic of CoT via Knowledge Probe
  • LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection
  • Knowledge is Not Enough: Injecting RL Skills for Continual Adaptation
  • Table Question Answering in the Era of Large Language Models: A Comprehensive Survey of Tasks, Methods, and Evaluation
  • Capability Decomposition for Unified Information Extraction via Hierarchical Mixture-of-Experts
  • A Scalable Multi-LLM Collaboration System with Retrieval-based Selection and Exploration-Exploitation-Driven Enhancement
  • Making Large Language Models Efficient Dense Retrievers
  • Rethinking Table Pruning in TableQA: From Sequential Revisions to Gold Trajectory-Supervised Parallel Search
  • RoboFailRing: Retrieval-Augmented and Language Grounding Failure Detection for VLM-enabled Robotic Manipulation
  • DORA: A Dual-Objective Reinforcement Learning Framework for Effective and Efficient Multimodal Agentic Search
  • Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
  • Reinforced Informativeness Optimization for Long-Form Retrieval-Augmented Generation
  • Can Factual Opinions Be Edited (Manipulated) in Large Language Models?
  • How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve Them
  • When Seeing Overrides Knowing: Disentangling Knowledge Conflicts in Vision-Language Models
  • When Identity Skews Debate: Anonymization for Bias-Reduced Multi-Agent Reasoning
  • ReFEree: Reference-Free and Fine-Grained Method for Evaluating Factual Consistency in Real-World Code Summarization
  • SARA: Selective and Adaptive Retrieval-augmented Generation with Context Compression
  • MATCH: Modulating Attention via In‑Context Retrieval for Long‑Context Transformers
  • LiveCLKTBench: Towards Reliable Evaluation of Cross-Lingual Knowledge Transfer in Multilingual LLMs
  • CoG: Controllable Graph Reasoning via Relational Blueprints and Failure-Aware Refinement over Knowledge Graphs
  • Autonomous Knowledge Graph Exploration with Adaptive Breadth-Depth Retrieval
  • Retrieval Heads are Dynamic
  • Model Internal Sleuthing: Finding Lexical Identity and Inflectional Features in Modern Language Models
  • DMRetriever: A Family of Models for Improved Text Retrieval in Disaster Management
  • Language Models Don’t Know What You Want: Evaluating Personalization in Deep Research Needs Real Users
  • MediEval: A Unified Medical Benchmark for Patient-Contextual and Knowledge-Grounded Reasoning in LLMs
  • Biomedical Question Answering via Multi-Level Summarization on a Local Knowledge Graph
  • RExBench: Can coding agents autonomously implement AI research extensions?
  • Text-Attributed Knowledge Graph Enrichment with Large Language Models for Medical Concept Representation
  • ViDoRe V3: A Comprehensive Evaluation of Retrieval Augmented Generation in Complex Real-World Scenarios
  • LOKA: Conflict-Aware LLM Knowledge Update with Adaptive Knowledge Memory
  • PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering
  • GS-Quant: Granular Semantic and Generative Structural Quantization for Knowledge Graph Completion
  • Bidirectional LMs are Better Knowledge Memorizers? A Benchmark for Real-world Knowledge Injection
  • Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering
  • CaseFacts: A Benchmark for Legal Fact-Checking and Precedent Retrieval
  • SpecAgent: A Speculative Retrieval and Forecasting Agent for Code Completion
  • DiZiNER: Disagreement-guided Instruction Refinement via Simulating Pilot Annotation for Zero-shot Named Entity Recognition
  • PaperRegister: Boosting Flexible-grained Paper Search via Hierarchical Register Indexing
  • LCR-RAG: Enhancing Logical Consistency in Retrieval-Augmented Generation via Neuro-symbolic Reinforcement Learning
  • HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering
  • Improving Autoformalization Using Direct Dependency Retrieval
  • Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering
  • Soft Orthogonal Low-Rank Adaptation for Knowledge Sharing in Large Language Model Continual Learning
  • ChartAgent: A Multimodal Agent for Visually Grounded Reasoning in Complex Chart Question Answering
  • Adaptive Planning for Multi-Attribute Controllable Summarization with Monte Carlo Tree Search
  • SearchGym: Bootstrapping Real-World Search Agents via Cost-Effective and High-Fidelity Environment Simulation
  • Learning to Select: Query-Aware Adaptive Dimension Selection for Dense Retrieval
  • REG: Retrieval via Emotion Similarity for Guiding Empathetic Dialogue Generation
  • Know the Known and the Unknown: Reasonable Answer Generation with Knowledge-Informed Citations
  • FACTrial: Factorized Clinical Contrastive Training for Scalable Patient-Trial Retrieval
  • Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism
  • Progressive Multimodal Search and Reasoning for Knowledge-Intensive Visual Question Answering
  • Identity-Robust Language Model Generation via Content Integrity Preservation
  • Knowledge Poisoning Attacks on Medical Multi-Modal Retrieval-Augmented Generation
  • NOSE: Neural Olfactory-Semantic Embedding with Tri-Modal Orthogonal Contrastive Learning
  • HiGoE: Hierarchical Graph of Evidence to Enhance Retrieval-Augmented Generation for Long-context Summarization
  • STK-Adapter: Incorporating Evolving Graph and Event Chain for Temporal Knowledge Graph Extrapolation
  • Why LLMs Hallucinate on Structured Knowledge: A Mechanistic Analysis of Reasoning over Linearized Representations
  • Benchmarking and Enabling Efficient Chinese Medical Retrieval via Asymmetric Encoders
  • CAKE: Causal-Guided Adaptive Knowledge Editing for LLMs
  • RARE: Redundancy-Aware Retrieval Evaluation Framework for High-Similarity Corpora
  • Zero-Shot Multimodal Retrieval with Multi-Scale Contextual Representations
  • THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA
  • R^3AG: Retriever Routing for Retrieval-Augmented Generation
  • ClusterRAG: Cluster-Based Collaborative Filtering for Personalized Retrieval-Augmented Generation
  • Toward Robust In-Context Learning: Leveraging Out-of-distribution Proxies for Target Inaccessible Demonstration Retrieval
  • AutoSchemaKG: Autonomous Knowledge Graph Construction through Dynamic Schema Induction from Web-Scale Corpora
  • Unveiling the Unknown: Open-Set Entity Typing via Two-Stage Generation
  • Chaining the Evidence: Robust Reinforcement Learning for Deep Search Agents with Citation-Aware Rubric Rewards
  • Language on Demand, Knowledge at Core: Composing LLMs with Encoder-Decoder Translation Models for Extensible Multilinguality
  • Fisher-Driven Adaptive Locating for Knowledge Editing in Large Language Models
  • Retrieving to Recover: Towards Incomplete Audio-Visual Question Answering via Semantic-consistent Purification
  • FlowSearch: Advancing Deep Research with Dynamic Structured Knowledge Flow
  • TamEdit: Trajectory-Aware Meta-Learning for Specificity-Preserving Continual Knowledge Editing
  • Query-Aware Knowledge Retrieval via Hyperbolic Structuring
  • ATIR: Towards Audio-Text Interleaved Contextual Retrieval
  • MAGIC: Deep Geometric Evolution with Structural Consensus for Temporal Knowledge Graph Reasoning
  • Efficient Prior-Guided Reasoning for Robust Retrieval-Augmented Generation under Conflicts
  • Bit-by-Bit: Progressive QAT Strategy with Outlier Channel Splitting for Stable Low-Bit LLMs
  • BrowseComp-Plus: A Fair and Disentangled Evaluation Benchmark for Deep Search Agents
  • LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked by Knowledge Points
  • Omni-Embed-Audio: Leveraging Multimodal LLMs for Robust Audio-Text Retrieval
  • MAB-DQA: Addressing Query Aspect Importance in Document Question Answering with Multi-Armed Bandits
  • How Training Data Shapes the Use of Parametric and In-Context Knowledge in Language Models
  • Learning What to Ignore: Mitigating Negative Transfer in Medical Knowledge Fusion via Clinical Task-Adaptive Selection
  • AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction
  • Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QA
  • ASTRA: Adaptive Semantic Tree Reasoning Architecture for Complex Table Question Answering
  • Mnemis: Dual-Route Retrieval on Hierarchical Graphs for Long-Term LLM Memory
  • When TableQA Meets Noise: A Dual Denoising Framework for Complex Questions and Large-scale Tables
  • Knowledge Beyond Language: Bridging the Gap in Multilingual Machine Unlearning Evaluation
  • AT²PO: Agentic Turn-based Policy Optimization via Tree Search
  • KoCo: Conditioning Language Model Pre-training on Knowledge Coordinates
  • Attention as Selector: Unlocking VLM Attention for Long Document Page Retrieval
  • TEMA: Anchor the Image, Follow the Text for Multi-Modification Composed Image Retrieval
  • WebAggregator: Enhancing Compositional Reasoning Capabilities of Deep Research Agent Foundation Models
  • Learning to Edit Knowledge via Instruction-based Chain-of-Thought Prompting
  • BadScientist: Can a Research Agent Write Convincing but Unsound Papers that Fool LLM Reviewers?
  • DFAMS: Dynamic-flow guided Federated Alignment based Multi-prototype Search
  • Locomo-Plus: Beyond-Factual Cognitive Memory Evaluation Framework for LLM Agents
  • WebSynthesis: World Model-Guided Monte Carlo Tree Search for Efficient WebAgent Trajectory Synthesis
  • Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based Documents
  • AIRCoder: Adaptive Integration of Multi-dimensional Retrieval for Repository-level Code Completion
  • Enhancing Agentic Textual Graph Retrieval with Synthetic Stepwise Supervision
  • Cognitive Scaffold: From Fluid Context to Crystallized Memory for Long-Horizon DeepResearch Agents
  • Learning to Think on Hypergraph: HyperCoT for Structure-Guided N-ary Knowledge Graph Completion
  • Re³: Relevance & Recency Retrieval for Mitigating Temporal Hallucination
  • TeCES: Collaborative Geometric Knowledge Representation Framework under Evolving Fact Snapshots
  • MicroC-KT: Modeling Community Effect via Learning Micro-Environment for Evidence-Grounded Explainable Knowledge Tracing
  • Temp-R1: A Unified Autonomous Agent for Complex Temporal KGQA via Reverse Curriculum Reinforcement Learning
  • S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA
  • Stable-RAG: Mitigating Retrieval-Permutation-Induced Hallucinations in Retrieval-Augmented Generation
  • Reliable Evaluation Protocol for Low-Precision Retrieval
  • VisRet: Visualization Improves Knowledge-Intensive Text-to-Image Retrieval
  • Selective Knowledge Distillation: Fusing LLM Semantic Strengths with DNN Efficiency for Binary Code Similarity Detection
  • SAFE-QAQ: End-to-End Slow-Thinking Audio-Text Fraud Detection via Reinforcement Learning
  • CIRAG: Construction–Integration Retrieval and Adaptive Generation for Multi-hop Question Answering
  • Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-Training
  • TrendFact: A Benchmark Towards Hotspot Perception in Automatic Fact-Checking
  • MPBoCo: Multimodal Prompt-based Boundary-enhanced Continual Framework for Joint Entity and Relation Extraction
  • SGVEF-LOOP: Coverage-Guided Progressive Topological Exploration and Fact-Grounded Metamorphic Evaluation for MCP Agents
  • Think Parallax: Solving Multi-Hop Problems via Multi-View Knowledge-Graph-Based Retrieval-Augmented Generation
  • Extending First-Order Logic for Factual Reasoning over Knowledge Graphs
  • ReCoQA: A Benchmark for Tool-Augmented and Multi-Step Reasoning in Real Estate Question and Answering
  • Attention Weights as an Indicator: Analyzing and Improving Document Utilization in Retrieval-Augmented Generation
  • DR-Arena: an Automated Evaluation Framework for Deep Research Agents
  • CaRL-EM: Cost-Aware Reinforcement Learning for Entity Matching with LLMs
  • ToMMeR - Efficient Entity Mention Detection from Large Language Models
  • Trust Within? Seek Beyond? Knowledge Boundary Aware Policy Optimization for Agentic Search
  • RPC-Bench: A Fine-grained Benchmark for Research Paper Comprehension
  • CompTab: A Comprehensive Benchmark for Real-World TableQA with Complex Reasoning and Irregular Tables
  • Learning More from Less: Exploiting Counterfactuals for Data-Efficient Chart Understanding
  • Video-MMMU: Evaluating Knowledge Acquisition from Multidisciplinary Professional Videos
  • Calibrated? Not for Everyone: How Sexual Orientation and Religious Markers Distort LLM Accuracy and Confidence in Medical QA
  • Collaboration of Fusion and Independence: Hypercomplex-driven Robust Multi-Modal Knowledge Graph Completion
  • MAGNET: Towards Adaptive GUI Agents with Memory-Driven Knowledge Evolution
  • Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval
  • KoCo-Bench: Can Large Language Models Leverage Domain Knowledge in Software Development?
  • FactVerse: A Benchmark for Factual Consistency in Interleaved Image–Text Generation
  • It’s High Time: A Survey of Temporal Question Answering
  • LLMs as Knowledge Graph Refiners: Mitigating Factual Inconsistencies in Generative Knowledge Extraction
  • Mitigating Structural Knowledge Collapse in Domain-Specific LLMs via Morpheme-Aware KV-Aggregation
  • Zero-shot Jianzi Recognition as Structured Visual Information Extraction in Open Compositional Symbolic Systems
  • Whose Facts Win? LLM Source Preferences under Knowledge Conflicts
  • Temporal Evidence Chain for Temporal Knowledge Graph Question Answering with Large Language Models
  • GLIER: Generative Legal Inference and Evidence Ranking for Legal Case Retrieval
  • HiChunk: Evaluating and Enhancing Retrieval Augmented Generation with Hierarchical Chunking
  • Can LLMs Act as Historians? Evaluating Historical Research Capabilities of LLMs via the Chinese Imperial Examination
  • Spectral Characterization and Mitigation of Sequential Knowledge Editing Collapse
  • Beyond Explicit Refusals: Soft-Failure Attacks on Retrieval-Augmented Generation
  • Compressing LLM Knowledge into Graph Representations for Text-attributed Graphs Learning
  • End-to-End Optimization of LLM-Driven Multi-Agent Search Systems via Heterogeneous-Group-Based Reinforcement Learning
  • CPR-RAG: Clinical Prior-Regularized Retrieval for Anatomy-Aware 3D CT Report Generation
  • EA-Agent: A Structured Multi-Step Reasoning Agent for Entity Alignment
  • QuantumQA: Enhancing Scientific Reasoning via Physics-Consistent Dataset and Verification-Aware Reinforcement Learning
  • The “Knowledge–Behavior Gap” in Cultural Taboo Safety of Large Language Models
  • MM-StanceDet: Retrieval-Augmented Multi-modal Multi-agent Stance Detection
  • Disentangling Reasoning Logic to Resolve Explicit Knowledge Conflicts
  • When Does Mixing Help? Analyzing Query Embedding Interpolation in Multilingual Dense Retrieval
  • Piece of Table: A Divide-and-Conquer Approach for Selecting Subtables in Table Question Answering
  • QA-MoE: Towards a Continuous Reliability Spectrum with Quality-Aware Mixture of Experts for Robust Multimodal Sentiment Analysis
  • Enhancing LLM-based Search Agents via Contribution Weighted Group Relative Policy Optimization
  • Optimizing User Profiles via Contextual Bandits for Retrieval-Augmented LLM Personalization
  • Teaching Language Models to Check Grounded Claim Factuality with Human Test-Taking Strategies
  • Stress Testing Factual Consistency Metrics for Long-Document Summarization
  • LLM-Generated Text May Harm Your Retrieval! A Robust Detection Strategy for Retrieval-Augmented Generation
  • The Retrieval Bottleneck: Scaling Laws for Reinforcement Learning in RAG
  • Generating then Refining for Reliable Knowledge Base Question Answering
  • StratMem-Bench: Evaluating Strategic Memory Use in Virtual Character Conversation Beyond Factual Recall
  • RSMeM: Knowledge-Enhanced Memory Evolution for Remote Sensing Agents with Systematic Evaluation
  • Behavior Knowledge Merge in Reinforced Agentic Models
  • MARS²: Scaling Multi-Agent Tree Search via Reinforcement Learning for Code Generation
  • AdabNER: Arabic Digital Archive Books with Nested Entity Recognition
  • Quantifying and Improving the Robustness of Retrieval-Augmented Language Models Against Spurious Features in Grounding Data
  • Beyond Static Artifacts: An Evolutionary Framework for Synthetic Claim Generation
  • VC-Inspector: Advancing Reference-free Evaluation of Video Captions with Factual Analysis
  • Defense Against Knowledge Poisoning Attack on GraphRAG
  • MM-PoisonRAG: Disrupting Multimodal RAG with Local and Global Knowledge Poisoning Attacks
  • IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering
  • MMSciCode: Real-world Evaluation of Multilingual Multi-Discipline Scientific Research Coding
  • ART: Attention Replacement Technique to Improve Factuality in LLMs
  • GameplayQA: A Benchmarking Framework for Decision-Dense POV-Synced Multi-Video Understanding of 3D Virtual Agents
  • DeepFact: Co-Evolving Benchmarks and Agents for Deep Research Factuality
  • ClaimDB: A Fact Verification Benchmark over Large Structured Data
  • New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
  • L2Dir: Integrating L₂-Norm and Directional Alignment for Unsupervised Contrastive Representation Learning in Multimodal Retrieval
  • TrustTable: A Neuro-Symbolic Auditing Framework for Faithful Table QA
  • Utility-Oriented Visual Evidence Selection for Multimodal Retrieval-Augmented Generation
  • KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering
  • Specializing Large Models for Oracle Bone Script Interpretation via Component-Grounded Multimodal Knowledge Augmentation
  • TIGER: Text-Informed Generalized Enzyme-Reaction Retrieval
  • Debate-of-Thoughts: Resolving Knowledge Conflicts in LLMs Through Internal Deliberation
  • GenomeQA: Benchmarking General Large Language Models for Genome Sequence Understanding
  • LLM-Powered Benchmark Factory: Reliable, Generic, and Efficient
  • COSMOS: Connectivity-Oriented Submodular Maximization for Optimal Subgraph Retrieval
  • Structured Semantic Information Helps Retrieve Better Examples for In-Context Learning Applied to Few-Shot Relation Extraction
  • Knowledge-driven Augmentation and Retrieval for Integrative Temporal Adaptation
  • PICTURE: Enhancing Theory-of-Mind in Large Language Models by Revealing, Not Hiding, Characters’ Lack of Knowledge
  • Flow-Based Page Unique Semantic Mapping Architecture for Document Visual Question Answering
  • LangSAE Editing: Improving Multilingual Information Retrieval via Post-hoc Language Identity Removal
  • Interpretable Traces, Unexpected Outcomes: Investigating the Disconnect in Trace-Based Knowledge Distillation
  • Towards Explainable Diagnosis: A Self-learned Explanatory Knowledge Base Approach
  • Joint Knowledge Base Completion and Question Answering by Combining Large Language Models and Small Language Models
  • Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
  • Verbal-R3: Verbal Reranker as the Missing Bridge between Retrieval and Reasoning
  • MTR-Suite: A Framework for Evaluating and Synthesizing Conversational Retrieval Benchmarks
  • Beyond Markovian Forgetfulness: Episodic Memory for Reasoning-Intensive Retrieval
  • Adaptive Retrieval for Reasoning
  • ToolOmni: Enabling Open-World Tool Use via Agentic learning with Proactive Retrieval and Grounded Execution
  • LegalGraphRAG: Multi-Agent Graph Retrieval-Augmented Generation for Reliable Legal Reasoning
  • TH-RAG : Topic-Based Hierarchical Knowledge Graphs for Robust Multi-hop Reasoning in Graph-based RAG Systems
  • Execution as Verification: Fine-Grained Self-Correcting Reasoning for Complex KBQA
  • Can Compact Language Models Search Like Agents? Distillation-Guided Policy Optimization for Preserving Agentic RAG Capabilities
  • Evolving Beyond Snapshots: Harmonizing Structure and Sequence via Entity State Tuning for Temporal Knowledge Graph Forecasting
  • CheckRLM: Effective Knowledge–Thought Coherence Checking in Retrieval-Augmented Reasoning
  • CAMEC: Complexity-Aware Multi-Expert Collaboration for Reliable Chinese Medical Question Answering
  • Frame-Semantic Knowledge Injection for Event-Level Inference in LLMs
  • QuDAR: Query-Wise Dual-Perspective Adaptive Retrieval
  • SPARKLE: A Structured and Plug-and-play Agentic Retrieval Policy for Adaptive RAG Models
  • SciCoQA: Quality Assurance for Scientific Paper–Code Alignment
  • ReTraceQA: Evaluating Reasoning Traces of Small Language Models in Commonsense Question Answering
  • MemCoRL: Alternating Co-Optimization of Memory Retrieval and Utilization via Collaborative Reinforcement Learning
  • Fast and Accurate Fisher-Guided Quantization via Efficient Kronecker Factorization
  • REAL: REtrieval-reAsoning and Logic-constructed Attention Behaviors for Long-Context KV Cache Compression
  • HowToNarrate: A General-Domain Benchmark for Synchronized Video Narration with External Knowledge
  • Knowledge Injection Exists in MoE? Exploring Expert-Aware Contrast Decoding in MoE for Mitigating LLMs’ Hallucinations
  • REaR : Retrieve, Expand and Refine for Effective Multitable Retrieval
  • Beyond Experience Retrieval: Learning to Generate Utility-Optimized Structured Experience for Frozen LLMs
  • KnowRL: Exploring Knowledgeable Reinforcement Learning for Factuality
  • ASTRA: An Automated Framework for Strategy Discovery, Retrieval, and Evolution for Jailbreaking LLMs
  • MedEinst: Benchmarking the Einstellung Effect in Medical LLMs through Counterfactual Differential Diagnosis
  • Decoding-Unlearning: Fact Forgetting via Entropy-Guided Inference
  • Measuring Social Bias in Vision-Language Models with Face-Only Counterfactuals from Real Photos
  • KASER: Knowledge-Aligned Student Error Simulator for Open-Ended Coding Tasks
  • Revisiting Evaluation of Question Answering Systems in Low-Resource Indic Languages: Bridging Human and Metric Alignment
  • Afri-MCQA: Multimodal Cultural Question Answering for African Languages
  • Structure Guided Retrieval-Augmented Generation for Factual Queries
  • MirrorQA: Benchmarking Multimodal LLMs on Mirror-Orientation Reasoning
  • M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question Answering
  • Knowledge-to-Verification: Exploring RLVR for LLMs in Knowledge-Intensive Domains
  • Mind Reader: Latent User Demand-Guided Content Optimization for Generative Search Engine
  • Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language Models
  • Deep-Reporter: Deep Research for Grounded Multimodal Long-Form Generation
  • Embodied-Reasoner: Synergizing Visual Search, Reasoning, and Action for Embodied Interactive Tasks
  • How Do Inpainting Artifacts Propagate to Language?
  • Failure Modes in Multi-Hop QA: The Weakest Link Effect and the Recognition Bottleneck
  • VeriTaS: The First Dynamic Benchmark for Multimodal Automated Fact-Checking
  • A Survey of Reasoning-Intensive Retrieval: Progress and Challenges
  • LOGICAL-COMMONSENSEQA: A Benchmark for Logical Commonsense Reasoning
  • Paper Circle: An Open-source Multi-agent Research Discovery and Analysis Framework
  • Can Large Language Models Keep Up? Benchmarking Online Adaptation to Continual Knowledge Streams
  • LogosKG: Hardware-Optimized Scalable and Interpretable Knowledge Graph Retrieval
  • GraphSynth: Resolving the Diversity-Reliability Trade-off with Probabilistic Factor Graphs
  • BioHiCL: Hierarchical Multi-Label Contrastive Learning for Biomedical Retrieval with MeSH Labels
  • Uncertainty-Aware Test-Time Search for Optimization Problem Solving
  • GUIDE: Towards Scalable Advising for Research Ideas
  • HeteroCache: A Dynamic Retrieval Approach to Heterogeneous KV Cache Compression for Long-Context LLM Inference
  • ViLL-E: Video LLM Embeddings for Retrieval
  • LEAF: Knowledge Distillation of Text Embedding Models with Teacher-Aligned Representations
  • Deep Research with Open-Domain Evaluation and Multi-Stage Guardrails for Safety
  • Understanding the Behaviors of Environment-aware Information Retrieval
  • Knowledge Vector of Logical Reasoning in Large Language Models
  • Investigating Counterfactual Unfairness in LLMs towards Identities through Humor
  • POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
  • INFACT: A Diagnostic Benchmark for Induced Faithfulness and Factuality Hallucinations in Video-LLMs
  • Temporal Leakage in Search-Engine Date-Filtered Web Retrieval: A Retrospective Forecasting Case Study
  • HSGraphAgent: Knowledge-Graph-Guided Large Language Models for Harmonized System Code Classification
  • Don’t Corrupt the Fact: A Trustworthy RAG Watermarking Framework based on Dual Factual Shield
  • How Long Reasoning Chains Influence LLMs’ Judgment of Answer Factuality
  • PolitNuggets: Benchmarking Agentic Discovery of Long-Tail Political Facts
  • Diagnosing and Remedying Representation Deficiencies for Deterministic Reasoning in KGQA
  • Explaining Sources of Uncertainty in Automated Fact-Checking
  • Evaluating Structure-Aware Retrieval and Safety in Statute-Centric Legal QA
  • SURE or Not? Investigating Semantic Understanding in Dense Retrieval Models
  • Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image Editing
  • FactCorrector: A Graph-Inspired Approach to Long-Form Factuality Correction of Large Language Models
  • What Makes AI Research Replicable? Executable Knowledge Graphs as Scientific Knowledge Representations
  • Factual Retrieval in LLMs Is a Redundant, Distributed and Non-Contiguous Process
  • Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
  • DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation
  • Revisiting Non-Verbatim Memorization in Large Language Models: The Role of Entity Surface Forms
  • ReasMark: A Robust Watermark for Attributing LLM Reasoning Under Knowledge Distillation Attacks
  • Awakening Dormant Experts:Counterfactual Routing to Mitigate MoE Hallucinations
  • KARL: Reinforcement Learning for LLM Agents on Multi-Turn Knowledge-Intensive Agentic Tasks
  • FIGMA: Towards FIne-Grained Music retrievAl
  • Diving into the Decoding Space of Non-Autoregressive Models via Lexically Constrained Search
  • Beyond Variance: Knowledge-Aware LLM Compression via Fisher-Aligned Subspace Diagnostics
  • MegaRAG: Multimodal Knowledge Graph-Based Retrieval Augmented Generation