Paper Digest: Recent Papers on Question Answering
Paper Digest Team extracted all recent Question Answering related papers on our radar, and generated highlight sentences for them. The results are then sorted by relevance & date. In addition to this ‘static’ page, we also provide a real-time version of this article, which has more coverage and is updated in real time to include the most recent updates on this topic.
Since 2018, Paper Digest has built a foundation of data spanning decades of conferences, journals, and research topics. The platform features a daily digest service that sifts through tens of thousands of new papers, clinical trials, news articles, and community posts, filtering the noise to highlight what matters most to specific interests. Beyond daily updates, dozens of built-in research tools streamline the academic workflow, supporting efficient reading and writing, comprehensive literature reviews, and automated research report generation.
Paper Digest Team
New York City, New York, 10017
team@paperdigest.org
TABLE 1: Paper Digest: Recent Papers on Question Answering
| Paper | Author(s) | Source | Date | |
|---|---|---|---|---|
| 1 | Bridging Legal Language Barriers Using Explainable AI: Outcome Prediction and Multilingual Knowledge Based Answer Retrieval for Indian Law Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study presents an integrated legal AI platform that combines interpretable case outcome prediction with multilingual, retrieval-grounded legal question answering to improve access to Indian law. |
Manish Thirunavu D; | International Journal of Latest Technology in Engineering … | 2026-07-16 |
| 2 | Expanding The Lexicon of Ge’ez Based African Languages: A Comparative Study of Amharic and Tigrinya Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our contributions are: (i) a vocabulary-extension and embedding-initialization procedure tailored to Ge’ez script; (ii) a two-stage training strategy under which vocabulary and continued-pretraining gains on Amharic/Tigrinya transfer to 17 typologically related, unaugmented African languages; and (iii) an evaluation spanning both intrinsic tokenization metrics (vocabulary coverage, fertility, OOV rate) and extrinsic task performance across all 19 languages. |
Hailay Kidu Teklehaymanot; Debela Desalegn Yadeta; Wolfgang Nejdl; | arxiv-cs.CL | 2026-07-16 |
| 3 | Stop Thinking, Start Looking: Efficient Post-Training for Multimodal Document Question Answering Via Reasoning-Free Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Perception-RFT, a training framework that applies Group Relative Policy Optimization (GRPO) to multimodal document QA, bypassing intermediate reasoning tokens to directly align visual features with structured grounding outputs. |
HARIKRISHNAN P M et. al. | arxiv-cs.AI | 2026-07-16 |
| 4 | EgoProceVQA: A Novel Egocentric Procedural Understanding Task with Self-Skill-Exploration Agent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such capabilities are crucial for building procedural AI assistants deployable on wearable devices. To bridge this gap, we introduce the Egocentric Procedural Understanding VQA task (EgoProceVQA), which systematically evaluates egocentric procedural reasoning abilities of current MLLMs and agents through six types of key-step-centric questions. |
JUNLONG LI et. al. | arxiv-cs.CV | 2026-07-15 |
| 5 | Automatic Question Answering Systems: A Comprehensive Review Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The Question Answering system combines research from several fields, including Natural Language Processing, Artificial Intelligence, Information Retrieval, and Information Extraction. |
Sanah Nashir Sayyed; Bharat Shelke; C. Namrata Mahender; | International Journal for Research in Applied Science and … | 2026-07-14 |
| 6 | Technical Report on The CVPR 2026@AdvML Workshop Challenge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This report presents the CVPR 2026@AdvML Workshop Challenge on adversarial multimodal attacks against autonomous-driving VLAs. |
TIANYUAN ZHANG et. al. | arxiv-cs.CV | 2026-07-13 |
| 7 | Confidently Wrong: Detecting Hallucinations in Financial Question Answering from LLM Internal States Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We ask how reliably such confidently wrong answers, or confident hallucinations, can be detected from a model’s internal activations, and whether those activations carry information beyond its observable outputs. We train linear probes on the residual stream and evaluate them on two established question-answering (QA) benchmarks built from real filings, FinQA and TAT-QA. |
Richard Zhe Wang; | arxiv-cs.CL | 2026-07-13 |
| 8 | LakeQuest: A Three-Domain Benchmark for Grounded Question Answering Across Data Lakes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current benchmarks abstract away this noisy discovery process, failing to evaluate end-to-end performance. To bridge this gap, we introduce LakeQuest, a human-validated benchmark of 9,846 QA pairs designed to evaluate the end-to-end retrieve-and-synthesize pipeline over realistic data lakes. |
MICHAEL SOLODKO et. al. | arxiv-cs.CL | 2026-07-13 |
| 9 | DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, such errors are often recalcitrant to tracing and calibration, posing a critical bottleneck for their practical deployment in geospatial tasks. To address this pressing challenge, this study proposes DM-KG (Direction-Metric Knowledge Graph), a structurally grounded spatial representation framework for street view imagery. |
XINYUE XU et. al. | arxiv-cs.CV | 2026-07-13 |
| 10 | Evidence-Backed Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Evidence-Backed Video Question Answering (E-VQA), a novel task requiring models to jointly output a semantic answer and precise spatio-temporal evidence: temporal segments and dense, tracked object segmentation masklets. |
SHIJIE WANG et. al. | arxiv-cs.CV | 2026-07-13 |
| 11 | STEC: Evidence Compression for Deep Search in Open-domain Multi-Hop QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Directly comparing raw trajectories exposes the verifier to noisy and unaligned content, while comparing answer strings ignores the evidence supporting each candidate, making reliable final selection difficult. To address this challenge, we propose STEC, an evidence compression framework for final answer selection in multi-hop QA. |
XINKANG LI et. al. | arxiv-cs.AI | 2026-07-12 |
| 12 | CRiT-QA: Evaluating Multi-hop Reasoning with Counterfactual Chains and Distractor Traps Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CRiT-QA (Counterfactual Reasoning with Traps), a dataset explicitly designed to address both limitations. |
JungMin Yun; JuneHyoung Kwon; YoungBin Kim; | arxiv-cs.AI | 2026-07-12 |
| 13 | Question Answering for Diagram-Rich Technical Meeting Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper reports our industrial experience developing and evaluating LMVQA, an LLM-based multimodal question-answering system for technical meeting videos. |
ZHUORAN XU et. al. | arxiv-cs.SE | 2026-07-11 |
| 14 | CLIR-Bench: Benchmarking Multimodal Question Answering Over Irregular Clinical Time Series Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks primarily focus on regularly sampled time-series QA or medical QA over static data, and therefore rarely assess whether models can faithfully ground their answers in irregular temporal observations. To fill this gap, we introduce CLIR-Bench, a benchmark for irregular clinical time series QA constructed from de-identified ICU records through a principled four-stage pipeline. |
FRANK NIE et. al. | arxiv-cs.CL | 2026-07-10 |
| 15 | Task-Specific Multimodal Question Answering Agents Via Confidence Calibration and Incremental Reasoning for QANTA 2026 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present our submission to the QANTA 2026 shared challenge at the ICML 2026 Workshop on Efficient Multimodal Question Answering (EMM-QA). |
Nirjhar Das; Md. Al-Mamun Provath; | arxiv-cs.CL | 2026-07-10 |
| 16 | Interpretable Uncertainty for Adaptive Retrieval and Reasoning in Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an uncertainty-aware framework for adaptive QA based on explicit signals derived from LLM internal representations. |
Ritajit Dey; Iadh Ounis; Graham McDonald; | arxiv-cs.IR | 2026-07-08 |
| 17 | Healthier LLMs: Retrieval-Augmented Generation for Public Health Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To assess realistic free-form answering, we introduce a rubric-based LLM-as-a-judge covering faithfulness, completeness, clarity, and factual consistency, and validate it against dual human annotations. |
FELIX FELDMAN et. al. | arxiv-cs.CL | 2026-07-07 |
| 18 | Mitigating Factual Hallucination in Large Reasoning Models Via Mixed-Mode Advantage Regularization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To explain this phenomenon, we formulate explicit thinking in factuality QA as a thinking residual over the model’s direct-answer tendency, which can either recover missing knowledge or introduce unsupported associations. Based on this formulation, we propose MARGO, \underline{\textit{M}}ixed-Mode \underline{\textit{A}}dvantage \underline{\textit{R}}egularization for \underline{\textit{G}}rounded \underline{\textit{O}}ptimization, a reinforcement learning framework that uses non-thinking rollouts as same-model references in advantage estimation. |
KAISHEN WANG et. al. | arxiv-cs.CL | 2026-07-07 |
| 19 | From Voting to Agent Collaboration: Answer-Type-Aware LLM Pipelines for BioASQ 14b Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study presents a question-type-specific large language model (LLM) framework for BioASQ 14b Task B, designed to improve answer robustness and evidence grounding in biomedical question answering. |
TAEYUN ROH et. al. | arxiv-cs.CL | 2026-07-07 |
| 20 | Retrieving A Set, Not Independent Passages: Set-Level Compatibility Learning for Efficient Set Exploration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: LLM-based set selection can model such interactions, but its computational cost limits practical use. We address this gap by formulating multi-hop retrieval as query-set compatibility scoring and propose a set-level retrieval framework. |
Mooho Song; Jay-Yoon Lee; | arxiv-cs.IR | 2026-07-06 |
| 21 | Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CIC, a confidence-interval-based calibration framework that converts arbitrary uncertainty scores into risk-controlled selective answering rules. |
Sijin Dong; Hiroyuki Shinnou; | arxiv-cs.CL | 2026-07-05 |
| 22 | Seeing Once Is Enough? Online Geometry-Aware Token Pruning for 3D Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To mitigate this issue, existing solutions rely on strategic frame selection or token-merging algorithms that require preprocessing in advance all frames of the scene, i.e., an offline fashion. In contrast, we propose the first online token-pruning method that can be integrated seamlessly with current MLLM models for 3D question answering tasks, without additional training and with lower memory usage.Our key insight is to project each input frame into a shared voxel space using depth information and camera pose, identifying spatially-overlapped regions across frames and selectively pruning redundant image tokens before they enter the language model. |
Ruei-Chi Lai; Bolivar Solarte; Chin-Hsuan Wu; Yi-Hsuan Tsai; Min Sun; | arxiv-cs.CV | 2026-07-04 |
| 23 | HETERQA: Benchmarking Record Retrieval Over Multiple Heterogeneous Sources Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing benchmarks are constructed from individual sources, and only a very few recent benchmarks have considered two or three sources. To alleviate this issue, we introduce HETERQA, a comprehensive benchmark with 857 QA pairs for record retrieval over five heterogeneous sources. |
Yaodong Su; Hanchang Li; Quanqing Xu; Chuanhui Yang; Yixiang Fang; | arxiv-cs.IR | 2026-07-03 |
| 24 | IDEAL-Bench: Indoor Dataset and Evaluation Suite for Analyzing 3D Layout Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce IDEAL-Bench, an evaluation suite that requires VLMs to predict structured 3D layouts on photorealistic indoor scenes across 10 room types, scored along five numerical dimensions and a perceptual render-and-compare protocol. |
Yuening Cai; Junwei Zhou; Youran Qu; Yu-Wing Tai; | arxiv-cs.CV | 2026-07-03 |
| 25 | Transformers and BiLSTM-Based Ensemble Modeling for Entity and Relation Aware Biomedical Extractive Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study introduces a novel transformer-based approach for biomedical EQA that incorporates Named Entity Recognition (NER) to identify key medical terms, Relation Extraction (RE) to understand their interconnections, and a BiLSTM layer to enhance contextual comprehension. |
Ahmed Ajmine Nehal; Md Mehedi Hasan; Farhana Elias; Rashedur M. Rahman; | Vietnam Journal of Computer Science | 2026-07-03 |
| 26 | LuxSQA: Ask Me in Luxembourgish with TTS-Augmented Spoken Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper investigates whether text-to-speech (TTS) can provide task-specific training data for Luxembourgish SQA without requiring a large human-recorded QA corpus. |
Nina Hosseini-Kivanani; Marco Matassoni; Alessio Brutti; | arxiv-cs.CL | 2026-07-02 |
| 27 | ReQuest: Rethinking-based Question-Aware Frame Selection for Long-Form Video QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ReQuest , an uncertainty-driven, question-adaptive keyframe selection pipeline that aligns question intent with relevant video content through selective computation. |
MINKUK KIM et. al. | arxiv-cs.CV | 2026-07-02 |
| 28 | MMTR: Strategy-Guided Multimodal Table Reasoning with Reflective Self-Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limitation primarily stems from the high density of structured information inherent in tables and the scarcity of high-quality instruction tuning data. To address these challenges and improve the model’s reasoning accuracy in tables, we propose MMTR, a strategy-guided multimodal table reasoning method with reflective self-correction. |
Lixin Bai; Yibo Ming; Yanmin Chen; | Information | 2026-07-01 |
| 29 | Dual-Confidence Contrastive Decoding for Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, we study the complementary problem of intra-context conflict in multi-document RAG. To evaluate this setting, we introduce DRQA, a factual-conflict question answering benchmark derived from enterprise deep-research scenarios, where answers are grounded in synthetic enterprise-specific facts that are designed not to be recoverable from the model’s internal memory. |
RAYMOND LI et. al. | arxiv-cs.CL | 2026-07-01 |
| 30 | Imprint: Online Memory Compression for Long-Horizon Egocentric QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Imprint, an interaction-centric memory framework that formulates long-horizon egocentric memory as an online memory compression problem rather than summarization. |
Kousik Das; Debaditya Roy; | arxiv-cs.CV | 2026-07-01 |
| 31 | Self-Study Reconsidered: The Hidden Fragility of Learning from Self-Generated QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Language models are increasingly taught from synthetic question–answer (QA) supervision: a model generates questions about a document, answers them from the same text, and the resulting pairs are used to fine-tune, distill, or compress knowledge into another model. |
EKATERINA ALIMASKINA et. al. | arxiv-cs.AI | 2026-06-30 |
| 32 | Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce an OpenMIC-derived diagnostic benchmark sequence for instrument grounding in music audio-language models, extending binary instrument-presence QA to genre-prior-reduced examples, confusable instrument discrimination, longer audio context, and temporal localization. |
Yujun Lee; Joonhyeok Shin; Hyoeun Kim; Kyuhong Shim; | arxiv-cs.SD | 2026-06-30 |
| 33 | Clinically Structured Rank-Gated LoRA for Cross-Benchmark Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose BiRG-LoRA, a single-adapter rank-gated LoRA method for medical question answering. |
Yining Huang; | arxiv-cs.CL | 2026-06-30 |
| 34 | A Reciprocal Interaction Framework for Collaborative Temporal Grounding and Question Answering in Egocentric Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, VQA models often generate ambiguous answers due to the lack of precise temporal cues, while VTG models fail to fully exploit the high-level semantic information embedded in the answers. To address these limitations, We propose a Reciprocal Interaction Framework (RIF). |
Jiaxu Wang; Tianshan Liu; Bing-Kun Bao; | ACM Transactions on Multimedia Computing, Communications, … | 2026-06-30 |
| 35 | JL1-CC&QA: Extending The JL1-CD Benchmark with Change Captioning and Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Remote sensing change detection (CD) traditionally focuses on pixel-level binary segmentation, which identifies where changes occur but neither what nor why. To bridge this semantic gap, we introduce JL1-CC&QA, a multi-task benchmark that extends the JL1-CD dataset with two complementary annotation layers: change captioning (CC) and change question answering (QA). |
Ziyuan Liu; Ruifei Zhu; Ouqiao Ma; Yuantao Gu; | arxiv-cs.CV | 2026-06-30 |
| 36 | Efficient Retrieval-Augmented Generation Via Token Co-occurrence Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent graph-based RAG methods improve the retrieval of interconnected chunks, they often rely on computationally expensive and error-prone LLM-based extraction pipelines. To address these issues, we propose TIGRAG (Token-Induced GraphRAG), an efficient graph-augmented RAG framework based on a token co-occurrence Knowledge Graph. |
GIANLUCA BONIFAZI et. al. | arxiv-cs.CL | 2026-06-29 |
| 37 | How Far Can You Get Without A GPU? A Systematic Benchmark of Lightweight Hallucination Detection Across Question Answering, Dialogue, and Summarisation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we explore a practical alternative: how well can hallucination detection perform using only lightweight, CPU-feasible methods built on publicly available models? |
Kriti Faujdar; Smit Kadvani; | arxiv-cs.CL | 2026-06-29 |
| 38 | ARMOR: Adaptive Retriever Optimization for Low-Resource Telecom Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we introduce ARMOR, Adaptive Regularized Mixture Optimization for Retrievers, which learns separate temperatures for the RAG retrieval distribution and InfoNCE softmax and regularizes the adapted query encoder toward the frozen base query encoder. |
Heshan Fernando; Quan Xiao; Yan Xin; Tianyi Chen; | arxiv-cs.IR | 2026-06-28 |
| 39 | Know The Known and The Unknown: Reasonable Answer Generation with Knowledge-Informed Citations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they often overlook key challenges such as citation granularity, the awareness of unknown information, and the adoption of effective training strategies. In this paper, we introduce Knowledge-informed Citation (KFC), which addresses these issues through a novel data construction pipeline, a new benchmark, and an innovative training strategy. |
YICHI ZHANG et. al. | acl | 2026-06-27 |
| 40 | Beyond Scaling: Measuring and Predicting The Upper Bound of Knowledge Retention in Language Model Pre-Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Size-dependent Mutual Information (SMI), an information-theoretic predictor that integrates knowledge frequency, knowledge specificity, and model size to forecast closed-book question answering (QA) accuracy. |
CHANGHAO JIANG et. al. | acl | 2026-06-27 |
| 41 | HiKEY: Hierarchical Multimodal Retrieval for Open-Domain Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches relying on flat text chunks or page-level images inherently struggle to (i) precisely pinpoint the target document among thousands of candidates and (ii) organically connect multimodal evidence, such as tables and figures, within a limited token budget. To address these challenges, we propose HiKEY, a hierarchical tree-based multimodal retrieval framework that elevates document hierarchy to a first-class retrieval signal. |
Joongmin Shin; Gyuho Shim; Jeongbae Park; Jaehyung Seo; Heuiseok Lim; | acl | 2026-06-27 |
| 42 | Towards Faithful Industrial RAG: A Reinforced Co-adaptation Framework for Advertising QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a reinforced co-adaptation framework that jointly optimizes retrieval and generation through two components: (1) Graph-aware Retrieval (GraphRAG), which models entity-relation structure over a high-citation knowledge subgraph for multi-hop, domain-specific evidence selection; and (2) evidence-constrained reinforcement learning via Group Relative Policy Optimization (GRPO) with multi-dimensional rewards covering faithfulness, style compliance, safety, and URL validity. |
WENWEI LI et. al. | acl | 2026-06-27 |
| 43 | EVE: A Domain-Specific LLM Framework for Earth Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Earth Virtual Expert (EVE), the first open-source, end-to-end initiative for developing and deploying domain-specialized LLMs for Earth Intelligence. |
ÀLEX R. ATRIO et. al. | acl | 2026-06-27 |
| 44 | AMATA: Adaptive Multi-Agent Trajectory Alignment for Knowledge-Intensive Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These difficulties are primarily due to hallucination and the limitations of LLMs in bridging long-tail knowledge gaps. To address this, we propose AMATA, an Adaptive Multi-Agent Trajectory Alignment framework that dynamically integrates external knowledge to improve response interpretability and factual grounding. |
TAOLIN ZHANG et. al. | acl | 2026-06-27 |
| 45 | Beyond Chunking: Discourse-Aware Hierarchical Retrieval for Long Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a discourse-aware hierarchical framework that leverages rhetorical structure theory (RST) for long document question answering. |
HUIYAO CHEN et. al. | acl | 2026-06-27 |
| 46 | Protecting Multimodal Large Language Models Against Misleading Visualizations Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We find that two methods, table-based QA and redrawing the visualization, are effective, with improvements of up to 19. |
Jonathan Tonglet; Tinne Tuytelaars; Marie-Francine Moens; Iryna Gurevych; | acl | 2026-06-27 |
| 47 | EASE: Entity-Aware Sub-table Generation for Real-world Multi-table QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenges of real-world table QA, we propose **EASE**: **E**ntity-**A**ware **S**ub-table Generation for R**E**al-world Multi-table QA framework. |
MYUNGHOON KANG et. al. | acl | 2026-06-27 |
| 48 | Chart-MRAG: Benchmarking Multimodal Retrieval Augmented Generation on Chart-based Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks primarily focus on simple image-text interactions, overlooking complex visual formats like charts that are prevalent in real-world applications. In this work, we introduce a novel task, Chart-based MRAG, to address this limitation. |
JIANG ZHONG et. al. | acl | 2026-06-27 |
| 49 | Biomedical Question Answering Via Multi-Level Summarization on A Local Knowledge Graph Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a novel method CLAIMS, which utilizes propositional claims to construct a local knowledge graph from retrieved documents. |
Lingxiao Guan; Yuanhao Huang; Jie Liu; | acl | 2026-06-27 |
| 50 | SAGE: A Search-AuGmented Evaluation of Large Language Models on Free-Form QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Search-AuGmented Evaluation (SAGE), a framework to assess LLM outputs without fixed ground-truth answers. |
Sher Badshah; Ali Emami; Hassan Sajjad; | acl | 2026-06-27 |
| 51 | Question Difficulty Estimation for Large Language Models Via Answer Plausibility Scoring Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Q-DAPS (Question Difficulty based on Answer Plausibility Scores) method, a novel approach that estimates question difficulty by computing the entropy of plausibility scores over candidate answers. |
Jamshid Mozafari; Bhawna Piryani; Adam Jatowt; | acl | 2026-06-27 |
| 52 | KG-MuLQA: A Framework for KG-based Multi-Level QA Extraction and Long-Context LLM Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce KG-MuLQA (Knowledge-Graph-based Multi-Level Question-Answer Extraction): a framework that (1) extracts QA pairs at multiple complexity levels (2) along three key dimensions – multi-hop retrieval, set operations, and answer plurality, (3) by leveraging knowledge-graph-based document representations. |
NIKITA TATARINOV et. al. | acl | 2026-06-27 |
| 53 | LinkQA: Synthesizing Diverse QA from Multiple Seeds Strongly Linked By Knowledge Points Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The advancement of large language models (LLMs) struggles with the scarcity of high-quality, diverse training data. To address this limitation, we propose LinkSyn, a KP-graph-based synthesis framework that for the first time enables flexible control over discipline and difficulty distributions while balancing KP coverage and popularity. |
XUEMIAO ZHANG et. al. | acl | 2026-06-27 |
| 54 | CAMEC: Complexity-Aware Multi-Expert Collaboration for Reliable Chinese Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CAMEC (Complexity-Aware Multi-Expert Collaboration), a framework that combines hierarchical medical adaptation with complexity-aware expert routing for reliable Chinese medical QA. |
YUKANG WU et. al. | acl | 2026-06-27 |
| 55 | AraVQA: Building A New Arabic Factoid Visual Question Answering Dataset from Wikipedia Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In addition, most of the existing Arabic VQA datasets focus on culturally-specific and dialect-aware domains. To address these limitations, we propose a new pipeline that leverages Wikipedia template tags to extract the relevant information for each image, which is subsequently utilized by the Large Language Model (LLM) to synthetically generate a new visual question answering dataset. |
Sultan Alrowili; Younes Samih; Abed Alhakim Freihat; Mathan Kumar Eswaran; | acl | 2026-06-27 |
| 56 | IterCOMP: Reasoning-aware Adaptive Prompt Compression for Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose IterCOMP, a unified, training-free prompt compression framework that incorporates multi-hop reasoning within an iterative compression loop. |
JungMin Yun; YoungBin Kim; | acl | 2026-06-27 |
| 57 | SAHM: A Benchmark for Arabic Financial and Shari’ah-Compliant Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SAHM, a document-grounded benchmark and instruction-tuning dataset for Arabic financial NLP and Shari’ah-compliant reasoning. |
RANIA ELBADRY et. al. | acl | 2026-06-27 |
| 58 | Beyond Timestamps: Bridging Forward and Backward Reasoning in Temporal Numerical and Relational Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a bi-directional evaluation framework consisting of forward generation via Question Answering (QA) and backward verification via Fact Verification (FV). |
XINYING QIAN et. al. | acl | 2026-06-27 |
| 59 | LoVeC: Reinforcement Learning for Better Verbalized Confidence in Long-Form Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose LoVeC (Long-form Verbalized Confidence), a novel reinforcement learning (RL)–based method that trains LLMs to append an on-the-fly numerical confidence score to each generated statement during long-form generation. |
Caiqi Zhang; Xiaochen Zhu; Chengzu Li; Nigel Collier; Andreas Vlachos; | acl | 2026-06-27 |
| 60 | AutoGraph-R1: End-to-End Reinforcement Learning for Knowledge Graph Construction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, its effectiveness is hindered by a fundamental disconnect: the knowledge graph (KG) construction process is decoupled from its downstream application, yielding suboptimal graph structures. To bridge this gap, we introduce AutoGraph-R1, the first framework to directly optimize KG construction for task performance using Reinforcement Learning (RL). |
HONG TING TSANG et. al. | acl | 2026-06-27 |
| 61 | POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Charts are a universally adopted medium for data communication, yet existing chart understanding benchmarks are overwhelmingly English-centric, limiting their accessibility and relevance to global audiences. To address this limitation, we introduce PolyChartQA, the first large-scale multilingual benchmark for chart question answering, comprising 22,606 charts and 26,151 QA pairs across 10 diverse languages. |
YICHEN XU et. al. | acl | 2026-06-27 |
| 62 | LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This often results in incomplete evidence retrieval and degraded answer quality for multi-page reasoning tasks. To address these limitations, we propose LAD-RAG, a novel Layout-Aware Dynamic RAG framework. |
ZHIVAR SOURATI et. al. | acl | 2026-06-27 |
| 63 | CT-FineBench: A Diagnostic Fidelity Benchmark for Fine-Grained Evaluation of CT Report Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Conventional evaluation metrics offer only coarse measures of lexical overlap or entity matching and fail to reflect the granular diagnostic accuracy required for clinical use. To address this gap, we propose CT-FineBench, a benchmark built from CT-RATE and Merlin to evaluate the fine-grained factual consistency of CT reports, constructed from CT-RATE and Merlin. |
RUIFENG YUAN et. al. | acl | 2026-06-27 |
| 64 | AgentRouter: A Knowledge-Graph-Guided LLM Router for Collaborative Multi-Agent Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose AgentRouter, a framework that formulates multi-agent QA as a knowledge-graph–guided routing problem supervised by empirical performance signals. |
ZHEYUAN ZHANG et. al. | acl | 2026-06-27 |
| 65 | THOR: A Theta-Gamma Hierarchical Oscillatory Reasoning Framework for Multi-hop QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by Theta-Gamma hierarchical oscillation which decouples global planning from local retrieval, enabling efficient attention transfer between hops and a verification and repair mechanism that interrupts the accumulation of errors in the wrong paths, we present **THOR**, a brain-inspired Theta-Gamma hierarchical oscillatory reasoning framework. |
Ziyang Ling; Ronald X. Xu; Mingzhai Sun; | acl | 2026-06-27 |
| 66 | MINTQA: A Multi-Hop Question Answering Benchmark for Evaluating LLMs on New and Long-tail Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce MINTQA (Multi-hop Question Answering on New and Tail Knowledge), a benchmark designed to evaluate multi-hop QA performance on questions involving 10,479 question-answer pairs for evaluating old/new knowledge and 17,887 pairs for assessing popular/unpopular knowledge, with each question equipped with corresponding sub-questions and answers. |
Jie He; Nan Hu; Wanqiu Long; Jiaoyan Chen; Jeff Z. Pan; | acl | 2026-06-27 |
| 67 | ODUTQA-MDC: A Task for Open-Domain Underspecified Tabular QA with Multi-turn Dialogue-based Clarification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The advancement of large language models (LLMs) has enhanced tabular question answering (Tabular QA), yet they struggle with open-domain queries exhibiting underspecified or uncertain expressions. To address this, we introduce the ODUTQA-MDC task and the first comprehensive benchmark to tackle it. |
ZHENSHENG WANG et. al. | acl | 2026-06-27 |
| 68 | Piece of Table: A Divide-and-Conquer Approach for Selecting Subtables in Table Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose *PieTa* (Piece of Table), a divide-and-conquer subtable selection framework that progressively aggregates locally selected evidence without requiring explicit global reasoning. |
Wonjin Lee; Kyumin Kim; Sungjae Lee; Jihun Lee; Kwang In Kim; | acl | 2026-06-27 |
| 69 | Revisiting Evaluation of Question Answering Systems in Low-Resource Indic Languages: Bridging Human and Metric Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These metrics often exhibit issues such as compressed scoring ranges, excessive zero scores, and weak alignment with human judgments. To overcome these limitations, this work introduces the LRM2QAS (Language Robust Multi-aspect Metrics for Question Answering Systems). |
Anuj Kumar; Satyadev Ahlawat; Yamuna Prasad; Virendra Singh; | acl | 2026-06-27 |
| 70 | Interpretable Traces, Unexpected Outcomes: Investigating The Disconnect in Trace-Based Knowledge Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To isolate the effect of trace semantics, we design experiments in the Question Answering (QA) domain using a rule-based problem decomposition method. This enables us to create Supervised Fine-Tuning (SFT) datasets for LLMs where – each QA problem is paired with either verifiably correct or incorrect CoT traces, while always providing the correct final solution. |
Siddhant Bhambri; Upasana Biswas; Subbarao Kambhampati; | acl | 2026-06-27 |
| 71 | Multimodal Graph RAG for Long-range Visually Rich Document Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing LLM-based KG construction methods handle only the language modality, leaving the automatic creation of multimodal KGs (MMKGs) for visually rich documents largely unexplored. In this paper, we introduce a multimodal graph-based RAG approach to tackle this problem. |
Yi-Cheng Wang; Chu-Song Chen; | arxiv-cs.IR | 2026-06-27 |
| 72 | MUSEG: Reinforcing Video Temporal Understanding Via Timestamp-Aware Multi-Segment Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose **MUSEG**, a novel RL-based method that enhances temporal understanding by introducing timestamp-aware multi-segment grounding. |
FUWEN LUO et. al. | acl | 2026-06-27 |
| 73 | Knowledge Is Not Enough: Injecting RL Skills for Continual Adaptation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We empirically observe that the parameter updates induced by SFT and RL are nearly orthogonal. Based on this observation, we propose **Parametric Skill Transfer (PaST)**, a framework that supports modular skill transfer for efficient and effective knowledge adaptation. |
Pingzhi Tang; Yiding Wang; Muhan Zhang; | acl | 2026-06-27 |
| 74 | Answering The Wrong Question: Reasoning Trace Inversion for Abstention in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Taking the vulnerabilities of reasoning models into account, we propose our Query Misalignment Framework. |
Abinitha Gourabathina; Inkit Padhi; Manish Nagireddy; Subhajit Chaudhury; Prasanna Sattigeri; | acl | 2026-06-27 |
| 75 | PubMed Reasoner: Dynamic Reasoning-based Retrieval for Evidence-Grounded Biomedical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Retrieval-augmented approaches partially address this gap but lack mechanisms to iteratively refine poor queries, whereas self-reflection methods kick in only after full retrieval is completed. In this context, we introduce PubMed Reasoner, a biomedical QA agent composed of three stages: **self-critic query refinement** evaluates MeSH terms for coverage, alignment, and redundancy to enhance PubMed queries based on partial (metadata) retrieval; **reflective retrieval** processes articles in batches until sufficient evidence is gathered; and **evidence-grounded response generation** produces answers with explicit citations. |
Yiqing Zhang; Xiaozhong Liu; Fabricio Murai; | acl | 2026-06-27 |
| 76 | It’s High Time: A Survey of Temporal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this survey, we provide a comprehensive overview of Temporal Question Answering (TQA), a research area that focuses on answering questions involving temporal constraints or context. |
Bhawna Piryani; Abdelrahman Abdallah; Jamshid Mozafari; Avishek Anand; Adam Jatowt; | acl | 2026-06-27 |
| 77 | Tailoring Diagnostic Modeling to Individual Learners: Personalized Distractor Generation Via MCTS-Guided Reasoning Reconstruction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While promising, this task is particularly demanding due to the limited number of QA records available for each student, which are insufficient for training, as well as the absence of their underlying reasoning process. To overcome this, we propose a novel, training-free two-stage framework. |
TAO WU et. al. | acl | 2026-06-27 |
| 78 | S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose S2G-RAG (Structured Sufficiency and Gap-judging RAG), an iterative framework with an explicit controller, S2G-Judge. |
Minghan Li; Junjie Zou; Xinxuan Lv; Chao Zhang; Guodong Zhou; | acl | 2026-06-27 |
| 79 | What Question Did You Answer? Refining Contact Center Evaluation Plans Via Backward Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Backward Question-based Refinement (BQR), a diagnostic framework that generates backward questions, revealing what a model understood rather than what was asked, to systematically distill implicit reasoning from large LMs into explicit evaluation plans. |
Prajwal Sood; Rushikesh Pawar; Digvijay Anil Ingle; Anup Pattnaik; | acl | 2026-06-27 |
| 80 | Latent Bridges for Multi-Table Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce GRAB, a constructor-encoder-bridge pipeline for table question answering. |
Simone Varriale; Tamara Cucumides; Floris Geerts; Paolo Papotti; | arxiv-cs.CL | 2026-06-27 |
| 81 | CRAFT: Training-Free Cascaded Retrieval for Tabular QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work establishes a scalable and adaptable paradigm for table retrieval, bridging the gap between fine-tuned architectures and lightweight, plug-and-play retrieval systems. |
Adarsh Singh; Kushal Raj Bhandari; Jianxi Gao; Soham Dan; Vivek Gupta; | acl | 2026-06-27 |
| 82 | ProMSA:Progressive Multimodal Search Agents for Knowledge-Based Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ProMSA, a progressive multimodal search agent for KB-VQA. |
ZHENGXIAN WU et. al. | arxiv-cs.CV | 2026-06-26 |
| 83 | QG-STR: Training-Time Optimized Question-Guided Scene Text Recognition Via Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To leverage multimodal reasoning for STR, we propose a training-time question-guided STR framework that integrates VQA, termed Q uestion- G uided S cene T ext R ecognition (QG-STR). |
Quanxing Xu; Ling Zhou; Xian Zhong; Feifei Zhang; Rubing Huang; | ACM Transactions on Multimedia Computing, Communications, … | 2026-06-26 |
| 84 | When The Aggregator Cheats: Data-Free Backdoors in Federated LLM-based QA Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we explore the potential vulnerability where a malicious aggregator, who may collude with a third-party vendor, stealthily implants advertisement-type backdoors into federated QA models, without ever accessing client data. |
Chenqing Zhu; Yanbo Dai; Yulong Tian; Qingming Li; Songze Li; | arxiv-cs.CR | 2026-06-25 |
| 85 | Causal Connections: Leveraging Multilingual Fine-Tuning for Financial QA@FinCausal 2026 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper describes team HSA_CORAL’s submission to the FinCausal 2026 shared task on extracting cause-effect relations from financial narratives via extractive question answering in English and Spanish. |
Akash Kumar Gautam; Serhii Hamotskyi; Christian Hänig; | arxiv-cs.CL | 2026-06-25 |
| 86 | TriViewBench: Controlled Complexity Scaling for Multi-View Structural Reasoning in MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TriViewBench, a controlled three-view visual reasoning benchmark constructed from synthetic 3D scenes with explicitly parameterized object count and occlusion. |
Yu-Yang Chen; Lan-Zhe Guo; | arxiv-cs.CV | 2026-06-24 |
| 87 | Automatic Prompt Generation Via Semantic Decomposition-and-recomposition for Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Seungyeon Lee; Dong-Gyu Lee; | Engineering Applications of Artificial Intelligence | 2026-06-24 |
| 88 | ViTexQA: A Multi-Frame Temporal Perception Dataset for Video Text Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This perception challenge fundamentally differs from static image text understanding, yet existing datasets fail to capture: the vast majority of questions remain answerable from single frames, inadequately reflecting real-world video text comprehension demands. To address this, we present ViTexQA, a large-scale video-text QA dataset, and FrameThinker for robust multi-frame temporal reasoning. |
ZHENTAO GUO et. al. | arxiv-cs.CV | 2026-06-23 |
| 89 | Corrigendum to “Multi-aspect Attentive Text Representations for Simple Question Answering Over Knowledge Base” [Natural Language Processing Journal, Volume 5, (2023), 100035] Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Zhixiang Zeng; Yuefeng Li; Jianming Yong; Xiaohui Tao; Vicky Liu; | Natural Language Processing Journal | 2026-06-23 |
| 90 | EG-VQA: Benchmarking Verifiable Video Question Answering with Grounded Temporal Evidence Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA). Nevertheless, existing benchmarks are … |
Linpeng Huang; Weixing Chen; Zexin Chen; Yang Liu; Liang Lin; | arxiv-cs.CV | 2026-06-23 |
| 91 | Video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present video-SALMONN-R$^3$, the first end-to-end video-LLM that enables re-watch through reinforcement learning without relying on chain-of-thought (CoT) cold-start. |
YIXUAN LI et. al. | arxiv-cs.CV | 2026-06-23 |
| 92 | AdaMem: Learning What to Remember for Personalized Long-Horizon LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that what is worth remembering is role-dependent, and propose \textbf{AdaMem} (Adaptive Memory), a method that \emph{learns what to remember} for each user from feedback. |
Xingyu Chen; Rui Wang; Zhaopeng Tu; Liefeng Bo; | arxiv-cs.CL | 2026-06-19 |
| 93 | CalVerT: Augmenting Agents with Calibrated Verifier Telemetry Improves Action and Learning in Knowledge-Intensive Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This produces two failure modes: committing to confident but unsupported answers, which hurts accuracy, and over-retrieving when the evidence in hand already suffices, resulting in wasted compute. To give agents a more complete picture of the state space they are operating in, we introduce calibrated verifier telemetry (CalVerT), which augments the agent’s state with additional telemetry: a calibrated self-confidence score and a grounding verifier score. |
Ashwin Vinod; Ying Ding; Elias Stengel-Eskin; | arxiv-cs.CL | 2026-06-19 |
| 94 | AgentFinVQA: A Deployable Multi-Agent Pipeline for Auditable Financial Chart QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present AgentFinVQA, a multi-agent pipeline that decomposes each query into planning, OCR, legend grounding, visual inspection, and verification, recording every step in a traceable Model Evaluation Packet (MEP) per sample. |
Aravind Narayanan; Shaina Raza; | arxiv-cs.AI | 2026-06-18 |
| 95 | Fine-grained Human Motion Understanding with Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose \methodname, an LLM-based model for fine-grained human motion understanding that represents motion as a sequence of skeletal poses with explicit timestamps for each pose. |
Thomas Markhorst; Zhi-Yi Lin; Jouh Yeong Chew; Jan van Gemert; Xucong Zhang; | arxiv-cs.CV | 2026-06-18 |
| 96 | WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance. |
YEHANG ZHANG et. al. | arxiv-cs.AI | 2026-06-17 |
| 97 | Code-Switching Reveals Language Anchoring in Multilingual LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To understand this degradation, we use grammar-forced CS as a controlled diagnostic setting for locating CS representations relative to their source and target counterparts. We introduce Anchor Bias, a geometric measure that quantifies language anchoring, whether a CS hidden state aligns closer to its source or target language counterpart. |
Jeonghyun Park; Seunghyun Yoon; Yonghyun Jun; Hwanhee Lee; | arxiv-cs.CL | 2026-06-17 |
| 98 | Beyond Tokenization: Direct Timestep Embedding and Contrastive Alignment for Time-Series Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Prior methods use patch-based encoders that split the series into fixed windows, locking in one granularity that breaks patterns and hides exact timesteps, through a separate module that rarely transfers across datasets with different lengths or sampling rates. To address this challenge, we propose CADE (Contrastive Alignment with Direct Embedding), a novel framework for TSQA built upon two key components: direct timestep embedding and semantic alignment. |
Yafeng Wu; Huu Hiep Nguyen; Thin Nguyen; Hung Le; | arxiv-cs.CL | 2026-06-17 |
| 99 | Trade-offs in Medical LLM Adaptation: An Empirical Study in French QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a study of medical domain adaptation using French medical question-answering (QA) as a case study. |
IKRAM BELMADANI et. al. | arxiv-cs.CL | 2026-06-17 |
| 100 | Optimal Scheduling in A Question-Answering Forum of Knowledge Workers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With this model, we calculate the capacity of the system for handling the requests while keeping the system stable, and design schedulers that achieve capacity. |
Rohit Negi; Mustafa Yilmaz; | arxiv-cs.AI | 2026-06-17 |
| 101 | From Drift to Coherence: Stabilizing Beliefs in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce prompted predictive resampling (PPR), where an LLM generates a sequence of answers to the same question. |
SongEun Kim; Seungyoo Lee; Edwin Fong; Hyungi Lee; Juho Lee; | arxiv-cs.LG | 2026-06-16 |
| 102 | Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work proposes BinTrack, a simple yet effective, fully open-source spatial-localization agent that leverages the temporal ordering of a robot’s trajectory. |
DONGBIN NA et. al. | arxiv-cs.RO | 2026-06-15 |
| 103 | Compositional Reasoning Depth Predicts Clinical AI Failure: Empirical Evidence Consistent with Transformer Compositionality Limits in Electronic Health Record Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by theoretical results on transformer compositionality limits, we introduce a pre-specified hop-count taxonomy — the number of distinct reasoning steps required to answer a clinical question from an EHR — as a principled predictor of model failure. |
Sanjay Basu; | arxiv-cs.CL | 2026-06-15 |
| 104 | JARVIS for HVAC: LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present JARVIS , a two-stage LLM-based QA framework tailored for sensor data-driven HVAC system interaction. |
SUNGMIN LEE et. al. | Proceedings of the ACM on Interactive, Mobile, Wearable and … | 2026-06-15 |
| 105 | Weaving Multi-Source Evidence for Biomedical Reasoning: The BioMedHop Benchmark and BioWeave Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing biomedical QA benchmarks mainly focus on exam-style knowledge, literature comprehension, or short-range multi-hop inference, leaving source-conditioned graph reasoning and evidence topology construction underexplored. To fill this gap, we introduce BioMedHop, a multi-source graph-grounded benchmark for evaluating biomedical reasoning over structured evidence topologies. |
XINGYU TAN et. al. | arxiv-cs.CL | 2026-06-15 |
| 106 | A Self Consistency Based Reranking for Narrative Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite recent advances in pretrained language models, most existing approaches rely on a single decoding output during inference, making them sensitive to generation variability and often resulting in incomplete or inconsistent answers . To address this limitation, we propose a self-ensemble Self-Consistency-Based reranking framework for narrative question answering. |
Molham Mohamed; Ali Hamdi; | arxiv-cs.CL | 2026-06-14 |
| 107 | EHRNote-ChatQA: A Benchmark for Evidence-Grounded Multi-Turn Clinical Question Answering Over Longitudinal Discharge Summaries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce EHRNote-ChatQA, the first benchmark for evidence-grounded multi-turn clinical question answering over patients’ multiple discharge summaries. |
JIYOUN KIM et. al. | arxiv-cs.CL | 2026-06-14 |
| 108 | MAGE-RAG: Multigranular Adaptive Graph Evidence for Agentic Multimodal RAG in Long-Document QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes MAGE-RAG, a multigranular adaptive graph evidence framework for long-document multimodal QA. |
YILONG ZUO et. al. | arxiv-cs.IR | 2026-06-14 |
| 109 | Object Tokens As A Bridge Between Segmentation and Visual Question Answering in Robotic Surgery Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a unified framework that jointly performs pixel-level segmentation and visual question answering within a single framework. |
YIPING LI et. al. | arxiv-cs.CV | 2026-06-14 |
| 110 | Medical Visual Question Answering with Multimodal: A Systematic Mini Review (2023–2026) Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study aims to systematically analyze recent developments in Med-VQA. |
MAIMUNA BISWAS NOSHIN et. al. | Frontiers in Digital Health | 2026-06-12 |
| 111 | OmniVideo-100K: A Dataset for Audio-Visual Reasoning Through Structured Scripts and Evidence Chains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, coupling long-text comprehension and QA synthesis into a single step often restricts models to localized events, yielding questions lacking long-term temporal connections and deep cross-modal reasoning. To address these issues, we propose an automated data engine featuring two mechanisms: (1) \textbf{Entity-Anchored Video Scripting} transforms videos into structured scripts, comprising summaries, main entity lists, and segment-wise audio-visual descriptions. |
Xinyue Cai; Chaoyou Fu; Yi-Fan Zhang; Ran He; Caifeng Shan; | arxiv-cs.CV | 2026-06-12 |
| 112 | ReportQA: QA-Based Radiology Report Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Clinicians use them to perform downstream diagnostic tasks without directly inspecting images. Based on this insight, we propose ReportQA, a clinical-related and flexible radiology report evaluation framework, supporting detailed quantitative analysis of radiology report generation systems. |
YIMING SHI et. al. | arxiv-cs.CL | 2026-06-12 |
| 113 | Operads for Compositional Reasoning in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose operads, mathematical structures that model many-in, one-out operations and compositions thereof, as a natural framework for describing question decomposition. |
Nathaniel Bottman; Kyle Richardson; | arxiv-cs.CL | 2026-06-11 |
| 114 | An Inference Framework for Integrating Visual Question Answering and Object Distance Estimation on Monocular Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Joon Cho; Choulsoo Jang; Changeun Lee; Kyunam Kim; | The Journal of Korea Robotics Society | 2026-06-11 |
| 115 | Lung-R1: A Knowledge Graph-Guided LLM for Pulmonary Diagnostic Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Built on LungKG, we propose Lung-R1, a LungKG-guided pulmonary LLM trained through KG-constrained reasoning-chain construction and KG-guided reinforcement learning. |
HAOYANG ZENG et. al. | arxiv-cs.AI | 2026-06-10 |
| 116 | How Fine-Grained Should A RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present HieraRAG, a hierarchical framework for studying granularity in RAG benchmark construction, defining optimal granularity as the level that maximizes discriminative power (the standard deviation of generation quality across categories) within a given RAG configuration. |
Chase M. Fensore; Kaustubh Dhole; Jason Fan; Eugene Agichtein; Joyce C. Ho; | arxiv-cs.CL | 2026-06-10 |
| 117 | One Token Per Multimodal Evidence: Latent Memory for Resource-Constrained QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Latent Memory, a latent-space memory paradigm that replaces each raw text or image evidence item with a single high-dimensional latent token produced by a small compressor LLM/VLM. |
Zhi Zheng; Ziqiao Meng; Hao Luan; Wei Liu; Wee Sun Lee; | arxiv-cs.AI | 2026-06-09 |
| 118 | LakeQA: An Exploratory QA Benchmark Over A Million-Scale Data Lake Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce LakeQA, a comprehensive benchmark for search-centric question answering over data lakes that jointly emphasizes searching and reasoning capabilities. |
HAONAN WANG et. al. | arxiv-cs.CL | 2026-06-09 |
| 119 | ScoutVLA: UAV-Centric Active Perception Via A Dual-Expert VLA Model for Open-World Embodied Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Drawing inspiration from the “waggle dance” of scout bees, which iteratively adjust their flight paths to verify target information, we propose ScoutVLA, an evidence-driven Vision-Language-Action model for outdoor EQA. |
WENHAO LU et. al. | arxiv-cs.CV | 2026-06-09 |
| 120 | Comparison of Three Large Language Models in Postoperative Rehabilitation Question Answering After Anterior Cruciate Ligament Reconstruction Based on Expert Ratings Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
TIANGE XIA et. al. | BMC Musculoskeletal Disorders | 2026-06-09 |
| 121 | Trace Only What You Need: Structure-Aware On-Demand Hypergraph Memory for Long-Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although retrieval-augmented generation (RAG) reduces the input context by retrieving relevant evidence, existing structured RAG methods still face three limitations: costly query-agnostic knowledge organization, insufficient use of original document structure, and no reuse of historical reasoning experience. To address these limitations, we propose DocTrace, a multi-agent RAG framework for long-document QA that supports query-triggered knowledge organization, document-structure-aware and experience-guided reasoning. |
Xiangjun Zai; Xingyu Tan; Chen Chen; Xiaoyang Wang; Wenjie Zhang; | arxiv-cs.CL | 2026-06-09 |
| 122 | TABVERSE: Benchmarking Cross-Format Table Understanding in LLMs and VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TABVERSE, a controlled multimodal table benchmark that aligns the same table content across multiple structural formats and rendered images, with question category and difficulty tags. |
Momina Ahsan; Sarfraz Ahmad; Ming Shan Hee; Roy Ka-Wei Lee; Preslav Nakov; | arxiv-cs.AI | 2026-06-08 |
| 123 | Where Does The Answer Come From? Benchmarking View-Level Visual Evidence Identification in Multi-View MLLMs for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a multi-view visual question answering benchmark for evaluating evidence-source identification: given six synchronized NuScenes views and a question, the model must identify the supporting camera view and answer the question. |
YIMU WANG et. al. | arxiv-cs.CL | 2026-06-08 |
| 124 | MMClima: A Framework for Multimodal Climate Science Data and Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MMClima, a large-scale multimodal climate question answering framework with 104k+ expert-validated question-answer pairs spanning articles, video transcriptions, and figures across five core climate science domains. |
Muhammad Umer Sheikh; Hassan Abid; Khawar Shehzad; Ufaq Khan; Muhammad Haris Khan; | arxiv-cs.LG | 2026-06-08 |
| 125 | ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. |
YI ZHANG et. al. | arxiv-cs.CV | 2026-06-07 |
| 126 | Ibcl: Instance-aware Bias Calibration with Contrastive Learning for Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Chuanfeng Liu; Hao Wu; Benxue Sun; Zhijun Fang; | Multimedia Systems | 2026-06-06 |
| 127 | SPARC: A Multi-Agent System for Electrical Circuit Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SPARC, a multi-agent system that answers questions over circuit diagrams by grounding reasoning in executable physics-based simulations. |
MUSHTARI SADIA et. al. | arxiv-cs.AI | 2026-06-05 |
| 128 | A COMPARATIVE EVALUATION OF LARGE LANGUAGE MODELS FOR PATIENT-FACING MEDICAL QUESTION ANSWERING: ACCURACY, BIAS, AND CLINICAL SAFETY ASSESSMENT Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Vinodhini Ravikumar; | INTERNATIONAL JOURNAL OF ARTIFICIAL INTELLIGENCE IN MEDICINE | 2026-06-05 |
| 129 | Scene Graph-guided Uncertainty Decomposition Improves Confidence Calibration in Surgical Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods typically model predictive uncertainty as a single global quantity, which fails to capture the hierarchical uncertainty arising from ambiguous objects, uncertain tool–tissue interactions, and complex scene context. To address this limitation, we propose a scene graph-guided framework that decomposes uncertainty into object-level, relation-level, and scene-level components and adaptively fuses them for confidence-aware surgical visual question answering. |
JUNZHUO SONG et. al. | Frontiers in Medicine | 2026-06-05 |
| 130 | Explicit Evidence Grounding Via Structured Inline Citation Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Properly attributing information through citations becomes, therefore, crucial. This work introduces FullCite, a framework that, in contrast to most previous works, generates structured inline citations linking each claim to both its source document and supporting evidence. |
Anar Yeginbergen; Amelie Wührl; Anna Rogers; Rodrigo Agerri; | arxiv-cs.CL | 2026-06-05 |
| 131 | EASE-TTT: Evidence-Aligned Selective Test-Time Training for Long-Context Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Evidence-Aligned SElective Test-Time Training (EASE-TTT), a within-context retrieval-augmented test-time training framework that converts selected evidence chunks into a soft attention supervision target over their token positions. |
XIAOPENG YUAN et. al. | arxiv-cs.CL | 2026-06-05 |
| 132 | Improving Answer Extraction in Context-based Question Answering Systems Using LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a question answering system based on large language models, where the input consists of a textual context and a corresponding question, and the output is a concise and accurate answer. |
Hafez Abdelghaffar; Ahmed Alansary; Ali Hamdi; | arxiv-cs.CL | 2026-06-04 |
| 133 | CRAFT: A Unified Counterfactual Reasoning Framework for Tabular Question Answering and Fact Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose CRAFT, a unified Counterfactual Reasoning Framework that reformulates Tabular question answering and fact verification into a general bidirectional verification process. |
CHENSHUO PAN et. al. | arxiv-cs.CL | 2026-06-04 |
| 134 | Reducing Hallucinations in Complex Question Answering Using Simple Graph-based Retrieval-Augmented Generation (long Version) Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore the idea of using a lightweight graph structure with a relatively simple graph schema, to support the RAG subsystem via a dedicated toolset. |
Christopher J. Wedge; Joshua Stutter; Danny Dixon; Jacek Cała; | arxiv-cs.CL | 2026-06-04 |
| 135 | MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they still rely on isolated and fragmented frames as the fundamental evidence units, limiting VLMs’ ability to effectively capture coherent event-level semantics. To address this limitation, we propose MemoryCard, a video-memory-based augmentation framework that organizes long videos into self-contained Memory Cards. |
QING YANG et. al. | arxiv-cs.CV | 2026-06-04 |
| 136 | MARDoc: A Memory-Aware Refinement Agent Framework for Multimodal Long Document QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MARDoc, a Memory-Aware Refinement Agent framework that decouples long-document QA into three specialized agents: an Explorer for multi-granularity multimodal retrieval, a Refiner for distilling interaction traces into structured evidence and reasoning memories, and a Reflector for checking evidence sufficiency and providing targeted feedback. |
KAIFENG CHEN et. al. | arxiv-cs.CL | 2026-06-04 |
| 137 | Chemical-Attribute Extraction Via Inverse Reinforcement Learning with Sub-Reward Matching for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study offers an efficient solution for mining implicit knowledge in chemical texts and provides insights into multi-objective generative tasks. |
Taiyu Zhang; Yuqing Ni; Xicheng Yang; Congyuan Xu; Xiaochen Liu; | Applied Sciences | 2026-06-03 |
| 138 | Can Small Language Models Handle Context-summarized Multi-turn Customer-service QA? A Synthetic Data-driven Comparative Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The main contributions of this work include the application of parameter-efficient fine-tuning to adapt SLMs for context-summarized multi-turn customer-service QA, a synthetic data construction pipeline for generating a context-summarized multi-turn QA dataset, and a structured evaluation framework combining quantitative metrics with human and LLM-as-a-judge assessments for customer-service QA evaluation. |
Lakshan Cooray; Deshan Sumanathilaka; Pattigadapa Venkatesh Raju; | Frontiers in Artificial Intelligence | 2026-06-02 |
| 139 | Hierarchical Context Enhancement for Long-tail Entity Retrieval Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our analysis reveals that the main cause is the absence of a dedicated mechanism for handling low-frequency terms. |
Yixuan Peng; Kewu Pan; | Frontiers in Artificial Intelligence | 2026-06-02 |
| 140 | When Retrieval Doesn’t Help: A Large-Scale Study of Biomedical RAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Across five models, ten biomedical QA datasets, four retrieval methods, and four retrieval corpora, we find that retrieval yields only small and inconsistent improvements over a no-retrieval baseline, typically within 1-2 points. |
Erfan Nourbakhsh; Rocky Slavin; Ke Yang; Anthony Rios; | arxiv-cs.CL | 2026-06-02 |
| 141 | Does Language Shift Break Medical Vision-Language Models? Indonesian Radiology Visual Question Answering Case Study Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce IndoRad-VQA, an Indonesian adaptation of VQA-RAD, to assess whether medical VLMs retain radiology reasoning ability when questions are asked in Bahasa Indonesia. |
Pieter Christy Yan Yudhistira; Dzaki Rafif Malik; Novanto Yudistira; | arxiv-cs.CL | 2026-06-02 |
| 142 | ODTQA-FoRe: An Open-Domain Tabular Question Answering Dataset for Future Data Forecasting and Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This task poses challenges in retrieving precise historical data, overcoming the forecasting limitations of LLMs, and standardizing responses for diverse queries. To solve the above challenges, we propose TimeFore, an LLM agent-based framework that decomposes the problem into three collaborative roles: a Retriever autonomously generates SQL to fetch data, a Forecaster invokes external time-series models for higher accuracy, and an Analyzer synthesizes the results to construct a precise and consistent final answer. |
ZHENSHENG WANG et. al. | arxiv-cs.IR | 2026-06-01 |
| 143 | RASER: Recoverability-Aware Selective Escalation Router for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce RASER (Recoverability-Aware Selective Escalation Router), a family of cheap routers built on one-shot RAG and six features from it. |
Yuyang Li; Zihe Yan; Tobias Käfer; | arxiv-cs.AI | 2026-06-01 |
| 144 | InsightVQA: High-Dimensional Emotion-Cognitive Visual Question Answering Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks mainly focus on emotion recognition, offering limited support for grounded understanding and response-oriented analysis. To address this gap, we introduce \textbf{InsightVQA}, a large-scale dataset for hierarchical visual question answering on emotion understanding and cognitive reasoning. |
SHIYU WANG et. al. | arxiv-cs.CV | 2026-06-01 |
| 145 | A Local Perturbation Theory for Cross-Domain Interference and Recovery in Multi-Domain RL Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that single-domain RL produces sparse, small-magnitude parameter edits with weak overlap among top-changed neurons, while different domains still share substantial active computation routes on which update directions determine whether they act synergistically or conflict. Guided by this observation, we prove under a local perturbation model of multi-domain RL that later-domain training harms an earlier domain mainly through a second-order damage term, which under the observed sparse route structure concentrates in a low-dimensional shared conflict subspace. |
Lei Yang; Siyu Ding; Deyi Xiong; | arxiv-cs.LG | 2026-06-01 |
| 146 | Question-Aware Evidence Ledgers for Video Relational Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a test-time reasoning pipeline built around a strong GPT-5.5 video QA solver and a set of question-aware evidence ledgers. |
Yilin Ou; Mengshi Qi; Huadong Ma; | arxiv-cs.CV | 2026-06-01 |
| 147 | Neuro-Symbolic Legal Guardian: A Hybrid RAG and Knowledge-Graph Framework for Reducing Hallucination in Legal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This manuscript puts forth the Neuro-Symbolic Legal Guardian (NSLG), an integrative model that synthesizes Retrieval-Augmented Generation (RAG) with knowledge graph reasoning to address the issue of hallucination in the domain of legal question answering. |
BHARATHI B et. al. | International Journal of Drug Delivery Technology | 2026-06-01 |
| 148 | Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video. |
Ali Alavi; | arxiv-cs.CV | 2026-05-31 |
| 149 | Adaptive Dense Evidence Refinement for Video Relational Reasoning for VRR-QA Challenge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present an inference-only system built around adaptive test-time computation. |
YUYANG SUN et. al. | arxiv-cs.CV | 2026-05-31 |
| 150 | OCC-RAG: Optimal Cognitive Core for Faithful Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this setting, task-specialized small language models (SLMs) offer a principled design choice. We introduce Optimal Cognitive Core (OCC), a family of SLMs built around this premise. |
MAKSIM SAVKIN et. al. | arxiv-cs.CL | 2026-05-30 |
| 151 | SPADER: Step-wise Peer Advantage with Diversity-Aware Exploration Rewards for Multi-Answer Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose SPADER, a reinforcement learning framework for long-horizon tool use in Multi-Answer QA. |
Qiming Shi; Zhaolu Kang; Yunfan Zhou; Di Weng; Yingcai Wu; | arxiv-cs.CL | 2026-05-30 |
| 152 | HypothesisMed: Inference-Time Answer Fusion and Structured Hypothesis-Space Reporting for Biomedical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents HypothesisMed, an inference-time reliability pipeline for biomedical multiple-choice question answering. |
Md Motaleb Hossen Manik; Ge Wang; | arxiv-cs.CL | 2026-05-30 |
| 153 | Towards Lightweight Reliability: Using Soft Prompts for Hallucination Mitigation in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In high-stakes domains, these errors can reduce trust and introduce real-world risk. To address this challenge, we present a parameter-efficient approach that uses soft prompts to mitigate hallucinated content and promote responsible abstention in generative question-answering (QA) tasks. |
S M Tahmid Siddiqui; Akib Jawad Ononto; Anoop Singhal; Latifur Khan; | arxiv-cs.CL | 2026-05-30 |
| 154 | KG-Guard: Graph-Based Hallucination Detection for Knowledge Base Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formulate hallucination detection in KBQA as an answer-node classification problem and propose a lightweight graph-based framework that treats the answering LLM as a black box. |
Albert Sawczyn; Piotr Bielak; Tomasz Kajdanowicz; | arxiv-cs.LG | 2026-05-29 |
| 155 | AdaptR1: Reinforcement Learning Based Adaptive Interleaved Thinking in Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This overlooks the dynamic nature of multi-step tasks, where the need for explicit reasoning varies across intermediate stages. To address this limitation, we introduce AdaptR1, a Reinforcement Learning (RL) based framework for adaptive interleaved thinking in multi-hop Question Answering (QA). |
YUXIN WANG et. al. | arxiv-cs.CL | 2026-05-29 |
| 156 | Fighting Numerical Hallucinations Via Data-centric Compilation for Online Financial QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we pioneer a data-centric paradigm and propose a novel framework, the Data-centric Reasoning Compiler (DCRC). |
HAO CHEN et. al. | arxiv-cs.IR | 2026-05-29 |
| 157 | Semantic Triplet Restoration: A Novel Protocol for Hierarchical Table Understanding in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Semantic Triplet Restoration (STR), a protocol that rewrites each cell as an atomic fact |
Yibin Zhao; Fangxin Shang; Dingrui Yang; Yuqi Wang; | arxiv-cs.CL | 2026-05-29 |
| 158 | ProtStructQA: A Denotation Threshold in Protein Structural Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ProtStructQA, an executable benchmark for protein structural question answering in which each natural-language question is generated from a hidden typed domain-specific language (DSL) program and the answer is obtained by executing that program on an AlphaFold-predicted structure. |
ARAVIND MANDIGA et. al. | arxiv-cs.CL | 2026-05-29 |
| 159 | Verifiable Rewards Beyond Math and Code: Lightweight Corpus-Grounded Process Supervision for Factual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CorVer (Corpus Verify), a lightweight, plug-in-ready process reward that replaces neural verifiers with a corpus-grounded signal derived from Wikipedia co-occurrence statistics. |
SHICHENG FAN et. al. | arxiv-cs.CL | 2026-05-28 |
| 160 | Brain-IT-VQA: From Brain Signals to Answers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Brain-IT-VQA, a framework for visual question answering from fMRI. |
Roman Beliy; Matias Cosarinsky; Oliver Heinimann; Navve Wasserman; Michal Irani; | arxiv-cs.CV | 2026-05-28 |
| 161 | VitalAgent: A Tool-Augmented Agent for Reactive and Proactive Physiological Monitoring Over Wearable Health Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose VitalAgent, a tool-augmented agentic framework for ECG/PPG-based mHealth that supports both reactive question answering and proactive monitoring. |
DI ZHU et. al. | arxiv-cs.AI | 2026-05-28 |
| 162 | Citation-Closure Retrieval and Per-Rule Attribution for Real-World Regulatory Compliance Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We formalize Regulatory Compliance QA with RegOps-Bench, a novel benchmark featuring an Operational Knowledge Graph derived from complex national R\&D regulations. To address these bottlenecks, we propose RefWalk, a unified framework driven by a shared topic anchor. |
Yeong-Joon Ju; Seong-Whan Lee; | arxiv-cs.AI | 2026-05-28 |
| 163 | OptoChat: A Large Language Model with Retrieval Augmented Generation for Optics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present OptoChat, a retrieval-augmented LLM purpose-built for the optical domain. |
XIAOQING BAO et. al. | Journal of Physics: Photonics | 2026-05-28 |
| 164 | Knowledge Dependency Estimation for Reliable Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textbf{Knot}, a structured rank-aware knowledge dependency estimator. |
Chaodong Tong; Qi Zhang; Nannan Sun; Lei Jiang; Yanbing Liu; | arxiv-cs.CL | 2026-05-27 |
| 165 | Better Later Than Sooner: Neuro-Symbolic Knowledge Graph Construction Via Ontology-grounded Post-extraction Correction Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a neuro-symbolic framework for ontology-grounded KG construction combining open-domain extraction, embedding-based canonicalization of types and predicates, and targeted LLM-based correction of ontology violations. |
Lorenzo Loconte; Timothy Hospedales; Cristina Cornelio; | arxiv-cs.AI | 2026-05-27 |
| 166 | Beyond Chunk-Local Extraction: Cross-Chunk Graph Augmentation for GraphRAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present CrossAug, a GNN-guided CROSS-Chunk Graph AUGmentation method that enriches GraphRAG indices with cross-chunk relational structure as an offline step before query-time retrieval. |
Jiaming Zhang; Yibo Zhao; Jing Yu; Jianxiang Yu; Xiang Li; | arxiv-cs.CL | 2026-05-27 |
| 167 | DocArena: Turning Raw Documents Into Controllable Training Environments for Document Search Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While prior works have made encouraging progress in improving training data quality, existing environments remain predominantly text-based and existing approaches can struggle to construct training environments that are controllable, scalable, and account for multimodal data. Given this, we propose DocArena, a fully automated data curation pipeline building on the practical need for multimodal document search and question-answering. |
JIAMIAN WANG et. al. | arxiv-cs.CV | 2026-05-27 |
| 168 | ConRAG: Consensus-Driven Multi-View Retrieval for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose ConRAG, a consensus-driven multi-view RAG framework that effectively boosts LLMs on complex multi-hop QA. |
Yikai Zhu; Kunfeng Chen; Qihuang Zhong; Juhua Liu; Bo Du; | arxiv-cs.CL | 2026-05-27 |
| 169 | EAPO: Entropy-Driven Adaptive Positive-Negative Sample Weighting for Policy Optimization in Open-Ended QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we systematically investigate the roles of positive and negative samples in reinforcement learning for open-ended QA. |
YUNSHENG ZENG et. al. | arxiv-cs.AI | 2026-05-26 |
| 170 | Query Symbolically or Retrieve Semantically? A Dataset and Method for Semi-Structured Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Symbolic approaches support such operations, but they are often brittle on noisy natural-language corpora. We address this gap with DualGraph, a RAG framework that represents documents through two complementary views: a Textual Knowledge Graph for semantic retrieval and a Symbolic Knowledge Graph for symbolic querying over typed subject–predicate–object triples. |
MATEUSZ CZYŻNIKIEWICZ et. al. | arxiv-cs.AI | 2026-05-26 |
| 171 | Granuscore: A Reference-Free Measure of Granularity for Text Analysis and Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Granuscore, a reference-free measure of granularity that leverages structural properties of a hierarchical embedding space. |
Lukas Ellinger; Alexander Fichtl; Miriam Anschütz; Georg Groh; | arxiv-cs.CL | 2026-05-26 |
| 172 | Chartographer: Counterfactual Chart Generation for Evaluating Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Chartographer, a framework to reverse engineer charts into executable code, validate reconstruction fidelity, generate seed-controlled counterfactual variants, and derive new answers from executable QA logic. |
Yifan Jiang; Dae Yon Hwang; Jesse C. Cresswell; Freda Shi; | arxiv-cs.CL | 2026-05-26 |
| 173 | ForestHG-Trace: Traceable Long-Horizon Ecological Reasoning Over Large-Scale Forest Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ForestHG-Trace, a framework for traceable long-horizon ecological reasoning over forest environments. |
ZIHANG CHENG et. al. | arxiv-cs.CV | 2026-05-26 |
| 174 | AVQANet : Amharic Visual Question and Answering Model Based on Deep Learning Approach for Ethiopian Museum Visitors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: ABSTRACT This study aims to design an effective deep learning‐based Amharic Visual Question Answering (AVQA) model for Ethiopian museum visitors by identifying a suitable algorithm architecture. |
Dires Workie Sefineh; Habtamu Ayalew Aycheh; Afework Abiye Jenber; Fentahun Yirsaw Tiruneh; | Computational Intelligence | 2026-05-26 |
| 175 | Simorgh at SemEval-2026 Task 7: Region-Aware Hybrid Retrieval for Low-Resource Cultural Reasoning in Multilingual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we investigate culturally grounded multiple-choice question answering with the BLEnD benchmark, which consists of a multilingual corpus of 30 languages and covers various socio-cultural domains, such as cuisine, sports, family, etc. |
Hadi Bayrami Asl Tekanlou; Mahdi Bakhtiyarzadeh; Jafar Razmara; | arxiv-cs.CL | 2026-05-26 |
| 176 | Extending Embodied Question Answering from Perception to Decision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present EQA-Decision, a large-scale embodied QA dataset that systematically covers four complementary dimensions of embodied reasoning: static scene construction, spatial understanding, task dynamics reasoning, and instant decision. |
Xicheng Gong; Qiwei Li; Peiran Xu; Yadong Mu; | arxiv-cs.RO | 2026-05-25 |
| 177 | MiRD: Reliable Set-Valued Prediction for Open-Ended Question Answering Via Miscoverage Risk Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce MiRD, a two-stage framework that decomposes overall miscoverage into sampling failure and conditional selection failure. |
Anqi Hu; Zhiyuan Wang; Zijun Jia; Bo Fu; | arxiv-cs.CL | 2026-05-25 |
| 178 | Clarification Is Not Enough: Post-Clarification Answering Remains The Bottleneck in Multi-Turn QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study the problem of preference elicitation in multi-turn question answering by decomposing the problem into two components: a \textbf{clarification policy}, which decides whether to ask a clarifying question or answer directly, and \textbf{post-clarification answering}, which produces the correct final answer once the missing information is provided. |
Jinyan Su; Jennifer Healey; | arxiv-cs.CL | 2026-05-24 |
| 179 | Knowing But Not Showing: LLMs Recognize Ambiguity But Rarely Ask Clarifying Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Doing so requires two abilities: recognizing that a query is ambiguous, and acting on that recognition by seeking clarification instead of answering directly. To study these abilities, we evaluate models on ambiguous, unambiguous, and disambiguated questions in three settings: standard question answering, explicit ambiguity judgment, and behavioral analysis, where a judge model classifies responses as direct answers, refusals, or clarifying questions. |
Jinyan Su; Claire Cardie; | arxiv-cs.CL | 2026-05-24 |
| 180 | AstroRAG — A Pagerank-Based Retrieval-Augmented Generation Pipeline for Question Answering in Astronomy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present AstroRAG — a PageRank-based retrieval-augmented generation (RAG) pipeline adapted for question answering in astronomy. |
Zhifeng Wang; Jason Jingshi Li; Kaihao Zhang; Ramesh Sankaranarayana; | arxiv-cs.CV | 2026-05-24 |
| 181 | CyberMaskQA: A Privacy-Aware Benchmark for Evaluating Large Language Models in Cybersecurity Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Furthermore, progress in privacy-preserving QA is hindered by the lack of annotated, context-rich datasets capable of jointly evaluating operational reasoning and privacy preservation. To address this gap, we introduce CYBERMASKQA, a privacy-aware QA benchmark covering key security domains. |
Matilda Gaddi; Jin Noh; Onat Gungor; Tajana Rosing; | arxiv-cs.CR | 2026-05-23 |
| 182 | StepGap: A Hybrid NLI-LLM Checker for Step-Level Evidence-Gap Detectionin Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present \textbf{StepGap}, a hybrid NLI-LLM decision tree that detects step-level evidence gaps in multi-hop QA and emits one of three typed labels: \textsc{Contradicted Claim} (CC), \textsc{Irrelevant Evidence} (IE), or \textsc{Missing Bridge} (MB), each tied to a concrete repair action. |
Yuelyu Ji; Zhuochun Li; Hui Ji; Daqing He; | arxiv-cs.CL | 2026-05-23 |
| 183 | Med-R2: An Adversarial Benchmark for Evidence-Grounded Reasoning in Medical VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Med-R2 Bench, a hierarchical benchmark aligned with the clinical workflow to evaluate adversarial robustness with visual grounding. |
WEN MA et. al. | arxiv-cs.CV | 2026-05-23 |
| 184 | TS-Skill: A Benchmark for Evaluating Analytical Skills in Time-Series Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TS-Skill, a controlled benchmark for evaluating three composable analytical skills in TSQA: temporal scale selection (SK1), temporal localization (SK2), and cross-interval integration (SK3). |
LIYING HAN et. al. | arxiv-cs.CL | 2026-05-23 |
| 185 | Decompose-and-Refine: Structured Legal Question Answering with Parametric Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite recent progress in multi-hop QA, existing approaches often rely on reasoning in natural language or retrieval without explicit query reformulation, leaving the vocabulary gap between user questions and statutory text largely unaddressed. To address this challenge, we propose Decompose-and-Refine (DaR), a statute-grounded LQA framework that tightly integrates step-wise question decomposition with parametric knowledge-based query refinement. |
Jihyung lee; Hyounghun Kim; Gary Lee; | arxiv-cs.CL | 2026-05-23 |
| 186 | Decomposing Queries Into Tool Calls for Long-Video Keyframe Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose ToolMerge, a keyframe retrieval method based on decomposition and merging: an Large Language Model (LLM) based planner decomposes the query into tool calls and specifies how their per-tool rankings are merged using boolean operators. |
Michal Shlapentokh-Rothman; Prachi Garg; Yu-Xiong Wang; Derek Hoiem; | arxiv-cs.CV | 2026-05-22 |
| 187 | Asking For An Old Friend: Diagnosing and Mitigating Temporal Failure Modes in LLM-based Statutory Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we present a benchmark of 312 expert-validated, time-sensitive German statutory QA pairs spanning three categories: Post-Cutoff Amendment Questions, Pre-Amendment Questions, and Multi-Provision Pre-Amendment Questions. |
Max Prior; Andreas Schultz; Matthias Grabmair; | arxiv-cs.CL | 2026-05-22 |
| 188 | DRIVESPATIAL: A Benchmark for Spatiotemporal Intelligence in VLMs for Autonomous Driving Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DriveSpatial, a benchmark of 15.6K human-verified QA pairs across 20 tasks from five large-scale AD datasets. |
HAO VO et. al. | arxiv-cs.CV | 2026-05-21 |
| 189 | DeferMem: Query-Time Evidence Distillation Via Reinforcement Learning for Long-Term Memory QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This workflow leaves downstream answerers to denoise retrieved candidates and reconstruct query-specific evidence. We present DeferMem, a long-term memory framework that decouples this problem into high-recall candidate retrieval and query-conditioned evidence distillation. |
Jianing Yin; Tan Tang; | arxiv-cs.CL | 2026-05-21 |
| 190 | GS-QA: A Benchmark for Geospatial Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GS-QA, an extensible geospatial QA benchmark with 2,800 question-answer pairs across 28 templates on top of OpenStreetMap and Wikipedia data, covering a wide range of spatial objects, predicates (including directional and towards filtering), and answer types (entity names, locations, distances, directions, counts, and aggregated areas/lengths). |
Majid Saeedan; Muhammad Shihab Rashid; Ahmed Eldawy; Vagelis Hristidis; | arxiv-cs.DB | 2026-05-21 |
| 191 | HyLoVQA: Dynamic Hypernetwork-Generated Low-Rank Adaptation for Continual Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This often leads to cross-level task interference, hindering accurate adaptation to the current task and object. To address this limitation, we propose HyLoVQA. |
YIRAN WANG et. al. | arxiv-cs.CV | 2026-05-21 |
| 192 | MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question-Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes MuKV, a method that features a multi-grained KV cache compression module and a semi-hierarchical retrieval approach to improve both efficiency and accuracy for long streaming VideoQA. |
Junbin Xiao; Jiajun Chen; Tianxiang Sun; Xun Yang; Angela Yao; | arxiv-cs.CV | 2026-05-21 |
| 193 | $M^3 QuestionIng$: Multi-modal Multi-span Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing systems fall short of aligning with real-world scenarios, where source documents often include both textual and visual content, requiring answers to incorporate images for better comprehension. To address this gap, we propose $M^3QAFrame$, a multi-modal, multi-span medical question-answering framework that leverages visual cues to enhance the generation of comprehensive answers drawn from diverse textual and visual spans. |
ANISHA SAHA et. al. | arxiv-cs.IR | 2026-05-19 |
| 194 | CaMo: Camera Motion Grounded Evaluation and Training for Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose the Spatial Narrative Score (SNS), an evaluation framework that requires VLMs to generate explicit spatial narratives capturing both scene semantics and camera motion, followed by reasoning with a frozen proxy LLM. |
HSIANG-WEI HUANG et. al. | arxiv-cs.CV | 2026-05-19 |
| 195 | FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our extensive evaluation reveals that while proprietary models like GPT-5 achieve respectable performance, current open-source VLMs significantly underperform, struggling particularly with spatial reasoning in multi-person scenes and distinguishing subtle differences in human movements and interactions. To address these identified weaknesses, we propose FineAgent, a modular framework that enhances VLMs by leveraging a Localizer and a Descriptor. |
Gueter Josmy Faure; Min-Hung Chen; Jia-Fong Yeh; Hung-Ting Su; Winston H. Hsu; | arxiv-cs.CV | 2026-05-19 |
| 196 | NeuroQA: A Large-Scale Image-Grounded Benchmark for 3D Brain MRI Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present NeuroQA, a large-scale benchmark for visual question answering in 3D brain magnetic resonance imaging (MRI), with 56,953 QA pairs from 12,977 subjects across 12 datasets. |
MOHAMMAD H. ABBASI et. al. | arxiv-cs.CV | 2026-05-19 |
| 197 | SADL: Sampling, Deliberation, and Pseudo-labeling for In-context Learning in Compositional Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Long Hoang Dang; Thao Minh Le; Vuong Le; Tu Minh Phuong; Truyen Tran; | Discover Artificial Intelligence | 2026-05-19 |
| 198 | Skyline Retrieval Meets Set-Cover Chunk Merging: A Cost-Effective RAG-Sketch for Long-Context LLM QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a cost-effective Retrieval-Augmented Generation Sketch (RAG-Sketch) for long-context LLM QA. |
Xinyi Zhu; Haoyang Li; Yongqi Zhang; Lei Chen; | Proceedings of the ACM on Management of Data | 2026-05-18 |
| 199 | The Evolution and Open Challenges of Text-based Visual Question Answering: A Review of Research and Data Trends Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Kobra Farshidi; Hassan Khotanlou; Elham Alighardash; | Multimedia Tools and Applications | 2026-05-18 |
| 200 | SVFSearch: A Multimodal Knowledge-Intensive Benchmark for Short-Video Frame Search in The Gaming Vertical Domain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SVFSearch, the first open benchmark for short-video frame search in the Chinese gaming domain. |
LINGTAO MAO et. al. | arxiv-cs.AI | 2026-05-18 |
| 201 | LogRouter: Adaptive Two-Level LLM Routing for Log Question Answering in Big Data Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present LogRouter, an end-to-end log question-answering system deployed on TUBITAK BILGEM’s national big data platform that combines a PySpark-based Drain3 ingestion pipeline, GPU-accelerated embeddings, and dual-index storage in Apache Druid and PostgreSQL with pgvector. |
Mert Coskuner; Merve Zeybel; Melik Mert Dolan; | arxiv-cs.LG | 2026-05-18 |
| 202 | Unlocking Biological Workflows for Robust Protein-Text Question Answering: A Dual-Dimensional RAG Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, constrained by the standard RAG pipeline, these models often rely on curated, static datasets instead of expert-proven biological workflows, lacking the fine-grained information processing and struggling to generalize to novel (OOD) proteins. To bridge this gap, we propose 2D-ProteinRAG, a novel framework that empowers LLMs to operate within the gold-standard biological research workflow (BLAST). |
LI DING et. al. | arxiv-cs.IR | 2026-05-17 |
| 203 | Analyzing Error Propagation in Korean Spoken QA with ASR-LLM Cascades Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We analyze how automatic speech recognition (ASR) errors propagate through ASR-LLM cascades in Korean spoken question answering (SQA), focusing on downstream semantic failures that conventional ASR metrics cannot fully capture. |
Donghyuk Jung; Youngwon Choi; | arxiv-cs.CL | 2026-05-17 |
| 204 | UCSF-PDGM-VQA: Visual Question Answering Dataset for Brain Tumor MRI Interpretation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a clinically relevant visual question answering (VQA) benchmark — the UCSF-PDGM-VQA dataset — consisting of 2,387 QA pairs from 473 glioma-related MRI studies in the public UCSF-PDGM dataset. |
Shiv Ghosh; Junayd Lateef; Yannan Yu; Andreas M. Rauschecker; Madhumita Sushil; | arxiv-cs.CV | 2026-05-16 |
| 205 | DICE: Disentangling Causal Evidence for Multimodal Textbook Question Answering Via Attentive Embedding Fusion Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Xu Gu; Yu Li; Bingke Zhu; Jinqiao Wang; Xiaolin Qin; | Pattern Recognition | 2026-05-15 |
| 206 | H-Mem: A Novel Memory Mechanism for Evolving and Retrieving Agent Memory Via A Hybrid Structure Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Memory data are ubiquitous in Large Language Model (LLM)-based agents (e.g., OpenClaw and Manus). A few recent works have attempted to exploit agents’memory for improving their … |
Jiawei Yu; Yixiang Fang; Xilin Liu; Yuchi Ma; | arxiv-cs.CL | 2026-05-15 |
| 207 | GRASP: Graph Agentic Search Over Propositions for Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Graph Agentic Search over Propositions (GRASP), an agentic system that simultaneously optimizes for high accuracy and minimal token usage in multi-hop question answering. |
Stockton Jenkins; Ramya Korlakai Vinayak; Junjie Hu; | arxiv-cs.MA | 2026-05-15 |
| 208 | Chain-of-Procedure: Hierarchical Visual-Language Reasoning for Procedural QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through comprehensive analysis, we identify two critical limitations in current VLMs: inadequate cross-modal retrieval of structured procedures given visual states, and misalignment between image sequence granularity and textual step decomposition. To address these issues, we present Chain-of-Procedure (CoP), a hierarchical reasoning framework that first retrieves relevant instructions using visual cues, then performs step refinement through semantic decomposition, and finally generates the next step. |
GUANHUA CHEN et. al. | arxiv-cs.CL | 2026-05-14 |
| 209 | FINESSE-Bench: A Hierarchical Benchmark Suite for Financial Domain Knowledge and Technical Analysis in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present FINESSE-Bench, a suite of eight specialized benchmarks comprising 3,993 questions for hierarchical evaluation of financial competencies in LLMs. |
DMITRY STANISHEVSKII et. al. | arxiv-cs.CL | 2026-05-14 |
| 210 | Distildoc: Deep Reinforcement Learning for Token-Efficient Multimodal Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The primary problem associated with the use of LLM’s is the increase in inference costs based on the number of tokens used for input. This issue tends to become more problematic when dealing with documents that include text, structured tables, and visual content. |
Patchipulusu Gayathri Asritha; Somala Kanth Mani Sai; Surampalli Bharat Sai; Dr. M. Sreelatha; | International Research Journal on Advanced Engineering Hub … | 2026-05-13 |
| 211 | Retrieval Is Cheap, Show Me The Code: Executable Multi-Hop Reasoning for Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We observe that multi-hop question answering is a typical form of step-by-step computation, and that this structured process aligns closely with how code-specialized language models are trained to operate. Motivated by this, we introduce \pyrag, a framework that reformulates multi-hop RAG as program synthesis and execution. |
JIASHUO SUN et. al. | arxiv-cs.AI | 2026-05-13 |
| 212 | Context Convergence Improves Answering Inferential Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Large Language Models (LLMs) are widely used in open-domain Question Answering (QA), their ability to handle inferential questions-where answers must be derived rather than directly retrieved-remains still underexplored. This study investigates how the structure and quality of passages influence LLM performance on such questions. |
Jamshid Mozafari; Bhawna Piryani; Adam Jatowt; | arxiv-cs.CL | 2026-05-12 |
| 213 | MedHopQA: A Disease-Centered Multi-Hop Reasoning Benchmark and Evaluation Framework for LLM-Based Biomedical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Evaluating large language models (LLMs) in the biomedical domain requires benchmarks that can distinguish reasoning from pattern matching and remain discriminative as model capabilities improve. |
REZARTA ISLAMAJ et. al. | arxiv-cs.CL | 2026-05-12 |
| 214 | Training LLMs with Reinforcement Learning for Intent-Aware Personalized Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This limits their effectiveness in single-turn settings, where the user’s latent goal must be inferred from minimal input and integrated into the thinking and reasoning process. To bridge this gap, we propose IAP (Intent-Aware Personalization), a reinforcement learning framework that trains models to infer implicit user intent directly from a single-turn question and incorporate it into thinking steps through a tag-based schema for generating personalized, intent-grounded answers. |
Maryam Amirizaniani; Benjamin Charles Germain Lee; Jevin West; Nicholas Weber; | arxiv-cs.CL | 2026-05-12 |
| 215 | Overview of The MedHopQA Track at BioCreative IX: Track Description, Participation and Evaluation of Systems for Multi-hop Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We developed a novel dataset of 1,000 challenging QA pairs spanning diseases, genes, and chemicals, with particular emphasis on rare diseases. |
REZARTA ISLAMAJ et. al. | arxiv-cs.CL | 2026-05-12 |
| 216 | SEP-LLM: Professional QA in The SEP Domain Using Retrieval-Augmented LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing general-purpose large language models show limitations in this field, mainly in knowledge retrieval accuracy, semantic matching, and legal compliance of generated content. Therefore, there is an urgent need to develop a specialized intelligent QA system tailored for the SEP domain.OBJECTIVES: This paper aims to develop an intelligent QA system for the SEP domain, SEP-LLM, to improve knowledge retrieval, semantic matching, and content compliance, providing high-quality automated answers to SEP-related questions.METHODS: We collected and curated a large set of SEP-related regulations, technical standards, and judicial cases to build a high-quality QA dataset. |
CHENCHEN GUO et. al. | ICST Transactions on Scalable Information Systems | 2026-05-11 |
| 217 | ASTRA-QA: A Benchmark for Abstract Question Answering Over Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this setting is still poorly supported by existing benchmarks and evaluation methods, which often lack stable abstract references or rely on coarse similarity metrics and unstable head-to-head comparisons. To alleviate this issue, we introduce ASTRA-QA, a benchmark for AbSTRAct Question Answering over documents. |
SHU WANG et. al. | arxiv-cs.CL | 2026-05-11 |
| 218 | Intravenous Extravasation Report Generation Using Deep Learning, Generative Artificial Intelligence, and Visual Question Answering Techniques Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
ADIREK MUNTHULI et. al. | Scientific Reports | 2026-05-11 |
| 219 | Neural at ArchEHR-QA 2026: One Method Fits All: Unified Prompt Optimization for Clinical QA Over EHRs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present Neural1.5, our method for the ArchEHR-QA 2026 shared task at CL4Health@LREC 2026, which comprises four subtasks: question interpretation, evidence identification, answer generation, and evidence alignment. |
Abrar Majeedi; Viswanatha Reddy Gajjala; Sai Prasanna Teja Reddy Bogireddy; Siddhant Rai; | arxiv-cs.CL | 2026-05-11 |
| 220 | AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. |
ROBIN LINZMAYER et. al. | arxiv-cs.AI | 2026-05-11 |
| 221 | How Should LLMs Listen While Speaking? A Study of User-Stream Routing in Full-Duplex Spoken Dialogue Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that how the user stream is routed into the LLM is therefore a key architectural question for full-duplex modeling. To study this question, we extend a text-only LLM into a unified full-duplex spoken dialogue system and compare two routing strategies under a shared training pipeline: (i) channel fusion, which injects the user stream directly into the LLM input, and (ii) cross-attention routing, which keeps the user stream as external memory accessed through cross-attention adapters. |
HUI LU et. al. | arxiv-cs.CL | 2026-05-11 |
| 222 | Assessment of RAG and Fine-Tuning for Industrial Question-Answering-Applications Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We extend the Cost-of-Pass framework proposed by Erol et al. (arXiv:2504.13359) to jointly assess output quality, generation cost, and user interaction cost. |
JAKOB STURM et. al. | arxiv-cs.CL | 2026-05-10 |
| 223 | SearchSkill: Teaching LLMs to Use Search Tools with Evolving Skill Banks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \Ours, a framework that makes query planning explicit through reusable search skills. |
Jinchao Hu; Meizhi Zhong; Kehai Chen; Min Zhang; | arxiv-cs.AI | 2026-05-09 |
| 224 | A Study on Question Answering in The Ceramics Domain Using Large Language Models with Retrieved Triples and Generated Textual Knowledge Prompts Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Qixian Zhang; Fubao He; Kaihua Hu; Juan Li; | Scientific Reports | 2026-05-09 |
| 225 | Lost in Volume: The CT-SpatialVQA Benchmark for Evaluating Semantic-Spatial Understanding of 3D Medical Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This uncertainty stems from the lack of systematic evaluation of semantic-spatial reasoning in volumetric medical VLMs for clinically reliable decision support. To address this gap, we introduce CT-SpatialVQA, a benchmark designed to evaluate semantic-spatial reasoning in 3D CT data. |
Mashrafi Monon; Umaima Rahman; Asif Hanif; Numan Saeed; Mohammad Yaqub; | arxiv-cs.CV | 2026-05-09 |
| 226 | A Semantic-Sampling Framework for Evaluating Calibration in Open-Ended Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Open-ended question answering (QA), the most common deployment setting for modern LLMs, is where existing evaluation methods fall short: logit-based metrics need restricted output formats and internal probabilities; verbalized confidence is self-reported and often overconfident; and sampling-based methods rely on task-specific extraction rules without a clear finite-sample target. We introduce Sem-ECE (Semantic-Sampling Expected Calibration Error), a calibration evaluation framework for open-ended QA that samples answers from the model, groups them into semantic classes, and uses the resulting frequencies as confidence. |
ZHANLIANG WANG et. al. | arxiv-cs.CL | 2026-05-08 |
| 227 | Multi-agent Decision Making: A Blackwell’s Informativeness Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we provide a principled approach to analyse decisions made in the multi-LLM setting using Blackwell’s informativeness framework. |
Zheng Zhang; Cuong C. Nguyen; Kevin Wells; Gustavo Carneiro; | arxiv-cs.LG | 2026-05-07 |
| 228 | Towards Self-Explainable Document Visual Question Answering with Chain-of-Explanation Predictions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CoExVQA, a self-explainable DocVQA framework with a grounded reasoning process through a chain-of-explanation design. |
Kjetil Indrehus; Adrian Duric; Changkyu Choi; Ali Ramezani-Kebrya; | arxiv-cs.LG | 2026-05-07 |
| 229 | Tatarstan Toponyms: A Bilingual Dataset and Hybrid RAG System for Geospatial Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: This paper addresses automatic geospatial question answering over multilingual toponymic data. An original bilingual dataset of toponyms of the Republic of Tatarstan is … |
Mullosharaf K. Arabov; | arxiv-cs.CL | 2026-05-07 |
| 230 | Evaluation of Large Language Models in Cardiovascular Surgery: A Comparative Study of Board-level Clinical Question Answering and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Mehmet Inanc Yesilkaya; Uzeyir Yilmaz; | Journal of Cardiothoracic Surgery | 2026-05-07 |
| 231 | Inference-Time Budget Control for LLM Search Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Under such dual budgets, better answers require not only stronger models, but also explicit control over which search action should receive the next budget unit and when the accumulated evidence is sufficient to commit a final answer. We study this problem in multi-hop question answering (QA) and formulate it as two-stage inference-time budget control. |
ZHENGRU FANG et. al. | arxiv-cs.AI | 2026-05-07 |
| 232 | Knowledge-Graph Paths As Intermediate Supervision for Self-Evolving Search Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, we observe that constructing and solving a multi-hop question can involve overlapping intermediate entities: the factual bridges used to formulate the question may provide approximate waypoints for answering it. Exploiting this overlap, we introduce Waypoint Coverage Reward (WCR), which grants graded partial credit to incorrect Solver trajectories according to their coverage of entities on the construction path, while preserving full reward for correct answers. |
HUYU WU et. al. | arxiv-cs.AI | 2026-05-07 |
| 233 | VTAgent: Agentic Keyframe Anchoring for Evidence-Aware Video TextVQA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The results suggest that the primary bottleneck lies in the localization of key question-relevant evidence, rather than in reasoning capacity itself. Building on this insight, we propose a question-guided agent framework that explicitly anchors the relevant keyframes before answering. |
Haibin He; Maoyuan Ye; Jing Zhang; Juhua Liu; Bo Du; | arxiv-cs.CV | 2026-05-06 |
| 234 | CANDI: Contextual Alignment for Niche Domains Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a robust baseline, we present MTSS-Net, a lightweight neuro-symbolic framework combining neural retrieval with rule-based reasoning. |
MEGHA CHAKRABORTY et. al. | arxiv-cs.CL | 2026-05-06 |
| 235 | Agentic Retrieval-Augmented Generation for Financial Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FinAgent-RAG, an agentic RAG framework that orchestrates iterative retrieval-reasoning loops with self-verification, specifically engineered for the precision requirements of financial numerical reasoning. |
Yang Shu; Yingmin Liu; Zequn Xie; | arxiv-cs.AI | 2026-05-06 |
| 236 | Temporal Reasoning Is Not The Bottleneck: A Probabilistic Inconsistency Framework for Neuro-Symbolic QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel neuro-symbolic question-answering framework governed by a Probabilistic Inconsistency Signal (PIS) that explicitly isolates perceptual errors from reasoning failures. |
Tran Quang Liem; | arxiv-cs.AI | 2026-05-05 |
| 237 | KARMA-MV: A Benchmark for Causal Question Answering on Music Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a causal knowledge graph (CKG) approach that augments vision-language models (VLMs) with structured retrieval of cross-modal dependencies. |
Archishman Ghosh; Abhinaba Roy; Dorien Herremans; | arxiv-cs.CV | 2026-05-05 |
| 238 | BIT.UA-AAUBS at ArchEHR-QA 2026: Evaluating Open-Source and Proprietary LLMs Via Prompting in Low-Resource QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents the joint participation of the BIT.UA and AAUBS groups in the ArchEHR-QA 2026 shared task, which focuses on clinical question answering and evidence grounding in a low-resource setting. |
RICHARD A. A. JONKER et. al. | arxiv-cs.CL | 2026-05-05 |
| 239 | DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the Dual Process Theory in cognitive science and Stanovich’s Cognitive Misers Theory, we propose an effective multi-hop QA framework DTKG (Dual-Track Knowledge Graph) through building a two-stage pipeline: i) Classification Stage (dynamic question categorization via few-shot prompting, emulating unconscious processing); and ii) Branch Processing Stage (tailored reasoning paths, emulating conscious processing). |
CHANGHAO WANG et. al. | icml | 2026-05-05 |
| 240 | TopBench: A Benchmark for Implicit Prediction and Reasoning Over Tabular Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These queries introduce two challenges: recognizing latent intent and reliable predictive reasoning over massive tables. To assess LLMs in such Tabular questiOn answering with implicit Prediction tasks, we introduce TopBench, a benchmark consisting of 779 samples across four sub-tasks, ranging from single-point prediction to decision making, treatment effect analysis, and complex filtering, requiring models to generate outputs spanning reasoning text and structured tables. |
An-Yang Ji; Jun-Peng Jiang; De-Chuan Zhan; Han-Jia Ye; | icml | 2026-05-05 |
| 241 | LakeQA: A Benchmark for Complex Exploratory QA Over A Million-Scale Data Lake Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce LakeQA, a comprehensive benchmark for search-centric question answering over data lakes that jointly emphasizes \emph{searching} and \emph{reasoning} capabilities. |
HAONAN WANG et. al. | icml | 2026-05-05 |
| 242 | Question Answering Dataset for Information Retrieval in Slovak Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study introduces a benchmark for evaluating information retrieval in the Slovak language, leveraging a question-answering dataset for fine-tuning and assessment of sentence transformers. |
Daniel Hládek; Michal Stromko; Kristián Sopkovič; Matúš Pleva; Ming-Hsiang Su; | PeerJ Computer Science | 2026-05-05 |
| 243 | Holi-Spatial: Evolving Video Streams Into Holistic 3D Spatial Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose \textbf{Holi-Spatial}, the first fully automated, large-scale, spatially-aware multimodal dataset, constructed from raw video inputs without human intervention, using the proposed data curation pipeline. |
YUANYUAN GAO et. al. | icml | 2026-05-05 |
| 244 | Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present $\textit{Video-in-the-Loop}$ (ViTL), a two-stage long-video QA framework that preserves a fixed token budget by first $\textit{localizing}$ question-relevant interval(s) with a low-fps skim and then $\textit{answering}$ via span-aware reallocation of visual tokens at higher effective frame rate, emitting an interleaved output with both spans and the final option for direct attribution. |
CHENDONG WANG et. al. | icml | 2026-05-05 |
| 245 | Solving Physics Olympiad Via Reinforcement Learning on Physics Simulators Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we show that physics simulators can serve as a powerful alternative source of supervision for training LLMs for physical reasoning. |
MIHIR PRABHUDESAI et. al. | icml | 2026-05-05 |
| 246 | A Linear Expectation Constraint for Selective Prediction and Routing with False-Discovery Control Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose LEC, a principled framework that reframes selective prediction as a decision problem governed by a linear expectation constraint over selection and error indicators. |
ZHIYUAN WANG et. al. | icml | 2026-05-05 |
| 247 | EEmo-Logic: A Unified Dataset and Multi-Stage Framework for Comprehensive Image-Evoked Emotion Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing models are still limited to coarse-grained emotion perception or deficient reasoning capabilities. To bridge this gap, we introduce **EEmoDB**, the largest image-evoked emotion understanding dataset to date. |
LANCHENG GAO et. al. | icml | 2026-05-05 |
| 248 | OpenTSLM: Time-Series Language Models for Reasoning Over Multivariate Medical Text- and Time-Series Data IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose OpenTSLM, a family of Time Series Language Models that integrate time-series as a native modality into pretrained LLMs, enabling natural-language prompting and reasoning over multiple time-series. |
PATRICK LANGER et. al. | icml | 2026-05-05 |
| 249 | ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We improve evaluation validity by introducing ReVSI, a benchmark and protocol that ensures each QA pair is answerable and correct under the model’s actual inputs. |
YIMING ZHANG et. al. | icml | 2026-05-05 |
| 250 | Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Flexi-LoRA, a novel framework that dynamically adjusts LoRA ranks based on input complexity during both training and inference. |
Z. Li; Y. Su; H. Zhou; Z. Fu; N. Collier; | icassp | 2026-05-04 |
| 251 | CE-GOCD: Central Entity-Guided Graph Optimization for Community Detection to Augment LLM Scientific Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This impairs the LLM’s comprehension of scientific literature, hindering the comprehensiveness and specificity of its responses. To address this, we propose Central Entity-Guided Graph Optimization for Community Detection (CE-GOCD), a method that augments LLMs’ scientific question answering by explicitly modeling and leveraging semantic substructures within academic knowledge graphs. |
J. Lan; | icassp | 2026-05-04 |
| 252 | Multi-Domain Audio Question Answering Benchmark Toward Acoustic Content Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a multi-domain Audio Question Answering (MD-Audio) benchmark designed to advance research in multi-domain sound understanding. |
C. -H. H. Yang; | icassp | 2026-05-04 |
| 253 | Knowledge Editing with Demonstration Selection for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, experiments indicate that existing KE methods for MQA often perform poorly, primarily because the pre-trained LLMs struggle to decompose sub-questions when relying on the same demonstration template. To address this issue, we propose Knowledge Editing with Demonstration Selection for Multi-hop Question Answering (KEDS), which dynamically selects the most suitable demonstrations using both similarity and latent concept modeling for any given MQA query. |
H. Zhao; J. Li; C. Tan; E. Yilmaz; S. Liang; | icassp | 2026-05-04 |
| 254 | Biomed-R2: Joint Diversity Retrieval and Evidence Reasoning for Biomedical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Typical retrieval-augmented generation (RAG) and its variants suffer from insufficient patient alignment, weak evidence integration, and fragile reasoning when dealing with biomedical question answering (QA). To address these challenges, we propose Biomed-R2, a two-stage RAG framework. |
H. GUAN et. al. | icassp | 2026-05-04 |
| 255 | Exploiting Latent and Implicit Chain of Thought for Efficient Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we examine whether implicit and latent CoT methods can be applied effectively to multi-hop QA, focusing on COCONUT (Chain of Continuous Thought) and implicit stepwise CoT internalization. |
H. Shakil; V. Srinivasan; H. Jeelani; K. Gunaratna; S. Chappidi; | icassp | 2026-05-04 |
| 256 | Compound-QA: A Benchmark for Evaluating LLMs on Compound Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Compound Question Synthesis (CQ-Syn) to build Compound-QA, a benchmark targeting questions composed of multiple interrelated sub-questions. |
Y. Hou; | icassp | 2026-05-04 |
| 257 | DaPT: A Dual-Path Framework for Multilingual Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Retrieval-augmented generation (RAG) systems have made significant progress in solving complex multi-hop question answering (QA) tasks in the English scenario. However, RAG … |
Y. Wang; | icassp | 2026-05-04 |
| 258 | KG2QA: Knowledge Graph-Enhanced Retrieval-Augmented Generation for Communication Standards Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The rapid evolution of communication technologies has led to an explosion of standards, rendering traditional expert-dependent consultation methods inefficient and slow. To address this challenge, we propose KG2QA, a question answering (QA) framework for communication standards that integrates fine-tuned large language models (LLMs) with a domain-specific knowledge graph (KG) via a retrieval-augmented generation (RAG) pipeline. |
Z. Luo; W. Wan; T. Zhang; D. Wang; X. Tang; | icassp | 2026-05-04 |
| 259 | ReTools: Reflection-Enhanced Tool Invocation for Domain-Specific QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches such as ReAct and RestGPT partially mitigate these issues yet remain limited in handling multi-step reliability, iterative recovery, and domain robustness. To address these gaps, we propose ReTools, a Tree-of-Thoughts (ToT) based framework that integrates three modules: (1) task planning, which decomposes complex queries into executable subtasks; (2) tool planning, which selects tools and generates accurate parameters while supporting reflective correction; and (3) reflective iteration, which monitors execution results and adapts to domain-specific requirements. |
F. Dong; | icassp | 2026-05-04 |
| 260 | SEARAG: Semantic Entropy-Guided Adaptive Retrieval for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing adaptive retrieval methods rely on output probabilities or heuristics, which poorly capture the model’s true knowledge needs. To address this, we propose Semantic Entropy-based Adaptive RAG (SEARAG), which trains a discriminative model to predict binary semantic entropy from hidden-layer states, quantifying uncertainty in real time. |
D. Yu; Q. Lin; Z. Yang; L. Zhou; | icassp | 2026-05-04 |
| 261 | M3GQA: A Multimodal Multi-Hop and Knowledge Graph-Based Framework for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Large Language Models (LLMs) have significantly advanced this field, effectively tackling these complex tasks remains a major challenge. To address this, we introduce M3GQA, a novel framework that constructs a unified knowledge graph from multimodal data and performs multi-hop reasoning over it. |
S. HU et. al. | icassp | 2026-05-04 |
| 262 | MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This makes it difficult to assess model robustness in real-world settings. We present MedStruct-S, a benchmark specifically designed to evaluate these tasks under unknown keys and OCR noise. |
Yingyun Li; Yu Wang; Haiyang Qian; | arxiv-cs.CL | 2026-05-04 |
| 263 | TRACE: Optimizing Multi-Hop Question Answering Via Confidence-Guided Retrieval Assimilation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by human cognition, we propose TRACE, which generates candidate answers at cognitive transfer points and evaluates their confidence. |
Y. Han; | icassp | 2026-05-04 |
| 264 | AQUA-Bench: Beyond Finding Answers to Knowing When There Are None in Audio Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Such cases are common in real-world settings, where questions may be misleading, ill-posed, or incompatible with the information. To address this gap, we present AQUA-Bench, a benchmark for Audio Question Unanswerability Assessment. |
C. -Y. Kuan; H. -Y. Lee; | icassp | 2026-05-04 |
| 265 | MedSpeak: A Knowledge Graph-Aided ASR Error Correction Framework for Spoken Medical QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose MedSpeak, a novel knowledge graph-aided ASR error correction framework that refines noisy transcripts and improves downstream answer prediction by leveraging both semantic relationships and phonetic information encoded in a medical knowledge graph, together with the reasoning power of LLMs. |
Y. Song; | icassp | 2026-05-04 |
| 266 | SubQRAG: Sub-Question Driven Dynamic Graph Rag Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this broad-view approach often lacks the deep structured reasoning required for complex multi-hop question answering (QA), leading to incomplete evidence and error accumulation. To address these issues, we propose SubQRAG1, a sub-question driven framework that enhances reasoning depth. |
J. Li; | icassp | 2026-05-04 |
| 267 | Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Semantic Reformulation Entropy (SRE), which improves uncertainty estimation in two ways. |
C. TONG et. al. | icassp | 2026-05-04 |
| 268 | Temporal-Aware Heterogeneous Graph Reasoning with Multi-view Fusion for Temporal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel framework with temporal-aware question encoding, multi-hop graph reasoning, and multi-view heterogeneous information fusion. |
W. Wen; | icassp | 2026-05-04 |
| 269 | Reliable Database Question Answering with Collaborative Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose RAISR, a master–auxiliary multi-agent framework for robust SQL parsing. |
M. ZHANG et. al. | icassp | 2026-05-04 |
| 270 | Game-Theoretic Insights Into Multi-Agent LLM Debate for Enhanced Clinical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We implement MAD for multiple-choice medical questions (MedQA) and evaluate it across five frontier language models, comparing single-shot prompting, chain-of-thought (CoT), and MAD with 2 or 3 agents. |
S. Sudhakara; S. Sudhakara; | icassp | 2026-05-04 |
| 271 | Enhancing Audio Question-Answering Performance Through Log-Likelihood Guided Reward Functions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a new reward framework for extracting learning signals from multiple-choice questions (MCQs). |
S. Blouir; | icassp | 2026-05-04 |
| 272 | Pathfinder: MCTS and LLM Feedback-Based Path Selection for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Hence, we propose PATHFINDER, an approach that: (i) uses Monte Carlo Tree Search to generate training path traces, (ii) improves training data quality by filtering erroneous and lengthy traces using sub-answer recall and LLM-as-a-judge verification, and (iii) reformulates sub-queries to handle failed retrieval cases. |
D. P. Maram; K. Gunaratna; V. Srinivasan; H. Jeelani; S. Chappidi; | icassp | 2026-05-04 |
| 273 | SCPRM: A Schema-aware Cumulative Process Reward Model for Knowledge Graph Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the issues, we propose a Schema-aware Cumulative Process Reward Model (SCPRM) that evaluates reasoning paths by conditioning on the reasoning prefix , and incorporating schema distance between current reasoning step and the implicit target parsed from the query, which provides cumulative and future rewards to guide the path explorations. |
Jiujiu Chen; Yazheng Liu; Sihong Xie; Hui Xiong; | arxiv-cs.AI | 2026-05-04 |
| 274 | CoreAnchor-QA: Center-Anchored and Self-Improving for Question-Answer Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Long document question-answering (QA) data is crucial for deploying domain models, yet high-quality corpora are scarce and cross-chunk evidence is hard to align, which makes generated questions and answers drift from the main theme. To address this, we propose CoreAnchor-QA, a framework that builds semantic anchors at document and chunk levels, asks from multiple perspectives around a fixed answer, and applies a two-stage gate based on semantic consistency and judgment scores together with a failure memory to form a self-improving loop. |
M. HUANG et. al. | icassp | 2026-05-04 |
| 275 | Reasoner-Assisted Planning: Enhance The Ability of Graph-Rag to Handle Complex Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently, multi-step retrieval methods informed by question planning have demonstrated notable improvements in performance; nevertheless, several challenges persist: omitting the crucial conditions of the question and some queries are not appropriate for planning. To mitigate these challenges, we introduce RPlanRAG. |
W. Hou; | icassp | 2026-05-04 |
| 276 | DocLayout: Elevating The Role of Complex Layout Understanding in Document Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: An empirical evaluation of 12 LVLMs shows that these challenges persist in complex layout understanding. To address this limitation, we introduce the Layout-aware Reasoning and Document Understanding (LRDU) framework, which combines semantic document segmentation and multi-agent collaboration. |
J. Zhang; | icassp | 2026-05-04 |
| 277 | UniPACT: A Multimodal Framework for Prognostic Question Answering on Raw ECG and Structured EHR Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Language Models (LLMs) offer a powerful reasoning engine for this task but struggle to natively process these heterogeneous, non-textual data types. To address this, we propose UniPACT (Unified Prognostic Question Answering for Clinical Time-series), a unified framework for prognostic question answering that bridges this modality gap. |
J. Tang; T. Xia; Y. Lu; A. Saeed; | icassp | 2026-05-04 |
| 278 | Maistros: A Greek Large Language Model Adapted Through Knowledge Distillation From Large Reasoning Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we focus on Modern Greek, for which only a limited number of question answering (QA) datasets have been proposed, most of which are intended for model evaluation. |
Nikolaos Giarelis; Charalampos Mastrokostas; Nikos Karacapilidis; | arxiv-cs.CL | 2026-05-03 |
| 279 | Flexi-LoRA with Input-Adaptive Ranks: Efficient Finetuning for Speech and Reasoning Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Flexi-LoRA, a novel framework that dynamically adjusts LoRA ranks based on input complexity during both training and inference. |
Zongqian Li; Yixuan Su; Han Zhou; Zihao Fu; Nigel Collier; | arxiv-cs.LG | 2026-05-03 |
| 280 | Dynamic Knowledge Correction Via Abductive for Domain Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Yulin Zhou; Ruizhang Huang; Chuan Lin; Lijuan Liu; Yongbin Qin; | Inf. Process. Manag. | |
| 281 | Video Active Perception: Effective Inference-Time Long-Form Video Understanding with Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by active perception theory, which posits that models gain information by acquiring data that differs from their expectations, we introduce Video Active Perception (VAP), a training-free method to enhance long-form video QA using VLMs. |
MARTIN Q. MA et. al. | arxiv-cs.CV | 2026-05-02 |
| 282 | SF20K Competition 2025: Summary and Findings Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This report presents the results and findings of the first edition of the Short-Films 20K (SF20K) Competition, held in conjunction with the SLoMO Workshop at ICCV 2025. |
Ridouane Ghermi; Xi Wang; Vicky Kalogeiton; Ivan Laptev; | arxiv-cs.CV | 2026-05-02 |
| 283 | SpecVQA: A Benchmark for Spectral Understanding and Visual Question Answering in Scientific Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To effectively reduce token length while preserving essential curve characteristics, we propose a spectral data sampling and interpolation reconstruction approach. |
JIALU SHEN et. al. | arxiv-cs.AI | 2026-04-30 |
| 284 | HealthNLP_Retrievers at ArchEHR-QA 2026: Cascaded LLM Pipeline for Grounded Clinical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The proposed approach uses a multi-stage cascaded pipeline powered by the Gemini 2.5 Pro large language model to interpret patient-authored questions and retrieve relevant evidence from lengthy clinical notes. |
MD BIPLOB HOSEN et. al. | arxiv-cs.CL | 2026-04-29 |
| 285 | DIAGRAMS: A Review Framework for Reasoning-Level Attribution in Diagram QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present DIAGRAMS, a lightweight, schema-driven review framework that decouples interface logic from dataset-specific JSON structures through an internal meta-schema and dataset adapters. |
ANIRUDH IYENGAR KANIYAR NARAYANA IYENGAR et. al. | arxiv-cs.CL | 2026-04-28 |
| 286 | DRAGON: A Benchmark for Evidence-Grounded Visual Reasoning Over Diagrams Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DRAGON, a benchmark for evaluating evidence-grounded visual reasoning in diagrams. |
ANIRUDH IYENGAR KANIYAR NARAYANA IYENGAR et. al. | arxiv-cs.CV | 2026-04-28 |
| 287 | Efficient and Accurate Medical AI: MediLore and MediOut Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing large language models (LLMs) achieve strong performance but are often limited by high memory usage, latency, and inconsistent behavior in handling rare or complex clinical queries. This study addresses these limitations by exploring efficient and robust modeling strategies for medical QA. |
S. Mohamed Rayhan; M. Hariprasath; K. Hemalatha; | Frontiers in Artificial Intelligence | 2026-04-28 |
| 288 | BioGraphletQA: Knowledge-Anchored Generation of Complex QA Datasets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents a principled and scalable framework for systematically generating complex Question Answering (QA) data. |
Richard A. A. Jonker; Bárbara Maria Ribeiro de Abreu Martins; Sérgio Matos; | arxiv-cs.CL | 2026-04-28 |
| 289 | CAN-QA: A Question-Answering Benchmark for Reasoning Over In-Vehicle CAN Traffic Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this formulation abstracts away the temporal and relational structure of CAN traffic and misaligns with real-world forensic workflows, which require systematic reasoning about traffic behavior. To address this gap, we introduce CAN-QA, the first benchmark that reformulates CAN traffic analysis as a question-answering (QA) task. |
Jing Chen; Abhijay Deevi; Onat Gungor; Tajana Rosing; | arxiv-cs.CR | 2026-04-27 |
| 290 | ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We improve evaluation validity by introducing ReVSI, a benchmark and protocol that ensures each QA pair is answerable and correct under the model’s actual inputs. |
YIMING ZHANG et. al. | arxiv-cs.CV | 2026-04-27 |
| 291 | Less Is More: Lightweight Prompt Compression for Question Answering Applications on Edge Devices Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose CORE, a two-stage sentence-level prompt compression method that eliminates the need for SLMs. |
ZIHUAI XU et. al. | arxiv-cs.CL | 2026-04-27 |
| 292 | Using Hierarchical Syntactic Transformer and Speech Act Identification for Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Jui-Feng Yeh; Kuei-Mei Lin; | Multimedia Tools and Applications | 2026-04-27 |
| 293 | An Intelligent Multi-Agent RAG System for Attributed Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multi-agent systems (MAS) offer an efficient design pattern for tackling complex distributed issues by utilising numerous independent agents that work together towards a shared goal. |
Dr. Bhagyashree Dharaskar; | International Journal for Research in Applied Science and … | 2026-04-25 |
| 294 | Development and Comparative Evaluation of Knowledge Graph–enhanced Large Language Models for Domain-specific Question Answering in Nursing Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
JIYUAN SHI et. al. | BMC Nursing | 2026-04-24 |
| 295 | Contexts Are Never Long Enough: Structured Reasoning for Scalable Question Answering Over Long Document Sets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SLIDERS, a framework for question answering over long document collections through structured reasoning. |
Harshit Joshi; Priyank Shethia; Jadelynn Dao; Monica S. Lam; | arxiv-cs.CL | 2026-04-24 |
| 296 | Navigating Large-Scale Document Collections: MuDABench for Multi-Document Analytical QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces the task of analytical question answering over large, semi-structured document collections. |
Zhanli Li; Yixuan Cao; Lvzhou Luo; Ping Luo; | arxiv-cs.CL | 2026-04-24 |
| 297 | BERAG: Bayesian Ensemble Retrieval-Augmented Generation for Knowledge-based Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Bayesian Ensemble Retrieval-Augmented Generation (BERAG), along with Bayesian Ensemble Fine-Tuning (BEFT), as a RAG framework in which language models are conditioned on individual retrieved documents rather than a single combined context. |
Jinghong Chen; Jingbiao Mei; Guangyu Yang; Bill Byrne; | arxiv-cs.CL | 2026-04-24 |
| 298 | Encoder-Free Human Motion Understanding Via Structured Motion Descriptions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by biomechanical analysis, where joint angles and body-part kinematics have long served as a precise descriptive language for human movement, we propose \textbf{Structured Motion Description (SMD)}, a rule-based, deterministic approach that converts joint position sequences into structured natural language descriptions of joint angles, body part movements, and global trajectory. |
Yao Zhang; Zhuchenyang Liu; Thomas Ploetz; Yu Xiao; | arxiv-cs.CV | 2026-04-23 |
| 299 | AUDITA: A New Dataset to Audit Humans Vs. AI Skill at Audio QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing audio question answering benchmarks largely emphasize sound event classification or caption-grounded queries, often enabling models to succeed through shortcut strategies, short-duration cues, lexical priors, dataset-specific biases, or even bypassing audio via metadata and captions rather than genuine reasoning Thus, we present AUDITA (Audio Understanding from Diverse Internet Trivia Authors), a large-scale, real-world benchmark to rigorously evaluate audio reasoning beyond surface-level acoustic recognition. |
TASNIM KABIR et. al. | arxiv-cs.CL | 2026-04-23 |
| 300 | Beyond Single Plots: A Benchmark for Question Answering on Multi-Charts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PolyChartQA, a mid-scale dataset specifically designed for question answering over multi-chart images. |
Azher Ahmed Efat; Seok Hwan Song; Wallapak Tavanapong; | arxiv-cs.CL | 2026-04-23 |
| 301 | An End-to-End Ukrainian RAG for Local Deployment. Optimized Hybrid Search and Lightweight Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents a highly efficient Retrieval-Augmented Generation (RAG) system built specifically for Ukrainian document question answering, which achieved 2nd place in the UNLP 2026 Shared Task. |
Mykola Trokhymovych; Yana Oliinyk; Nazarii Nyzhnyk; | arxiv-cs.CL | 2026-04-23 |
| 302 | HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by evidence that medical ontologies and patient trajectories exhibit hyperbolic geometry, we propose HypEHR, a compact Lorentzian model that embeds codes, visits, and questions in hyperbolic space and answers queries via geometry-consistent cross-attention with type-specific pointer heads. |
Yuyu Liu; Sarang Rajendra Patil; Mengjia Xu; Tengfei Ma; | arxiv-cs.AI | 2026-04-22 |
| 303 | RespondeoQA: A Benchmark for Bilingual Latin-English Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a benchmark dataset for question answering and translation in bilingual Latin and English settings, containing about 7,800 question-answer pairs. |
Marisa Hudspeth; Patrick J. Burns; Brendan O’Connor; | arxiv-cs.CL | 2026-04-22 |
| 304 | T2S-Metrics: Unified Library for Evaluating SPARQL Queries Generated From Natural Language Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present t2s-metrics, an open-source, extensible, and unified evaluation library designed specifically for SPARQL query comparison and execution-based assessment. |
YOUSOUF TAGHZOUTI et. al. | arxiv-cs.IR | 2026-04-22 |
| 305 | FPSBench: A Benchmark for Video Understanding at High Frame Rates Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Modern video-language models are typically trained on videos downsampled to low frames-per-second (FPS), and the most commonly used evaluation benchmarks are designed for low-FPS input as well. To address this shortcoming, we present FPS-Bench, a large video question-answering benchmark designed to evaluate VLMs’ capabilities to understand video at high-frame rates. |
ROHAN CHOUDHURY et. al. | cvpr | 2026-04-21 |
| 306 | ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current approaches struggle to decompose intricate questions into manageable sub-tasks and often fail to leverage specialized processing paths for different document elements. We present ORCA: Orchestrated Reasoning with Collaborative Agents for Document Visual Question Answering, a novel multi-agent framework that addresses these limitations through strategic agent coordination and iterative refinement. |
Aymen Lassoued; Mohamed Ali Souibgui; Yousri Kessentini; | cvpr | 2026-04-21 |
| 307 | Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While recent approaches have explored text-based chain-of-thought (CoT) reasoning for MLLMs, these methods often suffer from limited cross-modal interaction and increased hallucination, especially with longer videos or reasoning chains. To address these challenges, we propose Video Intelligence via Tool-Augmented Learning (VITAL), a novel end-to-end agentic video reasoning framework. |
HAOJI ZHANG et. al. | cvpr | 2026-04-21 |
| 308 | SAHM: A Benchmark for Arabic Financial and Shari’ah-Compliant Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SAHM, a document-grounded benchmark and instruction-tuning dataset for Arabic financial NLP and Shari’ah-compliant reasoning. |
RANIA ELBADRY et. al. | arxiv-cs.CL | 2026-04-21 |
| 309 | SafetyALFRED: Evaluating Safety-Conscious Planning of Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SafetyALFRED, built upon the embodied agent benchmark ALFRED, augmented with six categories of real-world kitchen hazards. |
JOSUE TORRES-FONSECA et. al. | arxiv-cs.AI | 2026-04-21 |
| 310 | MSJoE: Jointly Evolving MLLM and Sampler for Efficient Long-Form Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present MLLM-Sampler Joint-Evolution (MiSJoE), a novel framework that **jointly evolves** the MLLM and a lightweight key-frame sampler for efficient long-form video understanding. |
WENHUI TAN et. al. | cvpr | 2026-04-21 |
| 311 | VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by the advances of Large Multimodal Model to capture higher-order relational structures explicitly novel paradigm of \textit{visualized knowledge representation}, where knowledge graphs are transformed into graphical visualizations that LMMs can directly perceive and reason over. To systematically evaluate this capability, we introduce \textbf{VKG-QA}, a benchmark for \textit{Visual Knowledge Graph-based Question Answering}, covering three major categories and fourteen subtasks. |
YUNTAO DU et. al. | cvpr | 2026-04-21 |
| 312 | MuKV: Multi-Grained KV Cache Compression for Long Streaming Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes MuKV, a method that features a multi-grained KV cache compression module and a semi-hierarchical retrieval approach to improve both efficiency and accuracy for long streaming VideoQA. |
Junbin Xiao; Jiajun Chen; Tianxiang Sun; Xun Yang; Angela Yao; | cvpr | 2026-04-21 |
| 313 | HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce HanDyVQA, a fine-grained video question-answering benchmark that comprehensively covers both the manipulation and effect aspects of HOI. |
Masatoshi Tateno; Gido Kato; Hirokatsu Kataoka; Yoichi Sato; Takuma Yagi; | cvpr | 2026-04-21 |
| 314 | MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and are largely not open-ended, given the difficulty of evaluating free-form answers. In this paper, we introduce a novel open-ended multi-modal VideoQA benchmark, **MovieRecaps** created using movie recap videos—a distinctive type of YouTube content that summarizes a film by presenting its key events through synchronized visual (recap video) and textual (recap summary) modalities. |
Shaden Shaar; Bradon Michael Thymes; Sirawut Chaixanien; Claire Cardie; Bharath Hariharan; | cvpr | 2026-04-21 |
| 315 | Knowing Thyself: Ego-Grounding for Personalized Question-Answering in Egocentric Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding – the ability to understand the camera-wearer in egocentric videos. |
Junbin Xiao; Shenglang Zhang; Pengxiang Zhu; Angela Yao; | cvpr | 2026-04-21 |
| 316 | An Answer Is Just The Start: Related Insight Generation for Open-Ended Document-Grounded QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present InsightGen, a two-stage approach that first constructs a thematic representation of the document collection using clustering, and then selects related context based on neighborhood selection from the thematic graph to generate diverse and relevant insights using LLMs. |
Saransh Sharma; Pritika Ramu; Aparna Garimella; Koyel Mukherjee; | arxiv-cs.CL | 2026-04-21 |
| 317 | VideoAutoThink: Video Auto Reasoning Via Thinking Once, Answering Twice Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we first demonstrate that for RL-trained video models, direct answering often matches or even surpasses CoT performance, despite CoT producing step-by-step analyses at a higher compute cost. Motivated by this, we propose VideoAutoThink, a video understanding framework that adopts a “reason-when-necessary” strategy. |
SHUMING LIU et. al. | cvpr | 2026-04-21 |
| 318 | VRR-QA: Visual Relational Reasoning in Videos Beyond Explicit Cues Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current benchmarks fail to capture this essential aspect of video understanding. To address this gap, we introduce VRR-QA, a benchmark for Visual Relational Reasoning Beyond Explicit Cues. |
SWETHA SIRNAM et. al. | cvpr | 2026-04-21 |
| 319 | Video-Only ToM: Enhancing Theory of Mind in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The impact of LLM hallucinations on such tasks is also underexplored from an interpretability perspective. To address these issues, we introduce VisionToM, a vision-oriented intervention framework designed to strengthen task-aware reasoning. |
SIQI LIU et. al. | cvpr | 2026-04-21 |
| 320 | DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This one-size-fits-all” strategy often neglects model-specific and task-specific preferences, resulting in inaccurate or overly lengthy responses to graph-related queries. To address this, we propose the $\mbox{DynamicGTR}$ framework, which dynamically selects the optimal GTR for each query during inference, thereby enhancing the zero-shot graph QA capabilities of VLMs with a customizable accuracy and brevity trade-off. |
YANBIN WEI et. al. | cvpr | 2026-04-21 |
| 321 | Act Like A Pathologist: Tissue-Aware Whole Slide Image Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we try to bring models closer to how humans actually examine slides. |
WENTAO HUANG et. al. | cvpr | 2026-04-21 |
| 322 | Domain-oriented RAG Assessment (DoRA): Synthetic Benchmarking for RAG-based Question Answering on Defense Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present DoRA (Domain-oriented RAG Assessment), a domain-grounded benchmark built from defense documents that pairs synthetic, intent-conditioned QA (question answering) with auditable evidence passages for attribution. |
BAO GIA DOAN et. al. | arxiv-cs.CL | 2026-04-20 |
| 323 | Retrieval Augmented Generation Framework for The Nepali Legal Domain Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study presents the first application of a Retrieval Augmented Generation based model for Nepali legal question answering using case laws extracted from the Nepal Kanun Patrika digital archive. |
SAMIR WAGLE et. al. | arxiv-cs.CL | 2026-04-20 |
| 324 | PBSBench: A Multi-Level Vision-Language Framework and Benchmark for Hematopathology Whole Slide Image Interpretation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building upon PBSInstr, we develop PBS-VL, a hematopathology-tailored vision-language model for multi-level PBS interpretation at both cell and slide levels. |
YUANLONG WANG et. al. | arxiv-cs.CV | 2026-04-19 |
| 325 | A Hybrid TF-IDF and Knowledge Graph-Enhanced Retrieval-Augmented Generation Framework with Large Language Models for Domain-Aware Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Lilyani Asri Utami; | Journal of Applied Data Sciences | 2026-04-19 |
| 326 | CoGR-MoE: Concept-Guided Expert Routing with Consistent Selection and Flexible Reasoning for Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, unstable routing can lead to inconsistent expert selection in the same question type, while overly stable routing may reduce flexibility. To address this, we propose Concept-Guided Routing framework (CoGR-MoE), which incorporates semantics of the answer options to guide expert selection in the training phase. |
Xiyin Zeng; Yi Lu; Hao Wang; | arxiv-cs.CV | 2026-04-18 |
| 327 | Toward Auditable Urban Soil Management: A Knowledge Graph and LLM Approach Fusing Environmental and Geochemical Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study develops a domain knowledge graph (KG) and a KG-powered question-answering (KBQA) system for urban soil management to organize multi-source evidence and deliver precise, auditable answers to parcel- and pollutant-specific queries. |
XI QIN et. al. | Applied Sciences | 2026-04-17 |
| 328 | DiscoTrace: Representing and Comparing Answering Strategies of Humans and LLMs in Information-Seeking Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DiscoTrace, a method to identify the rhetorical strategies that answerers use when responding to information-seeking questions. |
Neha Srikanth; Jordan Boyd-Graber; Rachel Rudinger; | arxiv-cs.CL | 2026-04-16 |
| 329 | CoPA: Benchmarking Personalized Question Answering with Data-Informed Cognitive Factors Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Accordingly, we introduce CoPA, a benchmark with 1,985 user profiles for fine-grained, factor-level assessment. |
HANG SU et. al. | arxiv-cs.CL | 2026-04-16 |
| 330 | Improving Heart-Focused Medical Question Answering in LLMs Via Variance-Aware Rubric Rewards with GRPO Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate Group Relative Policy Optimization (GRPO) for post-training LLMs on heart-focused medical question answering with rubric-based supervision derived from RaR-Medicine. |
ARASH AHMADI et. al. | arxiv-cs.CL | 2026-04-16 |
| 331 | MM-Doc-R1: Training Agents for Long Document Visual Question Answering Through Multi-turn Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MM-Doc-R1, a novel framework that employs an agentic, vision-aware workflow to address long document visual question answering through iterative information discovery and synthesis. |
JIAHANG LIN et. al. | arxiv-cs.CL | 2026-04-15 |
| 332 | TSQA: Integrating Text Summarization and Question Answering to Improve Information Retrieval from Documents Using Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The aim of this paper is to develop an interaction between TS and QA in three stages to enhance IR performance. |
Ahmed Sami Jaddoa; Jaber Karimpour; Pedram Salehpour; | Information | 2026-04-15 |
| 333 | Leveraging LLM-GNN Integration for Open-World Question Answering Over Knowledge Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present GLOW, a hybrid system that combines a pre-trained GNN and an LLM for open-world KGQA. |
Hussein Abdallah; Ibrahim Abdelaziz; Panos Kalnis; Essam Mansour; | arxiv-cs.CL | 2026-04-15 |
| 334 | Pyramid Graph Neural Network Knowledge Distillation with Pre-trained Language Model for Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
XUENING LI et. al. | Engineering Applications of Artificial Intelligence | 2026-04-15 |
| 335 | Understanding The Fundamental Design Decisions of Retrieval-Augmented Generation Systems IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first comprehensive study of three universal RAG deployment decisions: whether to deploy RAG, how much information to retrieve, and how to integrate retrieved knowledge effectively. |
SHENGMING ZHAO et. al. | ACM Transactions on Software Engineering and Methodology | 2026-04-15 |
| 336 | Decoding The Delta: Unifying Remote Sensing Change Detection and Understanding with Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Methodologically, we propose Delta-LLaVA, a novel MLLM framework explicitly tailored for multi-temporal remote sensing interpretation. |
XIAOHE LI et. al. | arxiv-cs.CV | 2026-04-15 |
| 337 | EgoEsportsQA: An Egocentric Video Benchmark for Perception and Reasoning in Esports Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks focus on daily activities, yet lack a rigorous testbed for evaluating fast, rule-bound reasoning in virtual scenarios. To fill this gap, we introduce EgoEsportsQA, a pioneering video question-answering (QA) benchmark for grounding perception and reasoning in expert esports knowledge. |
JIANZHE MA et. al. | arxiv-cs.CV | 2026-04-14 |
| 338 | Calibrated Confidence Estimation for Tabular Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The paper proposes Multi-Format Agreement (MFA), which exploits the lossless and deterministic serialization variation unique to structured data (Markdown, HTML, JSON, CSV) to estimate confidence at 20% lower API cost than sampling baselines. |
Lukas Voss; | arxiv-cs.CL | 2026-04-14 |
| 339 | BoxTuning: Directly Injecting The Object Box for Multimodal Model Fine-Tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent work addresses this by serializing bounding box coordinates as text tokens, but this text-coordinate paradigm suffers from a fundamental modality mismatch: object information is inherently visual, yet encoding it as text incurs a high token cost that forces aggressive temporal downsampling. We propose BoxTuning, which resolves this mismatch by injecting object spatial-temporal information directly into the visual modality. |
Zekun Qian; Ruize Han; Wei Feng; | arxiv-cs.CV | 2026-04-13 |
| 340 | QFS-Composer: Query-focused Summarization Pipeline for Less Resourced Languages Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a novel QFS framework, QFS-Composer, that integrates query decomposition, question generation (QG), question answering (QA), and abstractive summarization to improve the factual alignment of a summary with user intent. |
Vuk Đuranović; Marko Robnik Šikonja; | arxiv-cs.CL | 2026-04-12 |
| 341 | S3Mem: Structured Spatiotemporal Scene-Event Memory for Long-Horizon Interactive Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose S3MEM, a structured scene-event episodic memory framework for long-horizon interactive question answering (QA). |
ENCHENG SU et. al. | arxiv-cs.CL | 2026-04-10 |
| 342 | Boosting LLM Performance with Generative Question-answer Pairs Via Wh-transformation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Abstract This study explores methods to enhance the performance of offline Large Language Models (LLMs) using generative question-answer (QA) pairs. |
Wen-Jet Peter Wang; Chen-Sheng Luther Liu; | Concentric. Studies in Linguistics | 2026-04-10 |
| 343 | DeAtt-LMCQA: A Deberta and Attention Based Model of Legal Multi-choice Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Guibin Chen; Xudong Luo; Yanling Li; Binxia Yang; Junlin Zhu; | Artificial Intelligence and Law | 2026-04-09 |
| 344 | Rag Performance Prediction for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address the task of predicting the gain of using RAG (retrieval augmented generation) for question answering with respect to not using it. |
Or Dado; David Carmel. Oren Kurland; | arxiv-cs.CL | 2026-04-09 |
| 345 | Yale-DM-Lab at ArchEHR-QA 2026: Deterministic Grounding and Multi-Pass Evidence Alignment for EHR Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We describe the Yale-DM-Lab system for the ArchEHR-QA 2026 shared task. |
Elyas Irankhah; Samah Fodeh; | arxiv-cs.CL | 2026-04-08 |
| 346 | DTCRS: Dynamic Tree Construction for Recursive Summarization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DTCRS, a method that dynamically generates summary trees based on document structure and query semantics. |
Guanran Luo; Zhongquan Jian; Wentao Qiu; Meihong Wang; Qingqiang Wu; | arxiv-cs.CL | 2026-04-08 |
| 347 | A Systematic Study of Retrieval Pipeline Design for Retrieval-Augmented Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study presents a systematic evaluation of retrieval-augmented medical question answering using the MedQA USMLE benchmark and a structured textbook-based knowledge corpus. |
Nusrat Sultana; Abdullah Muhammad Moosa; Kazi Afzalur Rahman; Sajal Chandra Banik; | arxiv-cs.CL | 2026-04-08 |
| 348 | Jamendo-MT-QA: A Benchmark for Multi-Track Comparative Music Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We describe an LLM-assisted pipeline for generating and filtering comparative questions, and benchmark representative audio-language models using both automatic metrics and LLM-as-a-Judge evaluation. |
JUNYOUNG KOH et. al. | arxiv-cs.IR | 2026-04-08 |
| 349 | GCoT-Decoding: Unlocking Deep Reasoning Paths for Universal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recently proposed CoT-decoding enables the model to generate CoT-style reasoning paths without prompts, but it is only applicable to problems with fixed answer sets. To address this limitation, we propose a general decoding strategy GCoT-decoding that extends applicability to a broader range of question-answering tasks. |
Guanran Luo; Wentao Qiu; Zhongquan Jian; Meihong Wang; Qingqiang Wu; | arxiv-cs.CL | 2026-04-08 |
| 350 | Evaluating Repository-level Software Documentation Via Question Answering and Feature-Driven Development Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Large Language Models (LLMs) advance documentation generation from code snippets to entire repositories, existing benchmarks have two key limitations: (1) they lack a holistic, repository-level assessment, and (2) they rely on unreliable evaluation strategies, such as LLM-as-a-judge, which suffers from vague criteria and limited repository-level knowledge. To address these issues, we introduce SWD-Bench, a novel benchmark for evaluating repository-level software documentation. |
Xinchen Wang; Ruida Hu; Cuiyun Gao; Pengfei Gao; Chao Peng; | arxiv-cs.SE | 2026-04-08 |
| 351 | Database Querying Under Missing Values Governed By Missingness Mechanisms Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The MG together with the observed DB allow to build a block-independent probabilistic DB, on which basis we propose two QA techniques that jointly capture probabilistic uncertainty and statistical plausibility of the implicit imputation of MVs. |
Leopoldo Bertossi; Farouk Toumani; Maxime Buron; | arxiv-cs.DB | 2026-04-07 |
| 352 | Data-Driven Function Calling Improvements in Large Language Model for Online Financial QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a data-driven pipeline to enhance function calling in LLM for our online, deployed financial QA, comprising dataset construction, data augmentation, and model training. |
XING TANG et. al. | arxiv-cs.IR | 2026-04-06 |
| 353 | EvolveRouter: Co-Evolving Routing and Prompt for Multi-Agent Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose EvolveRouter, a trainable framework that addresses both limitations by jointly improving agent quality and collaboration structure. |
Jiatan Huang; Zheyuan Zhang; Kaiwen Shi; Yanfang Ye; Chuxu Zhang; | arxiv-cs.CL | 2026-04-06 |
| 354 | Sports-QA: A Large-Scale Video Question Answering Benchmark for Complex and Professional Sports IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this paper, we introduce the first dataset, named Sports-QA, specifically designed for the sports VideoQA task. |
HAOPENG LI et. al. | International Journal of Computer Vision | 2026-04-06 |
| 355 | PassiveQA: A Three-Action Framework for Epistemically Calibrated Question Answering Via Supervised Finetuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we study decision-aware query resolution under incomplete information, where a model must determine whether to Answer, Ask for clarification, or Abstain. |
Madhav S Baidya; | arxiv-cs.CL | 2026-04-06 |
| 356 | This Treatment Works, Right? Evaluating LLM Sensitivity to Patient Question Framing in Medical QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Ideally, LLMs should respond consistently regardless of phrasing, particularly when grounded in the same underlying evidence. We investigate this through a systematic evaluation in a controlled retrieval-augmented generation (RAG) setting for medical question answering (QA), where expert-selected documents are used rather than retrieved automatically. |
HYE SUN YUN et. al. | arxiv-cs.CL | 2026-04-06 |
| 357 | An End-to-End Framework for Building Large Language Models for Software Operations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To explore the potential of LLMs in software operations, we propose OpsLLM, a domain-specific LLM that supports both knowledge-based question answering (QA) and root cause analysis (RCA). |
JINGKAI HE et. al. | arxiv-cs.LG | 2026-04-05 |
| 358 | GROUNDEDKG-RAG: Grounded Knowledge Graph Index for Long-document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we focus on RAG systems for long-document question answering. |
Tianyi Zhang; Andreas Marfurt; | arxiv-cs.CL | 2026-04-05 |
| 359 | PRAISE: Prefix-Based Rollout Reuse in Agentic Search Training Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Prefix-based Rollout reuse for Agentic search with Intermediate Step rEwards (PRAISE), a framework for improving both data efficiency and credit assignment in agentic search training. |
ERHAN ZHANG et. al. | arxiv-cs.AI | 2026-04-04 |
| 360 | Document-Level Numerical Reasoning Across Single and Multiple Tables in Financial Reports Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FinLongDocAgent, a Multi-Agent Multi-Round Retrieval-Augmented Generation (RAG) approach that iteratively retrieves evidence, performs intermediate calculations, and verifies results across rounds. |
Yi-Cheng Wang; Wei-An Wang; Chu-Song Chen; | arxiv-cs.CL | 2026-04-04 |
| 361 | PaveBench: A Versatile Benchmark for Pavement Distress Perception and Interactive Vision-Language Analysis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: They also do not connect perception with vision-language analysis. To address these limitations, we introduce PaveBench, a large-scale benchmark for pavement distress perception and interactive vision-language analysis on real-world highway inspection images. |
DEXIANG LI et. al. | arxiv-cs.CV | 2026-04-03 |
| 362 | Injecting Structured Biomedical Knowledge Into Language Models: Continual Pretraining Vs. GraphRAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We then derive a ~100-million-token textual corpus from this graph to continually pretrain two models: BERTUMLS (from BERT) and BioBERTUMLS (from BioBERT). We evaluate these models on six BLURB (Biomedical Language Understanding and Reasoning Benchmark) datasets spanning five task types and evaluate GraphRAG on the two QA (Question Answering) datasets (PubMedQA, BioASQ). |
Jaafer Klila; Sondes Bannour Souihi; Rahma Boujelben; Nasredine Semmar; Lamia Hadrich Belguith; | arxiv-cs.CL | 2026-04-03 |
| 363 | V2X-QA: A Comprehensive Reasoning Dataset and Benchmark for Multimodal Large Language Models in Autonomous Driving Across Ego, Infrastructure, and Cooperative Views Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce V2X-QA, a real-world dataset and benchmark for evaluating MLLMs across vehicle-side, infrastructure-side, and cooperative viewpoints. |
JUNWEI YOU et. al. | arxiv-cs.RO | 2026-04-03 |
| 364 | Ego-Grounding for Personalized Question-Answering in Egocentric Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding – the ability to understand the camera-wearer in egocentric videos. |
Junbin Xiao; Shenglang Zhang; Pengxiang Zhu; Angela Yao; | arxiv-cs.CV | 2026-04-02 |
| 365 | Do We Need Bigger Models for Science? Task-Aware Retrieval with Small Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we ask a simple question: do we need bigger models for scientific applications? |
Florian Kelber; Matthias Jobst; Yuni Susanti; Michael Färber; | arxiv-cs.IR | 2026-04-02 |
| 366 | VideoZeroBench: Probing The Limits of Video MLLMs with Spatio-Temporal Evidence Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current evaluations suffer from two critical limitations: (1) inflated scores can mask deficiencies in fine-grained visual understanding and reasoning, and (2) answer correctness is often measured without verifying whether models identify the precise spatio-temporal evidence supporting their predictions. To address this, we present VideoZeroBench, a hierarchical benchmark designed for challenging long-video question answering that rigorously verifies spatio-temporal evidence. |
JIAHAO MENG et. al. | arxiv-cs.CV | 2026-04-01 |
| 367 | Scalable Conversion of Stack Exchange Network-based Question Answering Sites’ Data Dump to SQL Server Database – A Case of Stack Overflow Data Dump Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Arjumand Fatima; O. Maqbool; | Array | 2026-04-01 |
| 368 | CP-RAG: Mitigating Distracting Content in Retrieval-Augmented Generation for Industrial Knowledge Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: With the increasing adoption of IIoT in industrial production, producing massive heterogeneous data, retrieval-augmented generation (RAG) has become a promising approach for … |
Cong Wang; Shuowen Chai; Tianyi Xu; Muhammad Adil; Tie Qiu; | IEEE Internet of Things Journal | 2026-04-01 |
| 369 | Can Large Language Models Self-Correct in Medical Question Answering? An Exploratory Study Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we conduct an exploratory analysis of self-reflective reasoning for medical multiple-choice question answering: using GPT-4o and GPT-4o-mini, we compare standard CoT prompting with an iterative self-reflection loop and track how predictions evolve across reflection steps on three widely used medical QA benchmarks (MedQA, HeadQA, and PubMedQA). |
Zaifu Zhan; Mengyuan Cui; Rui Zhang; | arxiv-cs.CL | 2026-03-31 |
| 370 | Long-Document QA with Chain-of-Structured-Thought and Fine-Tuned SLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a two-pillar framework, LiteCoST, to achieve both high accuracy and low latency with small language models (SLMs). |
ZHUOWEN LIANG et. al. | arxiv-cs.CL | 2026-03-31 |
| 371 | SeGPruner: Semantic-Geometric Visual Token Pruner for 3D Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose SeGPruner, a semantic-aware and geometry-guided token reduction framework for efficient 3D QA with multi-view images. |
WENLI LI et. al. | arxiv-cs.CV | 2026-03-31 |
| 372 | SurgTEMP: Temporal-Aware Surgical Video Question Answering with Text-guided Visual Memory for Laparoscopic Cholecystectomy Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Surgical video question answering is further challenged by low visual contrast, its highly knowledge-driven nature, diverse analytical needs spanning scattered temporal windows, and the hierarchy from basic perception to high-level intraoperative assessment. To address these challenges, we propose SurgTEMP, a multimodal LLM framework featuring (i) a query-guided token selection module that builds hierarchical visual memory (spatial and temporal memory banks) and (ii) a Surgical Competency Progression (SCP) training scheme. |
SHI LI et. al. | arxiv-cs.CV | 2026-03-31 |
| 373 | A Chinese Financial Event Knowledge Graph-based Retrieval-augmented Generation Framework for Financial Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Haitao Cheng; Ke Wang; Qi Wang; Tao Liu; Kai Sheng; | Engineering Applications of Artificial Intelligence | 2026-03-31 |
| 374 | PAR$^2$-RAG: Planned Active Retrieval and Reasoning for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textbf{Planned Active Retrieval and Reasoning RAG (PAR$^2$-RAG)}, a two-stage framework that separates \emph{coverage} from \emph{commitment}. |
XINGYU LI et. al. | arxiv-cs.AI | 2026-03-30 |
| 375 | When Choices Become Priors: Contrastive Decoding for Scientific Figure Multiple-Choice QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose SCICON, a training-free decoding method that scores each candidate by subtracting a text-only option score from its image-conditioned counterpart. |
Taeyun Roh; Eun-yeong Jo; Wonjune Jang; Jaewoo Kang; | arxiv-cs.AI | 2026-03-30 |
| 376 | CDH-Bench: A Commonsense-Driven Hallucination Benchmark for Evaluating Visual Fidelity in Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To evaluate it, we introduce \textbf{CDH-Bench}, a benchmark designed to create explicit \textbf{visual evidence–commonsense conflicts}. |
Kesheng Chen; Yamin Hu; Qi Zhou; Zhenqian Zhu; Wenjian Luo; | arxiv-cs.CV | 2026-03-29 |
| 377 | Beyond Code Snippets: Benchmarking LLMs on Repository-Level Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce StackRepoQA, the first multi-project, repository-level question answering dataset constructed from 1,318 real developer questions and accepted answers across 134 open-source Java projects. |
Yoseph Berhanu Alebachew; Hunter Leary; Swanand Vaishampayan; Chris Brown; | arxiv-cs.SE | 2026-03-27 |
| 378 | Do Hallucination Neurons Generalize? Evidence from Cross-Domain Transfer in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We ask a natural follow-up question: do H-neurons generalize across knowledge domains? Using a systematic cross-domain transfer protocol across 6 domains (general QA, legal, financial, science, moral reasoning, and code vulnerability) and 5 open-weight models (3B to 8B parameters), we find they do not. |
Snehit Vaddi; Pujith Vaddi; | arxiv-cs.CL | 2026-03-26 |
| 379 | QU-NLP at ArchEHR-QA 2026: Two-Stage QLoRA Fine-Tuning of Qwen3-4B for Patient-Oriented Clinical Question Answering and Evidence Sentence Alignment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a unified system addressing both Subtask 3 (answer generation) and Subtask 4 (evidence sentence alignment) of the ArchEHR-QA Shared Task. |
Mohammad AL-Smadi; | arxiv-cs.CL | 2026-03-26 |
| 380 | LLaVA-LE: Large Language-and-Vision Assistant for Lunar Exploration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce LLaVA-LE (Large Language-and-Vision Assistant for Lunar Exploration), a vision-language model specialized for lunar surface and subsurface characterization. |
Gokce Inal; Pouyan Navard; Alper Yilmaz; | arxiv-cs.CV | 2026-03-25 |
| 381 | Retrieval Improvements Do Not Guarantee Better Answers: A Study of RAG for AI Policy QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Retrieval-augmented generation (RAG) systems are increasingly used to analyze complex policy documents, but achieving sufficient reliability for expert usage remains challenging in domains characterized by dense legal language and evolving, overlapping regulatory frameworks. |
Saahil Mathur; Ryan David Rittner; Vedant Ajit Thakur; Daniel Stuart Schiff; Tunazzina Islam; | arxiv-cs.CL | 2026-03-25 |
| 382 | See, Remember, Explore: A Benchmark and Baselines for Streaming Spatial Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Spatial understanding is fundamental for embodied agents, yet most spatial VLMs and benchmarks remain offline-evaluating post-hoc QA over pre-recorded inputs and overlooking two crucial deployment-critical requirements: long-horizon streaming inference and active perception when the current view is insufficient. To address this gap, we introduce S3-Bench, a benchmark suite for streaming spatial question answering with active exploration, where queries are temporally grounded to specific timestamps and must be answered using only observations available up to that moment. |
Yuxi Wei; Wei Huang; Qirui Chen; Lu Hou; Xiaojuan Qi; | arxiv-cs.CV | 2026-03-24 |
| 383 | Mixture of Demonstrations for Textual Graph Understanding and Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose MixDemo, a novel GraphRAG framework enhanced with a Mixture-of-Experts (MoE) mechanism for selecting the most informative demonstrations under diverse question contexts. |
Yukun Wu; Lihui Liu; | arxiv-cs.IR | 2026-03-23 |
| 384 | Efficient Fine-Tuning Methods for Portuguese Question Answering: A Comparative Study of PEFT on BERTimbau and Exploratory Evaluation of Generative LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work presents a systematic evaluation of Parameter-Efficient Fine-Tuning (PEFT) and quantization techniques applied to BERTimbau for Question Answering on SQuAD-BR, the Brazilian Portuguese translation of SQuAD v1. |
Mariela M. Nina; Caio Veloso Costa; Lilian Berton; Didier A. Vega-Oliveros; | arxiv-cs.CL | 2026-03-22 |
| 385 | KLDrive: Fine-Grained 3D Scene Reasoning for Autonomous Driving Based on Knowledge Graph Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present KLDrive, the first knowledge-graph-augmented LLM reasoning framework for fine-grained question answering in autonomous driving. |
YE TIAN et. al. | arxiv-cs.AI | 2026-03-21 |
| 386 | PAVE: Premise-Aware Validation and Editing for Retrieval-Augmented LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present PAVE (Premise-Grounded Answer Validation and Editing), an inference-time validation layer for evidence-grounded question answering. |
Tianyi Huang; Caden Yang; Emily Yin; Eric Wang; Michael Zhang; | arxiv-cs.CL | 2026-03-21 |
| 387 | HORNet: Task-Guided Frame Selection for Video Question Answering with Vision-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce \textbf{HORNet}, a lightweight frame selection policy trained with Group Relative Policy Optimization (GRPO) to learn which frames a frozen VLM needs to answer questions correctly. |
Xiangyu Bai; Bishoy Galoaa; Sarah Ostadabbas; | arxiv-cs.CV | 2026-03-19 |
| 388 | Bypassing Document Ingestion: An MCP Approach to Financial Q&A Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study whether, and under which conditions, Model Context Protocol (MCP) offers a more reliable alternative to standard retrieval-augmented generation (RAG) by allowing large language models (LLMs) to interact directly with data rather than relying on document ingestion and chunk retrieval. |
Sasan Mansouri; Edoardo Pilla; Mark Wahrenburg; Fabian Woebbeking; | arxiv-cs.IR | 2026-03-19 |
| 389 | FDARxBench: Benchmarking Regulatory and Clinical Reasoning on FDA Generic Drug Assessment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce an expert curated, real-world benchmark for evaluating document-grounded question-answering (QA) motivated by generic drug assessment, using the U.S. Food and Drug Administration (FDA) drug label documents. |
BETTY XIONG et. al. | arxiv-cs.CL | 2026-03-19 |
| 390 | DaPT: A Dual-Path Framework for Multilingual Multi-hop Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Retrieval-augmented generation (RAG) systems have made significant progress in solving complex multi-hop question answering (QA) tasks in the English scenario. However, RAG … |
YILIN WANG et. al. | arxiv-cs.CL | 2026-03-19 |
| 391 | Halo: Domain-Aware Query Optimization for Long-Context Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Halo, a long-context QA framework that automatically extracts domain knowledge from user prompts and applies it as executable operators across a multi-stage query execution pipeline. |
Pramod Chunduri; Francisco Romero; Ali Payani; Kexin Rong; Joy Arulraj; | arxiv-cs.DB | 2026-03-18 |
| 392 | Multi-layer Biaffine Model and Transformer Question Answering System for Tamil Language Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
S Jeevit Davidson; M Murali; | Engineering Applications of Artificial Intelligence | 2026-03-18 |
| 393 | How Often Do Answers Change? Estimating Recency Requirements in Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks either periodically refresh answers or rely on fixed templates, but they do not reflect on how frequently answers change or whether a question inherently requires up-to-date information. To address this gap, we introduce a recency-stationarity taxonomy that categorizes questions by how often their answers change and whether this change frequency is time-invariant or context-dependent. |
Bhawna Piryani; Zehra Mert; Adam Jatowt; | arxiv-cs.CL | 2026-03-17 |
| 394 | IndexRAG: Bridging Facts for Cross-Document Reasoning at Index Time Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present IndexRAG, a novel approach that shifts cross-document reasoning from online inference to offline indexing. |
Zhenghua Bao; Yi Shi; | arxiv-cs.CL | 2026-03-17 |
| 395 | InViC: Intent-aware Visual Cues for Medical Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a lightweight plug-in framework, termed Intent-aware Visual Cues (InViC), to explicitly enhance image-based answer generation in medical VQA. |
ZHISONG WANG et. al. | arxiv-cs.CV | 2026-03-17 |
| 396 | Attention-guided Evidence Grounding for Spoken Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Attention-guided Evidence Grounding (AEG), a novel end-to-end framework that leverages the internal cross-modal attention of Speech Large Language Models (SpeechLLMs) to explicitly locate and ground key evidence in the model’s latent space. |
KE YANG et. al. | arxiv-cs.CL | 2026-03-17 |
| 397 | CounterRefine: Answer-Conditioned Counterevidence Retrieval for Inference-Time Knowledge Repair in Factual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present CounterRefine, a lightweight inference-time repair layer for retrieval-grounded question answering. |
Tianyi Huang; Ying Kai Deng; | arxiv-cs.CL | 2026-03-16 |
| 398 | Benchmarking Real-Time Question Answering Via Executable Code Workflows Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks are predominantly static and therefore fail to capture the temporal dynamics of information and the continuously evolving nature of real-world knowledge. To address this limitation, we propose RT-QA, a dynamic evaluation framework that leverages executable code workflows to retrieve up-to-date answers at evaluation time. |
WENJIE ZHOU et. al. | arxiv-cs.IR | 2026-03-16 |
| 399 | A Comprehensive Survey on Table Question Answering: Datasets, Methods and Future Directions Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Weiqiang Xu; Yang Liu; Lingfeng Lu; Huakang Li; Guozi Sun; | Engineering Applications of Artificial Intelligence | 2026-03-16 |
| 400 | Learning Question-Aware Keyframe Selection with Synthetic Supervision for Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a question-aware keyframe selection framework with two components: pseudo keyframe labels derived from LMMs that provide informative supervision and a coverage regularization that promotes diverse, complementary evidence across time. |
Minchan Kwon; Hyounguk Shon; Junmo Kim; | arxiv-cs.CV | 2026-03-16 |
| 401 | Information Asymmetry Across Language Varieties: A Case Study on Cantonese-Mandarin and Bavarian-German QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We manually construct a novel challenge question-answering (QA) dataset that captures knowledge conveyed on a local Wikipedia page, which is absent from their higher-resource counterparts-covering Mandarin Chinese vs. Cantonese and German vs. Bavarian. |
Renhao Pei; Siyao Peng; Verena Blaschke; Robert Litschko; Barbara Plank; | arxiv-cs.CL | 2026-03-15 |
| 402 | LongVidSearch: An Agentic Benchmark for Multi-hop Evidence Retrieval Planning in Long Videos Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing long-video benchmarks are largely static: they rarely enforce strict multi-hop retrieval and typically lack a standardized evidence-access interface, making it difficult to separate failures in retrieval planning from those in answer generation. To address this gap, we introduce LongVidSearch, a benchmark for evaluating agentic multi-hop evidence retrieval planning in long videos under standardized access constraints. |
Rongyi Yu; Chenyuan Duan; Wentao Zhang; | arxiv-cs.CV | 2026-03-15 |
| 403 | English Visual Question Answering: Building A Culturally Relevant Dataset from Image Captions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes an efficient and scalable VQA system using a pretrained CLIP (Contrastive Language–Image Pretraining) ViT-B/32 model for open-ended English language queries. |
Dr. Md Sirajul Huque; B. Dinesh kumar; Bhoomika Mandadi; C. Anil kumar; | International Journal of Scientific Research in Engineering … | 2026-03-15 |
| 404 | Automatic Inter-document Multi-hop Scientific QA Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present AIM-SciQA, an automated framework for generating multi-document, multi-hop scientific QA datasets. |
Seungmin Lee; Dongha Kim; Yuni Jeon; Junyoung Koh; Min Song; | arxiv-cs.CL | 2026-03-15 |
| 405 | The Reasoning Bottleneck in Graph-RAG: Structured Prompting and Context Compression for Multi-Hop QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose two augmentations: (i) SPARQL chain-of-thought prompting, which decomposes questions into triple-pattern queries aligned with the entity-relationship context, and (ii) graph-walk compression, which compresses the context by ~60% via knowledge-graph traversal with no LLM calls. |
Yasaman Zarinkia; Venkatesh Srinivasan; Alex Thomo; | arxiv-cs.IR | 2026-03-14 |
| 406 | Empowering LLMs with Symbolic Representation and Reasoning Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Large language models (LLMs) have achieved remarkable success in natural language processing tasks but still struggle with complex causal and logical reasoning. Previous … |
Fengxiang Cheng; | AAAI Conference on Artificial Intelligence | 2026-03-14 |
| 407 | Fine-Tuning Sample Order Matters in Propositional Logical Question-Answering (Student Abstract) Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Large language models (LLMs) have achieved impressive progress in natural language processing tasks but still struggle with complex logical reasoning. We observe that in … |
Fengxiang Cheng; Chuan-Cang Zhou; Fenrong Liu; R. Rooij; | AAAI Conference on Artificial Intelligence | 2026-03-14 |
| 408 | Sebis at ArchEHR-QA 2026: How Much Can You Do Locally? Evaluating Grounded EHR QA on A Single Notebook Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate how far grounded EHR question answering can be pushed when restricted to a single notebook. |
Ibrahim Ebrar Yurt; Fabian Karl; Tejaswi Choppa; Florian Matthes; | arxiv-cs.CL | 2026-03-14 |
| 409 | Consistency-Guided Decoding with Proof-Driven Disambiguation for Three-Way Logical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present CGD-PD, a lightweight test-time layer that (a) queries a single 3-way classifier on both $H$ and a mechanically negated form of $H$, (b) projects the pair onto a negation-consistent decision when possible, and (c) invokes a proof-driven disambiguation step that uses targeted binary entailment probes to selectively resolve $Unknown$ outcomes, requiring only an average of 4-5 model calls. |
Tianyi Huang; Ming Hou; Jiaheng Su; Yutong Zhang; Ziling Zhang; | arxiv-cs.CL | 2026-03-12 |
| 410 | MDER-DR: Multi-Hop Question Answering with Entity-Centric Summaries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a domain-agnostic, KG-based QA framework that covers both the indexing and retrieval/inference phases. |
Riccardo Campi; Nicolò Oreste Pinciroli Vago; Mathyas Giudici; Marco Brambilla; Piero Fraternali; | arxiv-cs.CL | 2026-03-11 |
| 411 | FinReflectKG — HalluBench: GraphRAG Hallucination Benchmark for Financial Question Answering Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce FinBench-QA-Hallucination, a benchmark for evaluating hallucination detection methods in KG-augmented financial QA over SEC 10-K filings. |
Mahesh Kumar; Bhaskarjit Sarmah; Stefano Pasquali; | arxiv-cs.CL | 2026-03-11 |
| 412 | ThReadMed-QA: A Multi-Turn Medical Dialogue Benchmark from Real Patient Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ThReadMed-QA, a benchmark of 2,437 fully-answered patient-physician conversation threads extracted from r/AskDocs, comprising 8,204 question-answer pairs across up to 9 turns. |
Monica Munnangi; Saiph Savage; | arxiv-cs.CL | 2026-03-11 |
| 413 | TIPS: Turn-Level Information-Potential Reward Shaping for Search-Augmented LLMs IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The optimization is often unstable due to sparse rewards and difficult credit assignments across reasoning and tool calls. To address this, we introduce Turn-Level Information Potential Reward Shaping (TIPS), a simple framework that assigns dense, turn-level rewards to each reasoning + tool-call segment based on the increased likelihood of the correct answer under a teacher model. |
YUTAO XIE et. al. | arxiv-cs.CL | 2026-03-11 |
| 414 | LLM As A Meta-Judge: Synthetic Data for NLP Evaluation Metric Validation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose \textit{LLM as a Meta-Judge}, a scalable framework that utilizes LLMs to generate synthetic evaluation datasets via controlled semantic degradation of real data, replacing human judgment. |
Lukáš Eigler; Jindřich Libovický; David Hurych; | arxiv-cs.CL | 2026-03-10 |
| 415 | A GraphRAG-Based Question-Answering System for Explainable and Advanced Reasoning Over Air Quality Insights Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The present work introduces an integrated GraphRAG-based Question Answering (QA) system that couples a domain-specific knowledge graph encoding fundamental IAQ concepts and relationships with a RAG-based natural language interface, thereby enabling explainable, context-aware, and advanced analytical reasoning over IAQ data. |
Christos Mountzouris; Grigorios Protopsaltis; John Gialelis; | Air | 2026-03-10 |
| 416 | AutoViVQA: A Large-Scale Automatically Constructed Dataset for Vietnamese Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore Vietnamese Visual Question Answering using transformer-based architectures, leveraging both textual and visual pre-training while systematically comparing automatic evaluation metrics under multilingual settings. |
NGUYEN ANH TUONG et. al. | arxiv-cs.CV | 2026-03-10 |
| 417 | Emotion Is Not Just A Label: Latent Emotional Factors in LLM Processing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We analyze how emotional tone systematically alters attention geometry in transformer models, showing that metrics such as locality, center-of-mass distance, and entropy vary across emotions and correlate with downstream question-answering performance. To facilitate controlled study of these effects, we introduce Affect-Uniform ReAding QA (AURA-QA), a question-answering dataset with emotionally balanced, human-authored context passages. |
Benjamin Reichman; Adar Avasian; Samuel Webster; Larry Heck; | arxiv-cs.CL | 2026-03-10 |
| 418 | SPD-RAG: Sub-Agent Per Document Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SPD-RAG, a hierarchical multi-agent framework for exhaustive cross-document question answering that decomposes the problem along the document axis. |
Yagiz Can Akay; Muhammed Yusuf Kartal; Esra Alparslan; Faruk Ortakoyluoglu; Arda Akpinar; | arxiv-cs.CL | 2026-03-09 |
| 419 | Gradually Excavating External Knowledge for Implicit Complex Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, this work proposes a gradual knowledge excavation framework for open-domain complex question answering, where LLMs iteratively and actively acquire external information, and then reason based on acquired historical knowledge. |
CHANG LIU et. al. | arxiv-cs.CL | 2026-03-09 |
| 420 | MEGC2026: Micro-Expression Grand Challenge on Visual Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial … |
XINQI FAN et. al. | arxiv-cs.CV | 2026-03-09 |
| 421 | Decomposition-Driven Multi-Table Retrieval and Reasoning for Numerical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we study the problem of numerical multi-table question answering (MTQA) over large-scale table collections (e.g., online data repositories). |
FENG LUO et. al. | arxiv-cs.DB | 2026-03-09 |
| 422 | BRIDGE: Benchmark for Multi-hop Reasoning In Long Multimodal Documents with Grounded Evidence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce BRIDGE, a benchmark for multi-hop reasoning over long scientific papers that require integrating evidence across text, tables, and figures. |
Biao Xiang; Soyeon Caren Han; Yihao Ding; | arxiv-cs.CL | 2026-03-08 |
| 423 | AQuA: Toward Strategic Response Generation for Ambiguous Visual Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Ambiguous Visual Question Answering (AQuA), a fine-grained dataset that classifies ambiguous VQA instances into four levels according to the nature and degree of ambiguity, along with the optimal response strategy for each case. |
Jihyoung Jang; Hyounghun Kim; | arxiv-cs.CV | 2026-03-07 |
| 424 | MAviS: A Multimodal Conversational Assistant For Avian Species Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing multimodal large language models face challenges when it comes to specialized topics like avian species, making it harder to provide accurate and contextually relevant information in these areas. To address this limitation, we introduce the MAviS-Dataset, a large-scale multimodal avian species dataset that integrates image, audio, and text modalities for over 1,000 bird species, comprising both pretraining and instruction-tuning subsets enriched with structured question-answer pairs. |
YEVHENIIA KRYKLYVETS et. al. | arxiv-cs.CV | 2026-03-07 |
| 425 | RAMoEA-QA: Hierarchical Specialization for Robust Respiratory Audio Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: They are also only validated in limited settings, leaving it unclear how reliably they handle the shifts encountered in real-world settings. To address these limitations, we introduce RAMoEA-QA, a hierarchically routed generative model for respiratory audio question answering that unifies multiple question types and supports both discrete and continuous targets within a single multimodal system. |
Gaia A. Bertolino; Yuwei Zhang; Tong Xia; Domenico Talia; Cecilia Mascolo; | arxiv-cs.SD | 2026-03-06 |
| 426 | Helpful or Harmful? Re-Evaluating Frugality in Retrieval-Augmented Generation for Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work introduces a frugality-based evaluation framework that jointly assesses accuracy improvements and computational cost to determine when retrieval-augmented generation is beneficial in medical question answering, rather than evaluating retrieval effectiveness through accuracy alone. |
Richard Coric; Ebenezer F. Oloyede; Heriberto Cuayáhuitl; | Machine Learning and Knowledge Extraction | 2026-03-06 |
| 427 | Confidence Before Answering: A Paradigm Shift for Efficient LLM Uncertainty Estimation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CoCA(Co-optimized Confidence and Answers), a GRPO reinforcement learning framework that jointly optimizes confidence calibration and answer accuracy via segmented credit assignment. |
CHANGCHENG LI et. al. | arxiv-cs.CL | 2026-03-05 |
| 428 | NCTB-QA: A Large-Scale Bangla Educational Question Answering Dataset and Benchmarking Performance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These systems tend to produce unreliable responses when correct answers are absent from context. To solve this problem, we introduce NCTB-QA, a large-scale Bangla question answering dataset comprising 87,805 question-answer pairs extracted from 50 textbooks published by Bangladesh’s National Curriculum and Textbook Board. |
Abrar Eyasir; Tahsin Ahmed; Muhammad Ibrahim; | arxiv-cs.CL | 2026-03-05 |
| 429 | Who Judges The Judge? Evaluating LLM-as-a-Judge for French Medical Open-ended QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluate whether large language models (LLMs) can act as judges of semantic equivalence in French medical OEQA, comparing closed-access, general-purpose, and biomedical domain-adapted models. |
Ikram Belmadani; Oumaima El Khettari; Pacôme Constant dit Beaufils; Richard Dufour; Benoit Favre; | arxiv-cs.CL | 2026-03-04 |
| 430 | RAG-X: Systematic Diagnosis of Retrieval-Augmented Generation for Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These approaches fail to diagnose whether an error stems from faulty retrieval or flawed generation, limiting developers from performing targeted improvement. To address this gap, we propose RAG-X, a diagnostic framework that evaluates the retriever and generator independently across a triad of QA tasks: information extraction, short-answer generation, and multiple-choice question (MCQ) answering. |
Aswini Sivakumar; Vijayan Sugumaran; Yao Qiang; | arxiv-cs.CL | 2026-03-03 |
| 431 | KGLMQA: Enhancing Medical Visual Question Answering with Knowledge Graphs and LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing models frequently encounter challenges such as restricted multimodal interaction, insufficient guidance from external medical knowledge, and a lack of rigorous diagnostic logic in their responses. To address these issues, we propose KGLMQA, a novel framework that integrates knowledge graphs with Large Language Models (LLMs). |
Wenhu Wang; Huina Liu; Changfa Wei; | PeerJ Computer Science | 2026-03-03 |
| 432 | Let The Agent Search: Autonomous Exploration Beats Rigid Workflows in Temporal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We show that simply granting an off-the-shelf LLM autonomy, that is, letting it decide what to do next, already yields substantial gains even in a strict zero-shot setting. Building on this insight, we propose AT2QA, an autonomous, training-free agent for temporal question answering that iteratively interacts with the temporal knowledge graph via a general search tool for dynamic retrieval. |
Xufei Lv; Jiahui Yang; Yifu Gao; Linbo Qiao; Houde Liu; | arxiv-cs.CL | 2026-03-02 |
| 433 | Co-MedGraphRAG: A Collaborative Large–Small Model Medical Question-Answering Framework Enhanced By Knowledge Graph Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work serves as a reference for researchers and developers designing medical question-answering frameworks and exploring decision-support applications. |
Sizhe Chen; Tao Chen; | Information | 2026-03-02 |
| 434 | A Self-Reflection Mechanism for Reducing Hallucination in Vietnamese Legal Question Answering Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a self-reflection mechanism that adds an iterative generate–evaluate–refine loop to a Graph-RAG pipeline for Vietnamese labor-law questions. |
THI VUONG PHAM et. al. | Scientific Journal of Computer Science | 2026-03-02 |
| 435 | Evaluating Prompting Strategies for Chart Question Answering with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a systematic evaluation of four widely used prompting paradigms (Zero-Shot, Few-Shot, Zero-Shot Chain-of-Thought, and Few-Shot Chain-of-Thought) across GPT-3.5, GPT-4, and GPT-4o on the ChartQA dataset. |
Ruthuparna Naikar; Ying Zhu; | arxiv-cs.CL | 2026-03-02 |
| 436 | DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Deep-research agents are capable of executing multi-step web exploration, targeted retrieval, and sophisticated question answering. Despite their powerful capabilities, … |
TONGZHOU WU et. al. | arxiv-cs.AI | 2026-03-01 |
| 437 | Spatial–Temporal Clue Reasoning Chain for Long Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
HAIBO GONG et. al. | IEEE Transactions on Circuits and Systems for Video … | 2026-03-01 |
| 438 | Prompt Sensitivity and Answer Consistency of Small Open-Source Large Language Models on Clinical Question Answering: Implications for Low-Resource Healthcare Deployment Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluated five open-source models (Gemma 2 2B, Phi-3 Mini 3.8B, Llama 3.2 3B, Mistral 7B, and Meditron-7B domain-pretrained without instruction tuning) across three clinical QA datasets (MedQA, MedMCQA, PubMedQA) using five prompt styles (original, formal, simplified, roleplay, direct). |
Shravani Hariprasad; | arxiv-cs.CL | 2026-02-28 |
| 439 | FHIRPath-QA: Executable Question Answering Over FHIR Electronic Health Records Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce FHIRPath-QA, the first open dataset and benchmark for patient-specific QA that includes open-standard FHIRPath queries over real-world clinical data. |
Michael Frew; Nishit Bheda; Bryan Tripp; | arxiv-cs.CL | 2026-02-26 |
| 440 | SPARTA: Scalable and Principled Benchmark of Tree-Structured Multi-hop QA Over Text and Tables Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SPARTA, an end-to-end construction framework that automatically generates large-scale Table-Text QA benchmarks with lightweight human validation, requiring only one quarter of the annotation time of HybridQA. |
Sungho Park; Jueun Kim; Wook-Shin Han; | arxiv-cs.CL | 2026-02-26 |
| 441 | A Novel Multi-modal Attentional Collaborative Learning Framework with Semantic Enhancement for Audio–visual Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
JIE YANG et. al. | Engineering Applications of Artificial Intelligence | 2026-02-25 |
| 442 | LiCQA : A Lightweight Complex Question Answering System Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present LiCQA, an unsupervised question answer- ing model that works primarily on the basis of corpus evidence. |
Sourav Saha; Dwaipayan Roy; Mandar Mitra; | arxiv-cs.CL | 2026-02-25 |
| 443 | A Dataset for Addressing Patient’s Information Needs Related to Clinical Course of Hospitalization IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, robust datasets to assess the factuality and relevance of AI-generated responses are lacking and, to our knowledge, none capture patient information needs in the context of their EHRs. To address this gap, we introduce ArchEHR-QA, an expert-annotated dataset of 134 cases from intensive care unit and emergency department settings. |
Sarvesh Soni; Dina Demner-Fushman; | Scientific Data | 2026-02-25 |
| 444 | Exploring Multimodal LMMs for Online Episodic Memory Question Answering on The Edge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate the feasibility of using Multimodal Large Language Models (MLLMs) for real-time online episodic memory question answering. |
Giuseppe Lando; Rosario Forte; Antonino Furnari; | arxiv-cs.CV | 2026-02-25 |
| 445 | Spelling Correction in Healthcare Query-Answer Systems: Methods, Retrieval Impact, and Empirical Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Healthcare question-answering (QA) systems face a persistent challenge: users submit queries with spelling errors at rates substantially higher than those found in the professional documents they search. |
Saurabh K Singh; | arxiv-cs.CL | 2026-02-24 |
| 446 | Controllable Evidence Selection in Retrieval-Augmented Question Answering Via Deterministic Utility Gating Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents a deterministic evidence selection framework for retrieval-augmented question answering. |
Victor P. Unda; | arxiv-cs.CL | 2026-02-23 |
| 447 | Retrieval-Augmented Generation for Multi-Hop Question Answering Based on Structured Planning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As the number of retrieval iterations increases, the generated queries can gradually drift from the correct reasoning path, and irrelevant or noisy information may accumulate, ultimately reducing reasoning accuracy. To address these challenges, we propose a novel retrieval-augmented generation method for multi-hop question answering based on structured planning. |
Yujiao Huang; Ling Yang; Xu-Hua Yang; Xinli Xu; | ACM Transactions on Knowledge Discovery from Data | 2026-02-23 |
| 448 | To Reason or Not To: Selective Chain-of-Thought in Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Methods: We propose Selective Chain-of-Thought (Selective CoT), an inference-time strategy that first predicts whether a question requires reasoning and generates a rationale only when needed. |
ZAIFU ZHAN et. al. | arxiv-cs.CL | 2026-02-23 |
| 449 | Temporal-Aware Heterogeneous Graph Reasoning with Multi-View Fusion for Temporal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel framework with temporal-aware question encoding, multi-hop graph reasoning, and multi-view heterogeneous information fusion. |
WUZHENGHONG WEN et. al. | arxiv-cs.CL | 2026-02-23 |
| 450 | Efficient Multimodal Learning Using BERT and Vision Transformers for Visual Question Answering on Peripheral Blood Cells Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Faheem Shehzad; Ciro Mennella; Massimo Esposito; Aniello Minutolo; | Discover Artificial Intelligence | 2026-02-22 |
| 451 | Learning to Reason for Multi-Step Retrieval of Personal Context in Personalized Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose PR2 (Personalized Retrieval-Augmented Reasoning), a reinforcement learning framework that integrates reasoning and retrieval from personal context for personalization. |
Maryam Amirizaniani; Alireza Salemi; Hamed Zamani; | arxiv-cs.CL | 2026-02-22 |
| 452 | A Large-scale Benchmark for Evaluating Large Language Models on Medical Question Answering in Romanian Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
ANA-CRISTINA ROGOZ et. al. | npj Digital Medicine | 2026-02-21 |
| 453 | Decomposing Retrieval Failures in RAG for Long-Document Financial Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Across methods, gains in document discovery tend to translate into stronger page recall, yet oracle performance still suggests headroom for page and chunk level retrieval. To target this gap, we introduce a domain fine-tuned page scorer that treats pages as an intermediate retrieval unit between documents and chunks. |
Amine Kobeissi; Philippe Langlais; | arxiv-cs.CL | 2026-02-19 |
| 454 | InsQABench: Benchmarking Chinese Insurance Domain Question Answering with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save |
LAI WEI et. al. | Information Processing & Management | 2026-02-19 |
| 455 | Condition-Gated Reasoning for Context-Dependent Biomedical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks do not evaluate such conditional reasoning, and retrieval-augmented or graph-based methods lack explicit mechanisms to ensure that retrieved knowledge is applicable to given context. To address this gap, we propose CondMedQA, the first benchmark for conditional biomedical QA, consisting of multi-hop questions whose answers vary with patient conditions. |
JASH RAJESH PAREKH et. al. | arxiv-cs.CL | 2026-02-19 |
| 456 | Robustness and Reasoning Fidelity of Large Language Models in Long-Context Code Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We conduct a systematic study of long-context code question answering using controlled ablations that test sensitivity to answer format, distractors, and context scale. |
Kishan Maharaj; Nandakishore Menon; Ashita Saxena; Srikanth Tamilselvam; | arxiv-cs.SE | 2026-02-19 |
| 457 | Evaluating Monolingual and Multilingual Large Language Models for Greek Question Answering: The DemosQA Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we address this research gap in Greek QA by contributing: (i) DemosQA, a novel dataset, which is constructed using social media user questions and community-reviewed answers to better capture the Greek social and cultural zeitgeist; (ii) a memory-efficient LLM evaluation framework adaptable to diverse QA datasets and languages; and (iii) an extensive evaluation of 11 monolingual and multilingual LLMs on 6 human-curated Greek QA datasets using 3 different prompting strategies. |
Charalampos Mastrokostas; Nikolaos Giarelis; Nikos Karacapilidis; | arxiv-cs.CL | 2026-02-18 |
| 458 | MoL: Adaptive Mixture-of-Length Reasoning for Efficient Question Answering with Context Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Mixture-of-Length (MoL), an approach for Question Answering (QA) with context that aims to improve the balance between reasoning quality and response efficiency. |
GUOCONG LI et. al. | iclr | 2026-02-17 |
| 459 | A Structured, Tagged, and Localized Visual Question Answering Dataset with Full Sentence Answers and Scene Graphs for Chest X-ray Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these limitations, we introduce MIMIC-Ext-CXR-QBA (abbr.We automatically generated our VQA dataset from scene graphs (also made available), which we constructed using LLM-based information extraction from radiology reports. |
Philip Müller; Friederike Jungmann; Georgios Kaissis; Daniel Rueckert; | iclr | 2026-02-17 |
| 460 | M4PQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose M4PQA, a human-annotated comprehensive paper QA dataset in the field of artificial intelligence, with 13,948 papers and 1,246 questions, that encompasses multi-task, multi-modal and instance-level evaluation. |
TIANCHENG HUANG et. al. | iclr | 2026-02-17 |
| 461 | Improving MLLMs in Embodied Exploration and Question Answering with Human-Inspired Memory Modeling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a non-parametric memory framework that explicitly disentangles episodic and semantic memory for embodied exploration and question answering. |
Ji Li; Jing Xia; Mingyi Li; Shiyan Hu; | arxiv-cs.RO | 2026-02-17 |
| 462 | Spatial Reasoning with Vision-Language Models in Ego-Centric Multi-View Scenes IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our results reveal a notable performance gap between human level scores and VLM performance, highlighting that current VLMs still fall short of human level spatial understanding (SU). To bridge this gap, we propose Ego3D-VLM, a post-training framework that enhances 3D spatial reasoning of VLMs. |
MOHSEN GHOLAMI et. al. | iclr | 2026-02-17 |
| 463 | QuRL: Rubrics As Judge For Open-Ended Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these limitations, we introduce a schema for generating case-wise rubrics that are question-specific, content-based and stylistically sensitive, thereby evaluating both factual soundness and writing quality. Building on this schema, we propose QuRL (Open-Ended QA with Rubric-guided Reinforcement Learning), a framework that automatically mines rubrics for each question from easily accessible online sources and leverages them as reward signals. |
Xiyu Wei; Qingwei Zong; Xiaoguang Li; Eugene J. Yu; Sujian Li; | iclr | 2026-02-17 |
| 464 | SmartChunk Retrieval: Query-Aware Chunk Compression with Planning for Efficient Document RAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SmartChunk retrieval, a query-adaptive framework for efficient and robust long-document question answering (QA). |
Xuechen Zhang; Koustava Goswami; Samet Oymak; Jiasi Chen; Nedim Lipka; | iclr | 2026-02-17 |
| 465 | FrugalRAG: Less Is More in RL Finetuning for Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FrugalRAG, a two-stage finetuning framework that adaptively _reduces_ the number of retrieval steps based on a question’s difficulty. |
Abhinav Java; Srivathsan Koundinyan; Nagarajan Natarajan; Amit Sharma; | iclr | 2026-02-17 |
| 466 | A$^2$Search: Ambiguity-Aware Question Answering with Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present A$^2$Search, an annotation-free, end-to-end training framework to recognize and handle ambiguity. |
FENGJI ZHANG et. al. | iclr | 2026-02-17 |
| 467 | VidBridge-R1: Bridging QA and Captioning for RL-based Video Understanding Models with Intermediate Proxy Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Naively combining reward signals from these tasks results in mutual performance degradation, which we attribute to a conflict between their opposing task natures. To address this challenge, we propose a novel training framework built upon two intermediate proxy tasks: DarkEventInfer, which presents videos with masked event segments, requiring models to infer the obscured content based on contextual video cues; and MixVidQA, which presents interleaved video sequences composed of two distinct clips, challenging models to isolate and reason about one while disregarding the other. |
XINLONG CHEN et. al. | iclr | 2026-02-17 |
| 468 | HIMM: Human-Inspired Long-Term Memory Modeling for Embodied Exploration and Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Deploying Multimodal Large Language Models as the brain of embodied agents remains challenging, particularly under long-horizon observations and limited context budgets. Existing … |
Ji Li; Bo Wang; Jingfan Xia; Mingyi Li; Shiyan Hu; | ArXiv | 2026-02-17 |
| 469 | EgoNight: Towards Egocentric Vision Understanding at Night with A Challenging Benchmark IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most existing benchmarks for egocentric vision understanding focus primarily on daytime scenarios, overlooking the low-light conditions that are inevitable in real-world applications. To investigate this gap, we present EgoNight, the first comprehensive benchmark for nighttime egocentric vision, with visual question answering (VQA) as the core task. |
DEHENG ZHANG et. al. | iclr | 2026-02-17 |
| 470 | EarthSE: A Benchmark Evaluating Earth Scientific Exploration Capability for Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present a comprehensive and professional benchmark for the Earth sciences, designed to evaluate the capabilities of LLMs in scientific exploration within this domain, spanning from fundamental to advanced levels.Leveraging a corpus of 100,000 research papers, we first construct two Question Answering (QA) datasets: Earth-Iron, which offers extensive question coverage for broad assessment, and Earth-Silver, which features a higher level of difficulty to evaluate professional depth. |
WANGHAN XU et. al. | iclr | 2026-02-17 |
| 471 | Are LLMs Really Not Knowledgeable? Mining The Submerged Knowledge in LLMs’ Memory Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By analyzing the token-level output distributions, we find that correct answers often appear among high-probability candidates, despite not being selected. Motivated by this, we propose Hits@k, a novel metric to evaluate latent knowledge retention independent of answer surface form. |
Xingjian Tao; Yiwei Wang; Yujun Cai; Zhicheng Yang; Jing Tang; | iclr | 2026-02-17 |
| 472 | CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present CounselBench, a large-scale benchmark developed with 100 mental health professionals to evaluate and stress-test large language models (LLMs) in realistic help-seeking scenarios. |
YAHAN LI et. al. | iclr | 2026-02-17 |
| 473 | Uncertainty As Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we focus on UQ for the contextual QA task and propose a theoretically grounded approach to quantify \emph{epistemic uncertainty}. |
YAVUZ FARUK BAKMAN et. al. | iclr | 2026-02-17 |
| 474 | IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address this gap, we introduce IndicVisionBench, the first large-scale benchmark centered on the Indian subcontinent.In addition, we release a paired parallel corpus of annotations across 10 Indic languages, creating a unique resource for analyzing cultural and linguistic biases in VLMs. |
ALI FARAZ et. al. | iclr | 2026-02-17 |
| 475 | AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by how humans link information associatively, we propose AssoMem, a novel framework constructing an associative memory graph that anchors dialogue utterances to automatically extracted clues. |
KAI ZHANG et. al. | iclr | 2026-02-17 |
| 476 | Query-Guided Spatial–Temporal–Frequency Interaction for Music Audio–Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in those methods, the audio input is primarily treated as complementary to video analysis, and the textual question information contributes minimally to audio–visual understanding, as it is typically integrated only in the final stages of reasoning. To address these limitations, we propose a novel Query-guided Spatial–Temporal–Frequency (QSTar) interaction method, which effectively incorporates question-guided clues and exploits the distinctive frequency-domain characteristics of audio signals, alongside spatial and temporal perception, to enhance audio–visual understanding. |
Kun Li; Michael Ying Yang; Sami Sebastian Brandt; | iclr | 2026-02-17 |
| 477 | SAFER: Risk-Constrained Sample-then-Filter in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, prior works unrealistically assume that admissible answers for all instances can be obtained via finite sampling, even for open-ended QA scenarios that lack a fixed and finite solution space. To address this, we introduce a two-stage risk control framework comprising abstention-aware **SA**mpling and conformalized **F**ilt**ER**ing (SAFER). |
Qingni Wang; Yue Fan; Xin Eric Wang; | iclr | 2026-02-17 |
| 478 | Addressing Pitfalls in The Evaluation of Uncertainty Estimation Methods for Natural Language Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose using several alternative risk indicators for risk correlation experiments that improve robustness of empirical assessment of UE algorithms for NLG. |
Mykyta Ielanskyi; Kajetan Schweighofer; Lukas Aichberger; Sepp Hochreiter; | iclr | 2026-02-17 |
| 479 | When Silence Is Golden: Can LLMs Learn to Abstain in Temporal QA and Beyond? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We instead frame abstention as a teachable skill and introduce a pipeline that couples Chain-of-Thought (CoT) supervision with Reinforcement Learning (RL) guided by abstention-aware rewards. |
Xinyu Zhou; Chang Jin; Carsten Eickhoff; Zhijiang Guo; Seyed Ali Bahrainian; | iclr | 2026-02-17 |
| 480 | Same Content, Different Representations: A Controlled Study for Table QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To support detailed analysis, we introduce a diagnostic benchmark with splits along table size, join requirements, query complexity, and schema quality. |
Yue Zhang; Seiji Maekawa; Nikita Bhutani; | iclr | 2026-02-17 |
| 481 | Automating Construction Contract Question Answering Using Large Language Model and Fine-tuning Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
MINGYU ZHANG et. al. | Expert Syst. Appl. | |
| 482 | KenLumachiQuAD – A Question Answering Dataset for Kenyan Luhya Lumarachi Language for Machine Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Question-answering (QA) datasets play a crucial role in testing and training machine learning models, from which we can develop practical end-user applications, such as internet search, dialogue systems, and chatbots. |
Barack Wamkaya Wanjawa; Lawrence Muchemi; Evans Miriti; | East African Journal of Information Technology | 2026-02-16 |
| 483 | IKIA: Image-Knowledge Internalization Assistance Model for Medical Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Yurun Bi; Xingang Wang; Yuteng Xiao; Yudong Zhang; | International Journal of Machine Learning and Cybernetics | 2026-02-16 |
| 484 | A Multi-Agent Framework for Medical AI: Leveraging Fine-Tuned GPT, LLaMA, and DeepSeek R1 for Evidence-Based and Bias-Aware Clinical Query Processing Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a multi-agent medical QA framework that combines complementary LLMs with evidence retrieval, uncertainty estimation, and bias checks to improve answer reliability. |
Naeimeh Nourmohammadi; Md Meem Hossain; The Anh Han; Safina Showkat Ara; Zia Ush Shamszaman; | arxiv-cs.CL | 2026-02-15 |
| 485 | Index Light, Reason Deep: Deferred Visual Ingestion for Visual-Dense Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes the Deferred Visual Ingestion (DVI) framework, adopting a demand-side ingestion strategy: the indexing phase performs only lightweight metadata extraction, deferring visual understanding to the moment users pose specific questions. |
Tao Xu; | arxiv-cs.CL | 2026-02-15 |
| 486 | Differentially Private Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Particularly for RAG systems, DP can reduce the usefulness of the augmented contexts leading to increase risk of hallucination from the LLMs. Motivated by these challenges, we present DP-KSA, a novel privacy-preserving RAG algorithm that integrates DP using the propose-test-release paradigm. |
Tingting Tang; James Flemings; Yongqin Wang; Murali Annavaram; | arxiv-cs.CR | 2026-02-15 |
| 487 | SRA: Semantic Relation-Aware Flowchart Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the issue, we propose a novel Semantic Relation-Aware (SRA) FlowchartQA approach. |
Xinyu Li; Bowei Zou; Yuchong Chen; Yifan Fan; Yu Hong; | arxiv-cs.MM | 2026-02-14 |
| 488 | ReFilter: Improving Robustness of Retrieval-Augmented Generation Via Gated Filter Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite their effectiveness at modest retrieval scales, these methods often fail to scale gracefully as the number of retrieved candidates k increases: Larger k improves evidence coverage, yet realistic top-k retrieval inevitably contains irrelevant or redundant content and increases the inference cost. To address these limitations, we propose ReFilter, a novel latent-based fusion framework that performs token-level filtering and fusion. |
YIXIN CHEN et. al. | arxiv-cs.CL | 2026-02-13 |
| 489 | Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose the Unified Memory Agent (UMA), an end-to-end reinforcement learning framework that unifies memory operations and question answering within a single policy. |
Kehao Zhang; Shangtong Gui; Sheng Yang; Wei Chen; Yang Feng; | arxiv-cs.LG | 2026-02-13 |
| 490 | RoadscapesQA: A Multitask, Multimodal Dataset for Visual Question Answering on Indian Roads Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we describe the data collection and annotation process, present key dataset statistics, and provide initial baselines for image QA tasks using vision-language models. |
Vijayasri Iyer; Maahin Rathinagiriswaran; Jyothikamalesh S; | arxiv-cs.CV | 2026-02-13 |
| 491 | TraceBack: Multi-Agent Decomposition for Fine-Grained Table Attribution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing table QA systems rarely provide fine-grained attribution, so even correct answers often lack verifiable grounding, limiting trust in high-stakes settings. We address this with TraceBack, a modular multi-agent framework for scalable, cell-level attribution in single-table QA. |
TEJAS ANVEKAR et. al. | arxiv-cs.CL | 2026-02-13 |
| 492 | Who Is The Richest Club in The Championship? Detecting and Rewriting Underspecified Questions Improve QA Performance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that this gap is partly due to underspecified questions – queries whose interpretation cannot be uniquely determined without additional context. To test this hypothesis, we introduce an LLM-based classifier to identify underspecified questions and apply it to several widely used QA datasets, finding that 16% to over 50% of benchmark questions are underspecified and that LLMs perform significantly worse on them. |
Yunchong Huang; Gianni Barlacchi; Sandro Pezzelle; | arxiv-cs.CL | 2026-02-12 |
| 493 | MultiCube-RAG for Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Built on the cube structure, we propose MultiCube-RAG, a training-free method consisting of multiple cubes for multi-step reasoning and retrieval. |
JIMENG SHI et. al. | arxiv-cs.CL | 2026-02-11 |
| 494 | RSHallu: Dual-Mode Hallucination Evaluation for Remote-Sensing Multimodal Large Language Models with Domain-Tailored Mitigation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal large language models (MLLMs) are increasingly adopted in remote sensing (RS) and have shown strong performance on tasks such as RS visual grounding (RSVG), RS visual question answering (RSVQA), and multimodal dialogue. |
ZIHUI ZHOU et. al. | arxiv-cs.CV | 2026-02-11 |
| 495 | Comprehensive Comparison of RAG Methods Across Multi-Domain Conversational QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a comprehensive empirical study of vanilla and advanced RAG methods across eight diverse conversational QA datasets spanning multiple domains. |
Klejda Alushi; Jan Strich; Chris Biemann; Martin Semmann; | arxiv-cs.CL | 2026-02-10 |
| 496 | AnalyticsGPT: An LLM Workflow for Scientometric Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces AnalyticsGPT, an intuitive and efficient large language model (LLM)-powered workflow for scientometric question answering. |
Khang Ly; Georgios Cheirmpos; Adrian Raudaschl; Christopher James; Seyed Amin Tabatabaei; | arxiv-cs.CL | 2026-02-10 |
| 497 | CAPID: Context-Aware PII Detection for Question-Answering Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To achieve privacy-preserving PII detection, we propose CAPID, a practical approach that fine-tunes a locally owned small language model (SLM) that filters sensitive information before it is passed to LLMs for QA. |
MARIIA PONOMARENKO et. al. | arxiv-cs.CR | 2026-02-10 |
| 498 | Engineering Trustworthy Retrieval-Augmented Generation for EU Electricity Market Regulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents the design of a domain-specific Retrieval-Augmented Generation (RAG) system for EU electricity market regulations, explicitly engineered to deliver source-grounded, traceable and low-hallucination answers. |
Șener Ali; Simona-Vasilica Oprea; Adela Bâra; | Electronics | 2026-02-10 |
| 499 | PulseLM: A Foundation Dataset and Benchmark for PPG-Text Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce PulseLM, a large-scale PPG-text dataset designed to bridge raw PPG waveforms and natural language through a unified, closed-ended question answering (QA) formulation. |
HUNG MANH PHAM et. al. | arxiv-cs.CL | 2026-02-10 |
| 500 | Vista: Scene-Aware Optimization for Streaming Video Question Answering Under Post-Hoc Queries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Vista, a novel framework for scene-aware streaming video QA that enables efficient and scalable reasoning over continuous video streams. |
HAOCHENG LU et. al. | arxiv-cs.CV | 2026-02-09 |
| 501 | CoRect: Context-Aware Logit Contrast for Hidden State Rectification to Resolve Knowledge Conflicts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through layer-wise analysis, we attribute this failure to a parametric suppression phenomenon: specifically, in deep layers, certain FFN layers overwrite context-sensitive representations with memorized priors. To address this, we propose CoRect (Context-Aware Logit Contrast for Hidden State Rectification). |
Xuhua Ma; Richong Zhang; Zhijie Nie; | arxiv-cs.CL | 2026-02-08 |
| 502 | Long-Context Long-Form Question Answering for Legal Domain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we address the challenges of long-context question answering in context of long-form answers given the idiosyncrasies of legal documents. |
ANAGHA KULKARNI et. al. | arxiv-cs.CL | 2026-02-06 |
| 503 | ViHERMES: A Graph-Grounded Multihop Question Answering Benchmark and System for Vietnamese Healthcare Regulations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce the Vietnamese Healthcare Regulations-Multihop Reasoning Dataset (ViHERMES), a benchmark designed for multihop QA over Vietnamese healthcare regulatory documents. |
LONG S. T. NGUYEN et. al. | arxiv-cs.CL | 2026-02-06 |
| 504 | IRPAPERS: A Visual Document Benchmark for Scientific Retrieval and Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce IRPAPERS, a benchmark of 3,230 pages from 166 scientific papers, with both an image and an OCR transcription for each page. |
CONNOR SHORTEN et. al. | arxiv-cs.IR | 2026-02-05 |
| 505 | CompactRAG: Reducing LLM Calls and Token Overhead in Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose CompactRAG, a simple yet effective framework that decouples offline corpus restructuring from online reasoning. |
HAO YANG et. al. | arxiv-cs.CL | 2026-02-05 |
| 506 | MediQAl: A French Medical Question Answering Dataset for Knowledge and Reasoning Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Abstract This work introduces MediQAl, a French medical question answering dataset designed to evaluate the capabilities of language models in factual medical recall and reasoning over real-world clinical scenarios. |
Adrien Bazoge; | Scientific Data | 2026-02-05 |
| 507 | Decomposed Prompting Does Not Fix Knowledge Gaps, But Helps Models Say I Don’t Know Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We find that although accuracy gains from decomposition diminish in frontier models, disagreements between prompting regimes remain highly indicative of potential errors. |
Dhruv Madhwal; Lyuxin David Zhang; Dan Roth; Tomer Wolfson; Vivek Gupta; | arxiv-cs.CL | 2026-02-04 |
| 508 | RA-QA: Towards Respiratory Audio-based Health Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This new data resource contains about 7.5 million QA pairs spanning more than 60 attributes and three question types: single verification, multiple choice, and open-ended questions. Building upon this dataset, we introduce a novel benchmark that compares audio-text generation models with traditional audio classifiers to evaluate their respective performance. |
Gaia A. Bertolino; Yuwei Zhang; Tong Xia; Domenico Talia; Cecilia Mascolo; | arxiv-cs.SD | 2026-02-04 |
| 509 | Knowing When to Answer: Adaptive Confidence Refinement for Reliable Audio-Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a formal problem formulation for \textit{Reliable} Audio-Visual Question Answering ($\mathcal{R}$-AVQA), where we prefer abstention over answering incorrectly. |
DINH PHU TRAN et. al. | arxiv-cs.LG | 2026-02-04 |
| 510 | No Shortcuts to Culture: Indonesian Multi-hop Question Answering for Complex Cultural Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a new framework that systematically transforms single-hop cultural questions into multi-hop reasoning chains spanning six clue types (e.g., commonsense, temporal, geographical). |
Vynska Amalia Permadi; Xingwei Tan; Nafise Sadat Moosavi; Nikos Aletras; | arxiv-cs.CL | 2026-02-03 |
| 511 | OmniRAG-Agent: Agentic Omnimodal Reasoning for Low-Resource Long Audio-Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Long-horizon omnimodal question answering answers questions by reasoning over text, images, audio, and video. Despite recent progress on OmniLLMs, low-resource long audio-video QA still suffers from costly dense encoding, weak fine-grained retrieval, limited proactive planning, and no clear end-to-end optimization.To address these issues, we propose OmniRAG-Agent, an agentic omnimodal QA method for budgeted long audio-video reasoning. |
YIFAN ZHU et. al. | arxiv-cs.CL | 2026-02-03 |
| 512 | JSynFlow: Japanese Synthesised Flowchart Visual Question Answering Dataset Built with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, developing VLMs with precise flowchart understanding requires large-scale datasets of flowchart images and corresponding text, the creation of which is highly time-consuming. To address this challenge, we introduce JSynFlow, a synthesised visual QA dataset for Japanese flowcharts, generated using large language models (LLMs). |
Hiroshi Sasaki; | arxiv-cs.CV | 2026-02-03 |
| 513 | ST-Raptor: An Agentic System for Semi-Structured Table QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Text-to-SQL methods typically require converting semi-structured tables into structured formats, inevitably leading to information loss, while approaches like Text-to-Code and multimodal LLM-based QA struggle with complex layouts and often yield inaccurate answers. To address these limitations, we present ST-Raptor, an agentic system for semi-structured table QA. |
JINXIU QU et. al. | arxiv-cs.AI | 2026-02-03 |
| 514 | P-RAG: Prompt-Enhanced Parametric RAG with LoRA and Selective CoT for Biomedical and Multi-Hop QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our contributions include: (1) LoRA-based fine-tuning of LLaMA-3.2-1B-Instruct for biomedical QA, (2) introduction of P-RAG with Chain-of-Thought prompting, and (3) state-of-the-art results on PubMedQA and 2WikiMultihopQA. |
Xingda Lyu; Gongfu Lyu; Zitai Yan; Yuxin Jiang; | arxiv-cs.CL | 2026-02-01 |
| 515 | CRAFT: Calibrated Reasoning with Answer-Faithful Traces Via Reinforcement Learning for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Traditional chain-of-thought generation often deviates from required structured output formats, leading to incomplete or malformed structured content. To address these challenges, we propose CRAFT (Calibrated Reasoning with Answer-Faithful Traces), a Group Relative Policy Optimization (GRPO) based reinforcement learning framework that trains models to perform faithful reasoning during response generation. |
YU LIU et. al. | arxiv-cs.CL | 2026-02-01 |
| 516 | PARSE: An Open-Domain Reasoning Question Answering Benchmark for Persian Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce PARSE, the first open-domain Persian reasoning QA benchmark, containing 10,800 questions across Boolean, multiple-choice, and factoid formats, with diverse reasoning types, difficulty levels, and answer structures. |
Jamshid Mozafari; Seyed Parsa Mousavinasab; Adam Jatowt; | arxiv-cs.CL | 2026-02-01 |
| 517 | Inferential Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Inferential QA — a new task that challenges models to infer answers from answer-supporting passages which provide only clues. |
Jamshid Mozafari; Hamed Zamani; Guido Zuccon; Adam Jatowt; | arxiv-cs.CL | 2026-02-01 |
| 518 | MARA: A Multimodal Adaptive Retrieval-Augmented Framework for Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, current approaches rely on query-agnostic document representations that overlook salient content and use static top-k evidence selection, which fails to adapt to the uncertain distribution of relevant information. To address these limitations, we propose the Multimodal Adaptive Retrieval-Augmented (MARA) framework, which introduces query-adaptive mechanisms to both retrieval and generation. |
HUI WU et. al. | arxiv-cs.IR | 2026-02-01 |
| 519 | Understanding QA Generation: Extracting Parametric and Contextual Knowledge with CQA for Low Resource Bangla Language Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our work not only introduces a novel framework for analyzing knowledge sources in Bangla QA but also uncovers critical findings that open up broader directions for counterfactual reasoning in low-resource language settings. |
Umme Abira Azmary; MD Ikramul Kayes; Swakkhar Shatabda; Farig Yousuf Sadeque; | arxiv-cs.CL | 2026-02-01 |
| 520 | Explainable Visual Question Answering: A Survey on Methods, Datasets and Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
YAXIAN WANG et. al. | Inf. Fusion | 2026-02-01 |
| 521 | A Framework of Knowledge Graph-Enhanced Large Language Model Based on Global Planning Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Knowledge graphs (KGs) can provide structured knowledge to assist large language models (LLMs) in interpretable reasoning. Knowledge graph question answering (KGQA) is a typical … |
YADING LI et. al. | IEEE Transactions on Knowledge and Data Engineering | 2026-02-01 |
| 522 | A Visual-textual Mutual Guidance Fusion Network for Remote Sensing Visual Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
HAOLIN LIU et. al. | Pattern Recognit. | 2026-02-01 |
| 523 | MedSpeak: A Knowledge Graph-Aided ASR Error Correction Framework for Spoken Medical QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose MedSpeak, a novel knowledge graph-aided ASR error correction framework that refines noisy transcripts and improves downstream answer prediction by leveraging both semantic relationships and phonetic information encoded in a medical knowledge graph, together with the reasoning power of LLMs. |
YUTONG SONG et. al. | arxiv-cs.CL | 2026-01-31 |
| 524 | DeALOG: Decentralized Multi-Agents Log-Mediated Reasoning Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DeALOG, a decentralized multi-agent framework for multimodal question answering. |
Abhijit Chakraborty; Ashish Raj Shekhar; Shiven Agarwal; Vivek Gupta; | arxiv-cs.CL | 2026-01-31 |
| 525 | Benchmarking Uncertainty Calibration in Large Language Model Long-Form Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce the first large-scale benchmark for evaluating UQ metrics in reasoning-demanding QA studying calibration of UQ methods, providing an extensible open-source framework to reproducibly assess calibration. |
Philip Müller; Nicholas Popovič; Michael Färber; Peter Steinbach; | arxiv-cs.CL | 2026-01-30 |
| 526 | TSAQA: Time Series Analysis Question And Answering Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TSAQA, a novel unified benchmark designed to broaden task coverage and evaluate diverse temporal analysis capabilities. |
BAOYU JING et. al. | arxiv-cs.AI | 2026-01-30 |
| 527 | Gender Disparities in StackOverflow’s Community-Based Question Answering: A Matter of Quantity Versus Quality Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we investigate whether answer quality is influenced by gender using a combination of human evaluations and automated assessments powered by Large Language Models. |
Maddalena Amendola; Cosimo Rulli; Carlos Castillo; Andrea Passarella; Raffaele Perego; | arxiv-cs.CY | 2026-01-30 |
| 528 | CE-GOCD: Central Entity-Guided Graph Optimization for Community Detection to Augment LLM Scientific Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This impairs the LLM’s comprehension of scientific literature, hindering the comprehensiveness and specificity of its responses. To address this, we propose Central Entity-Guided Graph Optimization for Community Detection (CE-GOCD), a method that augments LLMs’ scientific question answering by explicitly modeling and leveraging semantic substructures within academic knowledge graphs. |
JIAYIN LAN et. al. | arxiv-cs.CL | 2026-01-29 |
| 529 | Ontology-grounded Knowledge Graphs for Mitigating Hallucinations in Large Language Models for Clinical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Mohamed Ali; Zaki Taha; Mohamed Mabrouk Morsey; | Journal of Biomedical Informatics | 2026-01-28 |
| 530 | SRCR: Faithful Structured Reasoning with Curriculum Reinforcement Learning for Explainable Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
YUE FAN et. al. | Information Processing & Management | 2026-01-27 |
| 531 | MQADet: A Plug-and-play Paradigm for Enhancing Open-vocabulary Object Detection Via Multimodal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing open-vocabulary detectors often suffer from visual-textual misalignment and long-tailed category imbalance, leading to poor performance when handling objects described by complex, long-tailed textual queries. To overcome these challenges, we propose Multimodal Question Answering Detection (MQADet), a universal plug-and-play paradigm that enhances existing open-vocabulary detectors by leveraging the cross-modal reasoning capabilities of multimodal large language models (MLLMs). |
CAIXIONG LI et. al. | Scientific Reports | 2026-01-27 |
| 532 | Query-Guided Spatial-Temporal-Frequency Interaction for Music Audio-Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in those methods, the audio input is primarily treated as complementary to video analysis, and the textual question information contributes minimally to audio–visual understanding, as it is typically integrated only in the final stages of reasoning. To address these limitations, we propose a novel Query-guided Spatial–Temporal–Frequency (QSTar) interaction method, which effectively incorporates question-guided clues and exploits the distinctive frequency-domain characteristics of audio signals, alongside spatial and temporal perception, to enhance audio–visual understanding. |
Kun Li; Michael Ying Yang; Sami Sebastian Brandt; | arxiv-cs.CV | 2026-01-27 |
| 533 | Disaster Question Answering with LoRA Efficiency and Accurate End Position Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work introduces a disaster-focused question answering system based on Japanese disaster situations and response experiences. |
Takato Yasuno; | arxiv-cs.CL | 2026-01-27 |
| 534 | V-Loop: Visual Logical Loop Verification for Hallucination Detection in Medical Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent introspective detection methods, particularly uncertainty-based approaches, offer computational efficiency but are fundamentally indirect, as they estimate predictive uncertainty for an image-question pair rather than verifying the factual correctness of a specific answer. To address this limitation, we propose Visual Logical Loop Verification (V-Loop), a training-free and plug-and-play framework for hallucination detection in medical VQA. |
Mengyuan Jin; Zehui Liao; Yong Xia; | arxiv-cs.CV | 2026-01-26 |
| 535 | PaperSearchQA: Learning to Search and Reason Over Scientific Papers with RLVR Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Most RLVR search agents tackle general-domain QA, which limits their relevance to technical AI systems in science, engineering, and medicine. In this work we propose training agents to search and reason over scientific papers — this tests technical question-answering, it is directly relevant to real scientists, and the capabilities will be crucial to future AI Scientist systems. |
JAMES BURGESS et. al. | arxiv-cs.LG | 2026-01-26 |
| 536 | UniPACT: A Multimodal Framework for Prognostic Question Answering on Raw ECG and Structured EHR Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Language Models (LLMs) offer a powerful reasoning engine for this task but struggle to natively process these heterogeneous, non-textual data types. To address this, we propose UniPACT (Unified Prognostic Question Answering for Clinical Time-series), a unified framework for prognostic question answering that bridges this modality gap. |
Jialu Tang; Tong Xia; Yuan Lu; Aaqib Saeed; | arxiv-cs.LG | 2026-01-25 |
| 537 | DeepEra: A Deep Evidence Reranking Agent for Scientific Retrieval-Augmented Generated Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Retrieval-Augmented Generation (RAG) enhances LLMs by incorporating knowledge from external sources, thereby providing credible evidence for scientific question answering. But existing retrieval and reranking methods remain vulnerable to passages that are semantically similar but logically irrelevant, often reducing factual reliability and amplifying hallucinations.To address this challenge, we propose a Deep Evidence Reranking Agent (DeepEra) that integrates step-by-step reasoning, enabling more precise evaluation of candidate passages beyond surface-level semantics. |
HAOTIAN CHEN et. al. | arxiv-cs.CL | 2026-01-23 |
| 538 | Beyond Factual QA: Mentorship-Oriented Question Answering Over Long-Form Multilingual Content Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MentorQA, the first multilingual dataset and evaluation framework for mentorship-focused question answering from long-form videos, comprising nearly 9,000 QA pairs from 180 hours of content across four languages. |
Parth Bhalerao; Diola Dsouza; Ruiwen Guan; Oana Ignat; | arxiv-cs.CL | 2026-01-23 |
| 539 | DF-RAG: Query-Aware Diversity for Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, RAG is often challenged by reasoning-intensive question-answering (QA), since common retrieval methods like cosine similarity maximize relevance at the cost of introducing redundant content, which can reduce information recall. To address this, we introduce Diversity-Focused Retrieval-Augmented Generation (DF-RAG), which systematically incorporates diversity into the retrieval step to improve performance on complex, reasoning-intensive QA benchmarks. |
SAADAT HASAN KHAN et. al. | arxiv-cs.CL | 2026-01-23 |
| 540 | Mind The Ambiguity: Aleatoric Uncertainty Quantification in LLMs for Safe Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The deployment of Large Language Models in Medical Question Answering is severely hampered by ambiguous user queries, a significant safety risk that demonstrably reduces answer accuracy in high-stakes healthcare settings. In this paper, we formalize this challenge by linking input ambiguity to aleatoric uncertainty (AU), which is the irreducible uncertainty arising from underspecified input. |
YAOKUN LIU et. al. | arxiv-cs.CL | 2026-01-23 |
| 541 | ManuRAG: Multi-modal Retrieval Augmented Generation for Manufacturing Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ManuRAG, an innovative multi-modal RAG framework designed for manufacturing QA, incorporating specialized techniques to improve answer accuracy, reliability, and interpretability. |
Yunqing Li; Zihan Dong; Farhad Ameri; Jianbang Zhang; | arxiv-cs.CE | 2026-01-21 |
| 542 | LogicScore: Fine-grained Logic Evaluation of Conciseness, Completeness, and Determinateness in Attributed Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Current evaluation methods for Attributed Question Answering (AQA) suffer from \textit{attribution myopia}: they emphasize verification of isolated statements and their … |
ZHICHAO YAN et. al. | ArXiv | 2026-01-21 |
| 543 | \textsc{LogicScore}: Fine-grained Logic Evaluation of Conciseness, Completeness, and Determinateness in Attributed Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Consequently, Large Language Models (LLMs) often produce factually grounded yet logically incoherent responses with elusive deductive gaps. To mitigate this limitation, we present \textsc{LogicScore}, a unified evaluation framework that shifts the paradigm from local assessment to global reasoning scrutiny. |
ZHICHAO YAN et. al. | arxiv-cs.CL | 2026-01-21 |
| 544 | ManuRAG: Multi-modal Retrieval Augmented Generation for Manufacturing Question Answering (Early Version) Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: The evolution of digital manufacturing requires intelligent Question Answering (QA) systems that can seamlessly integrate and analyze complex multi-modal data, such as text, … |
Yunqing Li; Zihan Dong; Farhad Ameri; Jianbang Zhang; | ArXiv | 2026-01-21 |
| 545 | CCFQA: A Benchmark for Cross-Lingual and Cross-Modal Speech and Text Factuality Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing benchmarks for evaluating the reliability of Multimodal Large Language Models (MLLMs) predominantly focus on textual or visual modalities with a primary emphasis on English, which creates a gap in evaluation when processing multilingual input, especially in speech. To bridge this gap, we propose a novel Cross-lingual and Cross-modal Factuality benchmark (CCFQA). |
YEXING DU et. al. | aaai | 2026-01-20 |
| 546 | NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent works attempt to overcome these limitations through heuristic approaches; however, they lack explicit mechanisms for encoding temporal relationships and fail to provide any formal guarantees that the sampled context actually encodes the compositional or causal logic required by the question. To address these foundational gaps, we introduce NeuS-QA, a training-free, plug-and-play neuro-symbolic pipeline for LVQA. |
SAHIL SHAH et. al. | aaai | 2026-01-20 |
| 547 | QKVQA: Question-Focused Filtering for Knowledge-based VQA Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Visual Question Answering (VQA) is the task of answering questions based on image content. Building upon this, Knowledge-Based VQA (KB-VQA) requires models to answer questions … |
WEI YE et. al. | ArXiv | 2026-01-20 |
| 548 | Collaborative Enhancement of Large and Small Models for Question Answering Via Dual Knowledge Transfer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our statistical analysis reveals a complementary phenomenon between large language model-based question answering (QA) and small model-based QA. To facilitate dual knowledge transfer between these two paradigms, this paper introduces a collaborative enhancement method of large and small models for question answering. |
Shaofei Wang; Yunan Liu; Xiaolan Tang; Wenlong Chen; | aaai | 2026-01-20 |
| 549 | EgoCross: Benchmarking Multimodal Large Language Models for Cross-Domain Egocentric Video Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, real-world deployment inevitably encounters domain shifts, where target domains differ substantially in both visual style and semantic content. To bridge this gap, we introduce EgoCross, a comprehensive benchmark designed to evaluate the cross-domain generalization of MLLMs in EgocentricQA. |
YANJUN LI et. al. | aaai | 2026-01-20 |
| 550 | PKR-QA: A Benchmark for Procedural Knowledge Reasoning with Knowledge Module Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To generate question-answer pairs, we design graph traversal templates where each template is applied systematically over PKG. To enable interpretable reasoning, we propose a neurosymbolic approach called Knowledge Module Learning (KML), which learns procedural relations via neural modules and composes them for structured reasoning with LLMs. |
THANH-SON NGUYEN et. al. | aaai | 2026-01-20 |
| 551 | Improving Long-Context Summarization with Multi-Granularity Retrieval Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Notably, the human brain naturally integrates and summarizes prior knowledge upon reading a given text, progressively formulating a comprehensive understanding. Motivated by this cognitive process, we propose the Hierarchical Two-Stage Summarization-based Information Retrieval (HTSIR) method, which preprocesses the corpus prior to retrieval, summarizes continuous texts to obtain integrated information, and constructs a retrieval tree with varying summary granularities. |
Xueyu Chen; Kaitao Song; Zifan Song; Dongsheng Li; Cairong Zhao; | aaai | 2026-01-20 |
| 552 | Domain-Adaptation Through Synthetic Data: Fine-Tuning Large Language Models for German Law Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents an effective method for adapting advanced LLMs to German legal question answering through a novel synthetic data generation approach. |
ALI HAMZA BASHIR et. al. | arxiv-cs.CL | 2026-01-20 |
| 553 | DomainCQA: Crafting Knowledge-Intensive QA from Domain-Specific Charts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose DomainCQA, a framework for constructing domain-specific CQA benchmarks that emphasize both visual comprehension and knowledge-intensive reasoning. |
YUJING LU et. al. | aaai | 2026-01-20 |
| 554 | Mnemosyne: Accelerating Multi-Hop Question Answering Via Cache Hit Order Fitting Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing methods focus on the internal structure and ignore the misalignment between the queries’ arrival order and cache hit order. To tackle this, we propose Mnemosyne, a cache hit order fitting method designed to accelerate the RAG progress for MHQA. |
Haizhou Du; Jiujiu Li; Dongyang Li; Luobin Huang; Lisheng Wang; | aaai | 2026-01-20 |
| 555 | MIRAGE: Scaling Test-Time Inference with Parallel Graph-Retrieval-Augmented Reasoning Chains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: As a result, these approaches can lead to error accumulation throughout the reasoning chain, which significantly limits its effectiveness in medical question-answering (QA) tasks where both accuracy and traceability are critical requirements. To address these challenges, we propose MIRAGE (Multi-path Inference with Retrieval-Augmented Graph Exploration), a novel test-time scalable reasoning framework that performs dynamic multi-path inference over structured medical knowledge graphs. |
KAIWEN WEI et. al. | aaai | 2026-01-20 |
| 556 | Teaching Large Language Models to Maintain Contextual Faithfulness Via Synthetic Tasks and Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Teaching large language models (LLMs) to be faithful in the provided context is crucial for building reliable information-seeking systems. Therefore, we propose a systematic framework, CANOE, to reduce faithfulness hallucinations of LLMs across different downstream tasks without human annotations. |
SHUZHENG SI et. al. | aaai | 2026-01-20 |
| 557 | Beyond Chains: Bridging Large Language Models and Knowledge Bases in Complex Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by semantic parsing methods, we propose PDRR: a four-stage framework consisting of Predict, Decompose, Retrieve, and Reason. |
Yihua Zhu; Qianying Liu; Akiko Aizawa; Hidetoshi Shimodaira; | aaai | 2026-01-20 |
| 558 | DigimonGPT: An Evolvable Agent with Hierarchical Human-like Memory for Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the challenge, we introduce DigimonGPT, an evolvable VideoQA agent inspired by cognitive psychology. |
BORUI LI et. al. | aaai | 2026-01-20 |
| 559 | EvalQAG: A Framework for Automatic Complex QA Generation and A Benchmark QA Dataset for Policy Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce EvalQAG, a framework for generating high-quality QA pairs from renewable energy policy documents. |
Kirtan Brijeshbhai Soni; Krish Rupapara; Arpit Rana; Ghanshyam Verma; Paul Buitelaar; | aaai | 2026-01-20 |
| 560 | STRIDE-QA: Visual Question Answering Dataset for Spatiotemporal Reasoning in Urban Driving Scenes Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-Language Models (VLMs) have been applied to autonomous driving to support decision-making in complex real-world scenarios. |
Keishi Ishihara; Kento Sasaki; Tsubasa Takahashi; Daiki Shiono; Yu Yamaguchi; | aaai | 2026-01-20 |
| 561 | SAR: A Structure-Aligned Reasoning Framework for Temporal Knowledge Graph Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The core issue lies in structural misalignment: treating structured, temporally sensitive graph queries as plain text often causes LLMs to retrieve or reason with semantically similar but structurally incorrect facts, resulting in critical inaccuracies. To address this, we introduce SAR (Structure-Aligned Reasoning), a novel TKGQA framework that integrates LLM reasoning tightly with the explicit subject–predicate–object–time schema inherent in knowledge graphs. |
Qianyi Hu; Jiaxue Liu; Xinhui Tu; Shoujin Wang; | aaai | 2026-01-20 |
| 562 | StreamKV: Streaming Video Question-Answering with Segment-based KV Cache Retrieval and Compression IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose StreamKV, a training-free framework that seamlessly equips Video-LLMs with advanced KV cache retrieval and compression. |
YILONG CHEN et. al. | aaai | 2026-01-20 |
| 563 | OIDA-QA: A Multimodal Benchmark for Analyzing The Opioid Industry Documents Archive Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The complexity, multimodal nature, and specialized characteristics of these healthcare-related legal and corporate documents necessitate more advanced methods and models tailored to specific data types and detailed annotations, ensuring the precision and professionalism in the analysis. In this paper, we tackle this challenge by organizing the original dataset according to document attributes and constructing a benchmark with 400k training documents and 10k for testing. |
XUAN SHEN et. al. | aaai | 2026-01-20 |
| 564 | MacVQA: Adaptive Memory Allocation and Global Noise Filtering for Continual Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current methods often struggle with balancing knowledge retention, adaptation, and robust feature representation. To address these challenges, we propose a novel framework with adaptive memory allocation and global noise filtering called MacVQA for visual question answering. |
ZHIFEI LI et. al. | aaai | 2026-01-20 |
| 565 | Knowledge-based Question Answering Using Graph Neural Networks and Contextual Language Representations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Abstract This work introduces a novel question answering (QA) framework that integrates commonsense knowledge from ConceptNet with deep contextual embeddings from BERT using a graph neural network for structured reasoning. |
Mohamed Samir; Naglaa Fathy; Walaa Gad; | Scientific Reports | 2026-01-20 |
| 566 | Better Datasets Start from RefineLab: Automatic Optimization for High-Quality Dataset Refinement Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce RefineLab, the first LLM‑driven framework that automatically refines raw QA textual data into high-quality datasets under a controllable token‑budget constraint. |
Xiaonan Luo; Yue Huang; Ping He; Xiangliang Zhang; | aaai | 2026-01-20 |
| 567 | Cross-Granularity Hypergraph Retrieval-Augmented Generation for Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a novel RAG approach called HGRAG for MHQA that achieves cross-granularity integration of structural and semantic information via hypergraphs. |
Changjian Wang; Weihong Deng; Weili Guan; Quan Lu; Ning Jiang; | aaai | 2026-01-20 |
| 568 | GeoX-Bench: Benchmarking Cross-View Geo-Localization and Pose Estimation Capabilities of Large Multimodal Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large multimodal models (LMMs) have demonstrated remarkable capabilities across a wide range of tasks, however their knowledge and abilities in the cross-view geo-localization and pose estimation domains remain unexplored, despite potential benefits for navigation, autonomous driving, outdoor robotics, etc. To bridge this gap, we introduce GeoX-Bench, a comprehensive Benchmark designed to explore and evaluate the capabilities of LMMs in cross-view Geo-localization and pose estimation. |
YUSHUO ZHENG et. al. | aaai | 2026-01-20 |
| 569 | Self-Correction Distillation for Structured Data Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To improve the structured data QA ability of small-scale LLMs, we propose a self-correction distillation (SCD) method. |
YUSHAN ZHU et. al. | aaai | 2026-01-20 |
| 570 | ChartAttack: Testing The Vulnerability of LLMs to Malicious Prompting in Chart Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce ChartAttack, a novel framework for evaluating how MLLMs can be misused to generate misleading charts at scale. |
Jesus-German Ortiz-Barajas; Jonathan Tonglet; Vivek Gupta; Iryna Gurevych; | arxiv-cs.CL | 2026-01-19 |
| 571 | Pardon? Evaluating Conversational Repair in Large Audio-Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce a repair-aware evaluation setting that explicitly distinguishes between answerable and unanswerable audio inputs. |
SHUANGHONG HUANG et. al. | arxiv-cs.CL | 2026-01-19 |
| 572 | BioPulse-QA: A Dynamic Biomedical Question-Answering Benchmark for Evaluating Factuality, Robustness, and Bias in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: They also carry increasing risk of data leakage due to overlap with model pretraining corpora and often overlook critical dimensions such as robustness to linguistic variation and potential demographic biases. Materials and Methods: To address these gaps, we introduce BioPulse-QA, a benchmark that evaluates LLMs on answering questions from newly published biomedical documents including drug labels, trial protocols, and clinical guidelines. |
KRITI BHATTARAI et. al. | arxiv-cs.CL | 2026-01-18 |
| 573 | Augmenting Question Answering with A Hybrid RAG Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Structured-Semantic RAG (SSRAG), a hybrid architecture that enhances QA quality by integrating query augmentation, agentic routing, and a structured retrieval mechanism combining vector and graph based techniques with context unification. |
TIANYI YANG et. al. | arxiv-cs.CL | 2026-01-18 |
| 574 | SolarGPT-QA: A Domain-Adaptive Large Language Model for Educational Question Answering in Space Weather and Heliophysics Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce SolarGPT-QA, a question answering system based on a domain-adapted large language model built on the LLaMA-3 base model. |
Santosh Chapagain; MohammadReza EskandariNasab; Onur Vural; Shah Muhammad Hamdi; Soukaina Filali Boubrahimi; | arxiv-cs.LG | 2026-01-17 |
| 575 | AVIR: Adaptive Visual In-Document Retrieval for Efficient Multi-Page Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multi-page Document Visual Question Answering (MP-DocVQA) remains challenging because long documents not only strain computational resources but also reduce the effectiveness of the attention mechanism in large vision-language models (LVLMs). We tackle these issues with an Adaptive Visual In-document Retrieval (AVIR) framework. |
Zongmin Li; Yachuan Li; Lei Kang; Dimosthenis Karatzas; Wenkang Ma; | arxiv-cs.CV | 2026-01-17 |
| 576 | Reasoning in Trees: Improving Retrieval-Augmented Generation for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For multi-hop QA tasks, current iterative approaches predominantly rely on LLMs to self-guide and plan multi-step exploration paths during retrieval, leading to substantial challenges in maintaining reasoning coherence across steps from inaccurate query decomposition and error propagation. To address these issues, we introduce Reasoning Tree Guided RAG (RT-RAG), a novel hierarchical framework for complex multi-hop QA. |
YULING SHI et. al. | arxiv-cs.CL | 2026-01-16 |
| 577 | A Topic-aware Evaluation of ChatGPT’s Semantic Alignment with Community Answers Using BERTScore and BERTopic Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study introduces a topic-sensitive evaluation framework that enhances understanding of large language model behavior in real-world QA scenarios and supports the development of more effective and explainable conversational artificial intelligence (AI) systems. |
Mashael M. Alsulami; | PeerJ Computer Science | 2026-01-16 |
| 578 | From Single to Multi-Agent Reasoning: Advancing GeneGPT for Genomics QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We replicate GeneGPT and propose GenomAgent, a multi-agent framework that efficiently coordinates specialized agents for complex genomics queries. |
Kimia Abedini; Farzad Shami; Gianmaria Silvello; | arxiv-cs.AI | 2026-01-15 |
| 579 | ReGraM: Region-First Knowledge Graph Reasoning for Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We argue that the core challenge lies not in expanding access to knowledge, but in identifying and reasoning over the appropriate subset of evidence for each query. |
Chaerin Lee; Sohee Park; Hyunsik Na; Daseon Choi; | arxiv-cs.CL | 2026-01-14 |
| 580 | An Electronic Product Carbon Footprint Dataset for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work establishes a foundation for training advanced language models to automate aggregation and standardization of emissions data for ICT systems. |
Kaiwen Zhao; Ajesh Koyatan Chathoth; Bharathan Balaji; Stephen Lee; | Scientific Data | 2026-01-14 |
| 581 | EHRNavigator: A Multi-Agent System for Patient-Level Clinical Question Answering Over Heterogeneous Electronic Health Records Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Clinical decision-making increasingly relies on timely and context-aware access to patient information within Electronic Health Records (EHRs), yet most existing natural language question-answering (QA) systems are evaluated solely on benchmark datasets, limiting their practical relevance. To overcome this limitation, we introduce EHRNavigator, a multi-agent framework that harnesses AI agents to perform patient-level question answering across heterogeneous and multimodal EHR data. |
LINGFEI QIAN et. al. | arxiv-cs.CL | 2026-01-14 |
| 582 | Beyond Structured Knowledge: Performance Boundaries of ChatGPT in Geological-hazard Question Answering and The Need for Human-in-the-loop Oversight Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We conduct a version-specific evaluation of ChatGPT-4o for geological-hazard question answering using a transparent, rubric-based design. |
SAIER WU et. al. | Frontiers in Earth Science | 2026-01-13 |
| 583 | STAGE: A Benchmark for Knowledge Graph Construction, Question Answering, and In-Script Role-Playing Over Movie Screenplays Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce STAGE (Screenplay Text, Agents, Graphs and Evaluation), a unified benchmark for narrative understanding over full-length movie screenplays. |
QIUYU TIAN et. al. | arxiv-cs.CL | 2026-01-13 |
| 584 | Enhancing Large Language Models for Knowledge Graph Question Answering Via Multi-granularity Knowledge Injection and Structured Reasoning Path-augmented Prompting Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Chuanyang Gong; Zhihua Wei; Wenhao Tao; Duoqian Miao; | Information Processing & Management | 2026-01-13 |
| 585 | Judging Against The Reference: Uncovering Knowledge-Driven Failures in LLM-Judges on QA Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We identify a critical failure mode of such reference-based LLM QA evaluation: when the provided reference conflicts with the judge model’s parametric knowledge, the resulting scores become unreliable, substantially degrading evaluation fidelity. To study this phenomenon systematically, we introduce a controlled swapped-reference QA framework that induces reference-belief conflicts. |
DONGRYEOL LEE et. al. | arxiv-cs.CL | 2026-01-12 |
| 586 | Fine-Tuning Vs. RAG for Multi-Hop Question Answering with Novel Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we systematically compare parametric and non-parametric knowledge injection methods for open-domain multi-hop question answering. |
Zhuoyi Yang; Yurun Song; Iftekhar Ahmed; Ian Harris; | arxiv-cs.CL | 2026-01-11 |
| 587 | Efficient Visual Question Answering Pipeline for Autonomous Driving Via Scene Region Compression Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This focus limits their practical deployment in real-time autonomous driving scenarios. To tackle this issue, we propose an efficient VLM framework for autonomous driving VQA tasks, SRC-Pipeline. |
Yuliang Cai; Dongqiangzi Ye; Zitian Chen; Chongruo Wu; | arxiv-cs.CV | 2026-01-11 |
| 588 | UETQuintet at BioCreative IX – MedHopQA: Enhancing Biomedical QA with Selective Multi-hop Reasoning and Contextual Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a model designed to effectively address both direct and sequential questions. |
Quoc-An Nguyen; Thi-Minh-Thu Vu; Bich-Dat Nguyen; Dinh-Quang-Minh Tran; Hoang-Quynh Le; | arxiv-cs.CL | 2026-01-11 |
| 589 | FinCARDS: Card-Based Analyst Reranking for Financial Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose FinCards, a structured reranking framework that reframes financial evidence selection as constraint satisfaction under a finance-aware schema. |
YIXI ZHOU et. al. | arxiv-cs.IR | 2026-01-11 |
| 590 | N2N-GQA: Noise-to-Narrative for Graph-Based Table-Text Question Answering Using LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our key insight is that multi-hop reasoning requires understanding relationships between evidence pieces: by modeling documents as graph nodes with semantic relationships as edges, we identify bridge documents connecting reasoning steps, a capability absent in list-based retrieval. |
Mohamed Sharafath; Aravindh Annamalai; Ganesh Murugan; Aravindakumar Venugopalan; | arxiv-cs.CL | 2026-01-10 |
| 591 | Do Language Models Reason Across Languages? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a simple two-hop question answering setting, where answering a question requires making inferences over two multilingual documents. |
Yan Meng; Wafaa Mohammed; Christof Monz; | arxiv-cs.CL | 2026-01-10 |
| 592 | VideoAuto-R1: Video Auto Reasoning Via Thinking Once, Answering Twice IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we first demonstrate that for RL-trained video models, direct answering often matches or even surpasses CoT performance, despite CoT producing step-by-step analyses at a higher computational cost. Motivated by this, we propose VideoAuto-R1, a video understanding framework that adopts a reason-when-necessary strategy. |
SHUMING LIU et. al. | arxiv-cs.CV | 2026-01-08 |
| 593 | A Lightweight and Explainable Vision-Language Framework for Crop Disease Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work presents a lightweight vision-language framework for crop and disease identification from leaf images. |
Md. Zahid Hossain; Most. Sharmin Sultana Samu; Md. Rakibul Islam; Md. Siam Ansary; | arxiv-cs.CV | 2026-01-08 |
| 594 | From National Curricula to Cultural Awareness: Constructing Open-Ended Culture-Specific Question Answering Dataset Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CuCu, an automated multi-agent LLM framework that transforms national textbook curricula into open-ended, culture-specific question-answer pairs. |
Haneul Yoo; Won Ik Cho; Geunhye Kim; Jiyoon Han; | arxiv-cs.CL | 2026-01-08 |
| 595 | DisastQA: A Comprehensive Benchmark for Evaluating Question Answering in Disaster Management Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce DisastQA, a large-scale benchmark of 3,000 rigorously verified questions (2,000 multiple-choice and 1,000 open-ended) spanning eight disaster types. |
ZHITONG CHEN et. al. | arxiv-cs.CL | 2026-01-07 |
| 596 | From Chains to Graphs: Self-Structured Reasoning for General-Domain LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Self-Graph Reasoning (SGR), a framework that enables LLMs to explicitly represent their reasoning process as a structured graph before producing the final answer. |
YINGJIAN CHEN et. al. | arxiv-cs.CL | 2026-01-07 |
| 597 | Self-MedRAG: A Self-Reflective Hybrid Retrieval-Augmented Generation Framework for Reliable Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While Retrieval-Augmented Generation (RAG) mitigates these issues by incorporating external knowledge, conventional single-shot retrieval often fails to resolve complex biomedical queries requiring multi-step inference. To address this, we propose Self-MedRAG, a self-reflective hybrid framework designed to mimic the iterative hypothesis-verification process of clinical reasoning. |
Jessica Ryan; Alexander I. Gumilang; Robert Wiliam; Derwin Suhartono; | arxiv-cs.IR | 2026-01-07 |
| 598 | BioPIE: A Biomedical Protocol Information Extraction Dataset for High-Reasoning-Complexity Experiment Question Answer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing biomedical datasets focus on general or coarsegrained knowledge and thus fail to support the fine-grained experimental reasoning demanded by HID and MSR. To address this gap, we introduce Biomedical Protocol Information Extraction Dataset (BioPIE), a dataset that provides procedure-centric KGs of experimental entities, actions, and relations at a scale that supports reasoning over biomedical experiments across protocols. |
HAOFEI HOU et. al. | arxiv-cs.AI | 2026-01-07 |
| 599 | EpiQAL: Benchmarking Large Language Models in Epidemiological Question Answering for Enhanced Alignment and Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present EpiQAL, the first diagnostic benchmark for epidemiological question answering across diverse diseases, comprising three subsets built from open-access literature. |
MINGYANG WEI et. al. | arxiv-cs.CL | 2026-01-06 |
| 600 | LittiChoQA: Literary Texts in Indic Languages Chosen for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We address the scarcity of long-context QA resources for Indic languages by introducing LittiChoQA, the largest literary QA dataset to date covering many languages spoken in the Gangetic plains of India. |
Aarya Khandelwal; Ritwik Mishra; Rajiv Ratn Shah; | arxiv-cs.CL | 2026-01-06 |
| 601 | When Do Tools and Planning Help Large Language Models Think? A Cost- and Latency-Aware Benchmark Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Modern large language models (LLMs) increasingly rely on inference-time planning and external tools to improve reasoning. We benchmark this behavior on two real-world settings: … |
S. Ghoshal; Ali Al-Bustami; | SoutheastCon 2026 | 2026-01-06 |
| 602 | SentGraph: Hierarchical Sentence Graph for Multi-hop Retrieval-Augmented Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing chunk-based retrieval often provides irrelevant and logically incoherent context, leading to incomplete evidence chains and incorrect reasoning during answer generation. To address these challenges, we propose SentGraph, a sentence-level graph-based RAG framework that explicitly models fine-grained logical relationships between sentences for multi-hop question answering. |
JUNLI LIANG et. al. | arxiv-cs.CL | 2026-01-06 |
| 603 | DeCode: Decoupling Content and Delivery for Medical QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce DeCode, a training-free, model-agnostic framework that adapts existing LLMs to produce contextualized answers in clinical settings. |
Po-Jen Ko; Chen-Han Tsai; Yu-Shao Peng; | arxiv-cs.CL | 2026-01-05 |
| 604 | Question Answering for Multi-Release Systems: A Case Study at Ciena Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Motivated by the observed inaccuracy of state-of-the-art question-answering techniques on multi-release system documents, we propose QAMR, a chatbot designed to answer questions across multi-release system documentation. |
PARHAM KHAMSEPOUR et. al. | arxiv-cs.SE | 2026-01-05 |
| 605 | Adversarial Question Answering Robustness: A Multi-Level Error Analysis and Mitigation Study Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We perform comprehensive multi-level error analysis using five complementary categorization schemes, identifying negation confusion and entity substitution as the primary failure modes. |
Agniv Roy Choudhury; Vignesh Ponselvan Rajasingh; | arxiv-cs.CL | 2026-01-05 |
| 606 | When Do Tools and Planning Help LLMs Think? A Cost- and Latency-Aware Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Modern large language models (LLMs) increasingly rely on inference-time planning and external tools to improve reasoning. We benchmark this behavior on two real-world settings: event-centric question answering over graph-structured knowledge (Event-QA) and persuasive response generation in Reddit ChangeMyView (CMV). |
Subha Ghoshal; Ali Al-Bustami; | arxiv-cs.CL | 2026-01-05 |
| 607 | Crop GraphRAG: Pest and Disease Knowledge Base Q&A System for Sustainable Crop Protection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By mitigating the limitations of large language models in specialized agricultural contexts, this study provides a pragmatic tool for intelligent QA in the agricultural domain and advances the application of AI in crop protection. |
HAO WU et. al. | Frontiers in Plant Science | 2026-01-05 |
| 608 | MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet existing VideoQA benchmarks struggle to capture this multimodal reasoning and are largely not open-ended, given the difficulty of evaluating free-form answers. In this paper, we introduce a novel open-ended multi-modal VideoQA benchmark, MovieRecapsQA created using movie recap videos–a distinctive type of YouTube content that summarizes a film by presenting its key events through synchronized visual (recap video) and textual (recap summary) modalities. |
Shaden Shaar; Bradon Thymes; Sirawut Chaixanien; Claire Cardie; Bharath Hariharan; | arxiv-cs.CV | 2026-01-05 |
| 609 | Augmenting Medical Visual Question Answering with Mixup, Label Smoothing, and Layer-wise Relevance Propagation EXplainable Artificial Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, an imbalance in the number and distribution of image and Question–Answer (QA) pairs poses challenges for developing robust models. This study proposes improving existing MVQA datasets using data augmentation techniques specifically Mixup and Label Smoothing—to address this issue. |
Sheerin Sitara Noor Mohamed; Kavitha Srinivasan; | PeerJ Computer Science | 2026-01-05 |
| 610 | PdfQA: Diverse, Challenging, and Realistic Question Answering Over PDFs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present pdfQA, a multi-domain 2K human-annotated (real-pdfQA) and 2K synthetic dataset (syn-pdfQA) differentiating QA pairs in ten complexity dimensions (e.g., file type, source modality, source position, answer type). |
TOBIAS SCHIMANSKI et. al. | arxiv-cs.CL | 2026-01-05 |
| 611 | CTIS-QA: Clinical Template-Informed Slide-level Question Answering for Pathology Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a clinical diagnosis template-based pipeline to systematically collect and structure pathological information. |
HAO LU et. al. | arxiv-cs.CV | 2026-01-04 |
| 612 | MambaFormer: Token-Level Guided Routing Mixture-of-Experts for Accurate and Efficient Clinical Assistance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The deployment of large language models (LLMs) in real-world clinical applications is constrained by the fundamental trade-off between computational cost and the efficiency of linear-time models. To address this, we propose an LLM-based MambaFormer hybrid Mixture-of-Experts (MoE) framework for efficient medical question-answering (QA) and clinical assistance. |
Hamad Khan; Saddam Hussain Khan; | arxiv-cs.CV | 2026-01-03 |
| 613 | Retrieval-Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Multi-hop question answering (QA) requires systems to iteratively retrieve evidence and reason across multiple hops. While recent RAG and agentic methods report strong results, … |
Yuelyu Ji; Zhuochun Li; Rui Meng; Daqing He; | ArXiv | 2026-01-02 |
| 614 | RAGPPI: Retrieval-Augmented Generation Benchmark for Protein-Protein Interactions in Drug Discovery Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Retrieving the biological impacts of protein-protein interactions (PPIs) is essential for target identification (Target ID) in drug development. Given the vast number of proteins … |
YOUNGSEUNG JEON et. al. | Conference of the European Chapter of the Association for … | 2026-01-01 |
| 615 | Task-Level Instructions Induction for Audio Question Answering from Few Examples Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Po-Chun Chen; Hen-Hsen Huang; Hsin-Hsi Chen; | Conference of the European Chapter of the Association for … | 2026-01-01 |
| 616 | DoKE: Domain Knowledge-enhanced Multimodal Large Language Model-based Framework for Chest X-ray Follow-up Medical Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Zhichuan Wang; Qiao Deng; Tiffany So; Wan Hang Keith Chiu; Edward S. Hui; | Biomed. Signal Process. Control. | 2026-01-01 |
| 617 | Dual-space Intervention for Mitigating Bias in Robust Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
RUNMIN WANG et. al. | Expert Syst. Appl. | 2026-01-01 |
| 618 | Combining Retrieved with Generated Contexts Via A Listwise Reranker for Open-domain Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Yi Zhao; Yanxiang He; Xiaotong Zhang; Weidong Wen; | Neurocomputing | 2026-01-01 |
| 619 | The Problem of Ambiguity in Table Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Jorge Osés Grijalba; L. Urena-Lopez; Eugenio Martínez Cámara; J. Camacho-Collados; | Conference of the European Chapter of the Association for … | 2026-01-01 |
| 620 | GlintLM: Graph-Layered Integration with Nodal Topology with Language Models – A Bipartite Approach to Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Zhuofan Chen; Yao Hui Hoon; Renne Ye Kai Ong; J. Wong; | Inf. Syst. | 2026-01-01 |
| 621 | Semantic Event Graphs for Long-Form Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Semantic Event Graphs (SEG), a lightweight symbolic interface between video and language that replaces raw frames with compact temporal interaction logs. |
Aradhya Dixit; Tianxi Liang; | arxiv-cs.CV | 2026-01-01 |
| 622 | Retrieval–Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This survey takes the execution procedure as the unit of analysis and introduces a four-axis framework covering (A) overall execution plan, (B) index structure, (C) next-step control (strategies and triggers), and (D) stop/continue criteria. |
Yuelyu Ji; Zhuochun Li; Rui Meng; Daqing He; | arxiv-cs.CL | 2026-01-01 |
| 623 | Enhancing The QA Model Through A Multi-domain Debiasing Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: By identifying errors related to lexical bias, numerical reasoning, and entity recognition, we develop a multi-domain debiasing framework incorporating knowledge distillation, debiasing techniques, and domain expansion. |
Yuefeng Wang; ChangJae Lee; | arxiv-cs.CL | 2026-01-01 |
| 624 | Explicit Abstention Knobs for Predictable Reliability in Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate whether confidence-based abstention provides reliable control over error rates in video question answering, and whether that control remains robust under distribution shift. |
Jorge Ortiz; | arxiv-cs.AI | 2025-12-31 |
| 625 | Intelligent Diabetes Question Answering System Based on Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
步翔 徐; | Artificial Intelligence and Robotics Research | 2025-12-31 |
| 626 | DermaVQA-DAS: Dermatology Assessment Schema (DAS) & Datasets for Closed-Ended Question Answering & Segmentation in Patient-Generated Dermatology Images Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent advances in dermatological image analysis have been driven by large-scale annotated datasets; however, most existing benchmarks focus on dermatoscopic images and lack patient-authored queries and clinical context, limiting their applicability to patient-centered care. To address this gap, we introduce DermaVQA-DAS, an extension of the DermaVQA dataset that supports two complementary tasks: closed-ended question answering (QA) and dermatological lesion segmentation. |
WEN-WAI YIM et. al. | arxiv-cs.CV | 2025-12-30 |
| 627 | HaluNet: Multi-Granular Uncertainty Modeling for Efficient Hallucination Detection in LLM Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present \textbf{HaluNet}, a lightweight and trainable neural framework that integrates multi granular token level uncertainties by combining semantic embeddings with probabilistic confidence and distributional uncertainty. |
CHAODONG TONG et. al. | arxiv-cs.CL | 2025-12-30 |
| 628 | Comparative Analysis of The Risk of Hadith Errors in Question-Answering Systems Based on Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study aims to conduct a comparative analysis of the risk of errors generated by Large Language Model-based Question-Answering systems in answering hadith-related questions. |
Hakkun Elmunsyah; | International Journal of Research and Scientific Innovation | 2025-12-30 |
| 629 | EdgeJury: Cross-Reviewed Small-Model Ensembles for Truthful Question Answering on Serverless Edge Inference Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present EdgeJury, a lightweight ensemble framework that improves truthfulness and robustness using only small instruction-tuned language models (3B-8B) suitable for serverless edge inference. |
Aayush Kumar; | arxiv-cs.LG | 2025-12-29 |
| 630 | Integrating Domain Knowledge for Financial QA: A Multi-Retriever RAG Approach with LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through comprehensive ablation experiments and error analysis, we find that domain-specific training with the SecBERT encoder significantly contributes to our best neural symbolic model surpassing the FinQA paper’s top model, which serves as our baseline. |
Yukun Zhang; Stefan Elbl Droguett; Samyak Jain; | arxiv-cs.CL | 2025-12-29 |
| 631 | Retrieval Augmented Question Answering: When Should LLMs Admit Ignorance? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We investigate the use of LLMs for retrieval augmented question answering. |
Dingmin Wang; Ji Ma; Shankar Kumar; | arxiv-cs.CL | 2025-12-29 |
| 632 | Chain-of-thought Reviewing and Correction for Time Series Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Different from purely textual tasks, time series data are inherently verifiable, enabling consistency checking between reasoning steps and the original input. Motivated by this property, we propose T3LLM, which performs multi-step reasoning with an explicit correction mechanism for time series question answering. |
Chen Su; Yuanhe Tian; Yan Song; | arxiv-cs.CL | 2025-12-27 |
| 633 | Uncertainty-Aware Dynamic Knowledge Graphs for Reliable Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a demonstration of uncertainty-aware dynamic KGs, a framework that combines (i) dynamic construction of evolving KGs, (ii) confidence scoring and uncertainty-aware retrieval, and (iii) an interactive interface for reliable and interpretable QA. |
YU TAKAHASHI et. al. | arxiv-cs.CL | 2025-12-26 |
| 634 | KG20C & KG20C-QA: Scholarly Knowledge Graph Benchmarks for Link Prediction and Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present KG20C and KG20C-QA, two curated datasets for advancing question answering (QA) research on scholarly data. |
Hung-Nghiep Tran; Atsuhiro Takasu; | arxiv-cs.IR | 2025-12-25 |
| 635 | Medical QA Dialogue Datasets in RAG Systems Performance Evaluation and ChatGPT Optimization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Using ChatGPT-3.5 as a baseline and extending to GPT-4o and GPT-5, we compare multiple retrieval pipelines, including dense retrieval, Cross-Encoder reranking, Reciprocal Rank Fusion (RRF), and Cascade RRF→Rerank. |
Muretijiang Muhetaer; Ailimulati Yusupu; Wang Yifan; Munire Mutalipu; Fan Hao; | Scientific Reports | 2025-12-24 |
| 636 | MAR:Multi-Agent Reflexion Improves Reasoning Abilities in LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, continual reflections of the same LLM onto itself exhibit degeneration of thought, where the LLM continues to repeat the same errors again and again even with the knowledge that its wrong. To address this problem, we instead introduce multi-agent with multi-persona debators as the method to generate reflections. |
ONAT OZER et. al. | arxiv-cs.AI | 2025-12-23 |
| 637 | CycleChart: A Unified Consistency-Based Learning Framework for Bidirectional Chart Understanding and Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CycleChart, a consistency-based learning framework for bidirectional chart understanding and generation. |
Dazhen Deng; Sen Yang; Yuchen He; Yuan Tian; Yingcai Wu; | arxiv-cs.CL | 2025-12-22 |
| 638 | Simulated Patient Systems Powered By Large Language Model-based AI Agents Offer Potential for Transforming Medical Education IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: A user study with medical students shows that AIPatient delivers high fidelity, usability, and educational value, matching or exceeding human-simulated patients in history-taking. |
HUIZI YU et. al. | Communications Medicine | 2025-12-19 |
| 639 | Toward Ethical AI Through Bayesian Uncertainty in Neural Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We explore Bayesian reasoning as a means to quantify uncertainty in neural networks for question answering. |
Riccardo Di Sipio; | arxiv-cs.CL | 2025-12-19 |
| 640 | DrugRAG: Enhancing Pharmacy LLM Performance Through A Novel Retrieval-Augmented Generation Pipeline Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Methods: We benchmarked eleven existing LLMs with varying parameter sizes (8 billion to 70+ billion) using a 141-question pharmacy dataset. |
HOUMAN KAZEMZADEH et. al. | arxiv-cs.CL | 2025-12-16 |
| 641 | Evaluating The Capability of Video Question Generation for Expert Knowledge Elicitation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: For a continuous improvement of VQG models, we propose a protocol that evaluates the ability by simulating question-answering communication with experts using a question-to-answer retrieval. |
Huaying Zhang; Atsushi Hashimoto; Tosho Hirasawa; | arxiv-cs.CV | 2025-12-16 |
| 642 | Code-Driven LLM Agent for One-Shot Explanatory Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose OneCoLA (One-shot and training-free Code-driven LLM Agent), a novel framework for Multimodal Explanatory Visual Question Answering (MEVQA). |
Zuyi Zhou; Dizhan Xue; Baoyuan Qi; Shengsheng Qian; Changsheng Xu; | ACM Transactions on Multimedia Computing, Communications, … | 2025-12-16 |
| 643 | Parameter Efficient Multimodal Instruction Tuning for Romanian Vision Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we contribute to reducing the multimodal NLP resource gap for Romanian. |
George-Andrei Dima; Dumitru-Clementin Cercel; | arxiv-cs.CL | 2025-12-16 |
| 644 | Improving VQA Reliability: A Dual-Assessment Approach with Self-Reflection and Cross-Model Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the susceptibility of VLMs to hallucinations can lead to overconfident yet incorrect answers, severely undermining answer reliability. To address this, we propose Dual-Assessment for VLM Reliability (DAVR), a novel framework that integrates Self-Reflection and Cross-Model Verification for comprehensive uncertainty estimation. |
XIXIAN WU et. al. | arxiv-cs.CV | 2025-12-16 |
| 645 | Toward Ambulatory Vision: Learning Visually-Grounded Active View Selection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Visually Grounded Active View Selection (VG-AVS), a task that selects the most informative next viewpoint using only the visual information in the current image, without relying on scene memory or external knowledge. |
Juil Koo; Daehyeon Choi; Sangwoo Youn; Phillip Y. Lee; Minhyuk Sung; | arxiv-cs.CV | 2025-12-15 |
| 646 | Anatomy-Aware Adaptation of Pre-Trained Models for Medical Difference Visual Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Medical Difference Visual Question Answering (Med-Diff-VQA) is a challenging and clinically significant task that requires identifying and interpreting subtle anatomical … |
Qianying Zhou; Yuhan Gao; Hua Zou; Fei Luo; Xiwen Bai; | 2025 IEEE International Conference on Bioinformatics and … | 2025-12-15 |
| 647 | KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose KFS-Bench, the first benchmark for key frame sampling in long video question answering (QA), featuring multi-scene annotations to enable direct and robust evaluation of sampling strategies. |
Zongyao Li; Kengo Ishida; Satoshi Yamazaki; Xiaotong Ji; Jianquan Liu; | arxiv-cs.CV | 2025-12-15 |
| 648 | Hybrid Retrieval-Augmented Generation for Robust Multilingual Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We develop and evaluate a multilingual Retrieval-Augmented Generation pipeline specifically designed for question answering on noisy historical documents. |
Anthony Mudet; Souhail Bakkali; | arxiv-cs.DL | 2025-12-14 |
| 649 | Time-Aware Complex Question Answering Over Temporal Knowledge Graph Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Luyi Bai; Tongyue Zhang; Guangchen Feng; | Data Knowl. Eng. | |
| 650 | Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking By Contrasting Layers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing RAG methods for simple and multi-hop question answering (QA) are still prone to incorrect retrievals and hallucinations. To address these limitations, we propose CoopRAG, a novel RAG framework for the question answering task in which a retriever and an LLM work cooperatively with each other by exchanging informative knowledge, and the earlier and later layers of the retriever model work cooperatively with each other to accurately rank the retrieved documents relevant to a given query. |
Youmin Ko; Sungjong Seo; Hyunjoon Kim; | arxiv-cs.CL | 2025-12-11 |
| 651 | NeuReg: Neuro-Symbolic QA Generation from Regulatory Compliance Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Education providers face increasing challenges in complying with complex and evolving funding regulations. While large language models (LLMs) have the potential to support … |
Umair Arshad; D. Corsar; I. Nkisi-Orji; | Proceedings of the 13th Knowledge Capture Conference 2025 | 2025-12-10 |
| 652 | MedBioRAG: Semantic Search and Retrieval-Augmented Generation with Large Language Models for Medical and Biological QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce MedBioRAG, a retrieval-augmented model designed to improve biomedical QA performance through a combination of semantic and lexical search, document retrieval, and supervised fine-tuning. |
Seonok Kim; | arxiv-cs.CL | 2025-12-10 |
| 653 | AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Here, we introduce AutoMedic, a multi-agent simulation framework that enables automated evaluation of LLMs as clinical conversational agents. |
Gyutaek Oh; Sangjoon Park; Byung-Hoon Kim; | arxiv-cs.CL | 2025-12-10 |
| 654 | DIANGPT: A Review on A Domain-Specific Question Answering System Using LORA and RAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This review evaluates DianGPT’s architectural pipeline, including dataset processing, supervised fine-tuning, retrieval workflows, and automated evaluation strate- gies. |
Pallavi SK; Pradeep Nayak; Nischitha Nischitha; Omkar KS; Omkar JK; | International Journal of Scientific Research in Engineering … | 2025-12-10 |
| 655 | SimpleDevQA: Benchmarking Large Language Models on Development Knowledge QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through this pipeline, we introduce SimpleDevQA, a multilingual benchmark derived from real user dialogues. |
JING ZHANG et. al. | arxiv-cs.SE | 2025-12-09 |
| 656 | BoundingDocs: A Unified Dataset for Document Question Answering with Spatial Annotations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Abstract We present a unified dataset for document Question-Answering (QA), which is obtained combining several public datasets related to Document AI and visually rich document understanding (VRDU). |
Simone Giovannini; Fabio Coppini; Andrea Gemelli; Simone Marinai; | International Journal on Document Analysis and Recognition … | 2025-12-06 |
| 657 | Modeling Contextual Passage Utility for Multihop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a lightweight approach to model contextual passage utility, accounting for inter-passage dependencies. |
Akriti Jain; Aparna Garimella; | arxiv-cs.CL | 2025-12-06 |
| 658 | Knowing What’s Missing: Assessing Information Sufficiency in Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we propose a structured Identify-then-Verify framework for robust sufficiency modeling. |
Akriti Jain; Aparna Garimella; | arxiv-cs.CL | 2025-12-06 |
| 659 | Collective Narrative Grounding: Community-Coordinated Data Contributions to Improve Local AI Systems Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large language model (LLM) question-answering systems often fail on community-specific queries, creating knowledge blind spots that marginalize local voices and reinforce epistemic injustice. We present Collective Narrative Grounding, a participatory protocol that transforms community stories into structured narrative units and integrates them into AI systems under community governance. |
Zihan Gao; Mohsin Y. K. Yousufi; Jacob Thebault-Spieker; | arxiv-cs.CL | 2025-12-05 |
| 660 | ArtistMus: A Globally Diverse, Artist-Centric Benchmark for Retrieval-Augmented Music Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce MusWikiDB, a vector database of 3.2M passages from 144K music-related Wikipedia pages, and ArtistMus, a benchmark of 1,000 questions on 500 diverse artists with metadata such as genre, debut year, and topic. |
Daeyong Kwon; SeungHeon Doh; Juhan Nam; | arxiv-cs.CL | 2025-12-05 |
| 661 | Optimizing Medical Question-Answering Systems: A Comparative Study of Fine-Tuned and Zero-Shot Large Language Models with RAG Framework Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present a retrieval-augmented generation (RAG) based medical QA system that combines domain-specific knowledge retrieval with open-source LLMs to answer medical questions. |
TASNIMUL HASSAN et. al. | arxiv-cs.CL | 2025-12-05 |
| 662 | Grounded Multilingual Medical Reasoning for Question Answering with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a method to generate multilingual reasoning traces grounded in factual medical knowledge. |
Pietro Ferrazzi; Aitor Soroa; Rodrigo Agerri; | arxiv-cs.CL | 2025-12-05 |
| 663 | PathFinder: MCTS and LLM Feedback-based Path Selection for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Hence, we propose PATHFINDER, an approach that: (i) uses Monte Carlo Tree Search to generate training path traces, (ii) improves training data quality by filtering erroneous and lengthy traces using sub-answer recall and LLM-as-a-judge verification, and (iii) reformulates sub-queries to handle failed retrieval cases. |
Durga Prasad Maram; Kalpa Gunaratna; Vijay Srinivasan; Haris Jeelani; Srinivas Chappidi; | arxiv-cs.LG | 2025-12-04 |
| 664 | Fine-Tuning BERT for Domain-Specific Question Answering: Toward Educational NLP Resources at University Scale Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we developed a chatbot for the University of Limerick’s Department of Electronic and Computer Engineering to provide course information to students. |
Aurélie Montfrond; | arxiv-cs.CL | 2025-12-04 |
| 665 | BookRAG: A Hierarchical Structure-aware Index-based Approach for Retrieval-Augmented Generation on Complex Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing RAG approaches often focus on general documents, and they overlook the fact that many real-world documents (such as books, booklets, handbooks, etc.) have a hierarchical structure, which organizes their content from different granularity levels, leading to poor performance for the QA task. To address these limitations, we introduce BookRAG, a novel RAG approach targeted for documents with a hierarchical structure, which exploits logical hierarchies and traces entity relations to query the highly relevant information. |
Shu Wang; Yingli Zhou; Yixiang Fang; | arxiv-cs.IR | 2025-12-02 |
| 666 | CryptoQA: A Large-scale Question-answering Dataset for AI-assisted Cryptography Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, their ability to perform deep reasoning and mathematical analysis, particularly for complex tasks as required in cryptography, remains poorly understood, largely due to the lack of suitable data for evaluation and training. To address this gap, we present CryptoQA, the first large-scale question-answering (QA) dataset specifically designed for cryptography. |
MAYAR ELFARES et. al. | arxiv-cs.CR | 2025-12-02 |
| 667 | Retrieving on A Topic Graph for Long Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
SHIWEI CHEN et. al. | Neurocomputing | 2025-12-01 |
| 668 | CRAFT Your Dataset: Task-Specific Synthetic Dataset Generation Through Corpus Retrieval and Augmentation IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We propose Corpus Retrieval and Augmentation for Fine-Tuning (CRAFT), a method for generating synthetic datasets, given a small number of user-written few-shots that demonstrate the task to be performed. |
Ingo Ziegler; Abdullatif Köksal; Desmond Elliott; Hinrich Schütze; | Transactions of the Association for Computational … | 2025-12-01 |
| 669 | Memory-Augmented Knowledge Fusion with Safety-Aware Decoding for Domain-Adaptive Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce Knowledge-Aware Reasoning and Memory-Augmented Adaptation (KARMA), a novel framework designed to enhance QA performance in care scenarios. |
Lei Fu; Xiang Chen; Kaige Gao Xinyue Huang; Kejian Tong; | arxiv-cs.CL | 2025-12-01 |
| 670 | Memory-enriched Thought-by-thought Framework for Complex Diagram Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
XINYU ZHANG et. al. | Comput. Vis. Image Underst. | 2025-12-01 |
| 671 | CAIRNS: Balancing Readability and Scientific Accuracy in Climate Adaptation Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present Climate Adaptation question-answering with Improved Readability and Noted Sources (CAIRNS), a framework that enables experts — farmer advisors — to obtain credible preliminary answers from complex evidence sources from the web. |
Liangji Kong; Aditya Joshi; Sarvnaz Karimi; | arxiv-cs.CL | 2025-12-01 |
| 672 | Zero-Shot Knowledge-Based Visual Question Answering with Frozen Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
JING LIU et. al. | Big Data Min. Anal. | 2025-12-01 |
| 673 | Advancing Academic Chatbots: Evaluation of Non Traditional Outputs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We implemented a prototype combining Meta’s LLaMA 3 70B open weight and OpenAI’s GPT 4o mini API based. |
Nicole Favero; Francesca Salute; Daniel Hardt; | arxiv-cs.CL | 2025-11-30 |
| 674 | ORCA: Open-ended Response Correctness Assessment for Audio Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ORCA (Open-ended Response Correctness Assessment), a framework that models the variability in human judgments using Beta distributions to predict both expected correctness and uncertainty. |
ŠIMON SEDLÁČEK et. al. | arxiv-cs.SD | 2025-11-28 |
| 675 | Tourism Question Answer System in Indian Language Using Domain-Adapted Foundation Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, a dataset comprising 7,715 Hindi QA pairs pertaining to Varanasi tourism was constructed and subsequently augmented with 27,455 pairs generated via Llama zero-shot prompting. |
Praveen Gatla; Nikita Kanwar; Gouri Sahoo; Rajesh Kumar Mundotiya; | arxiv-cs.CL | 2025-11-28 |
| 676 | Machine Learning and Deep Learning Techniques in Arabic Question Answering Systems: Innovations and Challenges Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, we recognize that the impact of AQAS extends beyond academia; it has significant implications for various sectors, including education, technology, and information access. Through this comprehensive examination, we aim to lay the groundwork for ongoing innovation and development in AQAS. |
Azza Mohamed; Khaled Abdelqader; Khaled Shaalan; | PeerJ Computer Science | 2025-11-28 |
| 677 | WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world Scenarios Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce WearVQA, the first benchmark specifically designed to evaluate the Visual Question Answering (VQA) capabilities of multi-model AI assistant on wearable devices like smart glasses. |
EUN CHANG et. al. | arxiv-cs.AI | 2025-11-27 |
| 678 | Unlocking Electronic Health Records: A Hybrid Graph RAG Approach to Safe Clinical AI for Patient QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current solutions typically isolate retrieval methods focusing either on structured data (SQL/Cypher) or unstructured semantic search but fail to integrate both simultaneously. This work presents MediGRAF (Medical Graph Retrieval Augmented Framework), a novel hybrid Graph RAG system that bridges this gap. |
Samuel Thio; Matthew Lewis; Spiros Denaxas; Richard JB Dobson; | arxiv-cs.CL | 2025-11-27 |
| 679 | JBE-QA: Japanese Bar Exam QA Dataset for Assessing Legal Domain Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce JBE-QA, a Japanese Bar Exam Question-Answering dataset to evaluate large language models’ legal knowledge. |
ZHIHAN CAO et. al. | arxiv-cs.CL | 2025-11-27 |
| 680 | KA-RAG: Integrating Knowledge Graphs and Agentic Retrieval-Augmented Generation for An Intelligent Educational Question-Answering Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Generative artificial intelligence (AI) and large language models (LLMs) are reshaping the landscape of intelligent educational systems; however, existing solutions often suffer from unstructured resource organization, limited interpretability, and suboptimal retrieval precision. To address these challenges, this study introduces KA-RAG, a course-oriented question answering (QA) framework that integrates a structured Knowledge Graph (KG) with an Agentic Retrieval-Augmented Generation (Agentic-RAG) workflow. |
Fangqun Gao; Shu Xu; Weiyan Hao; Tao Lu; | Applied Sciences | 2025-11-26 |
| 681 | LongVT: Incentivizing Thinking with Long Videos Via Native Tool Calling IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by how humans comprehend long videos – by first skimming globally and then examining relevant clips for details – we introduce LongVT, an end-to-end agentic framework that enables Thinking with Long Videos via interleaved Multimodal Chain-of-Tool-Thought. |
ZUHAO YANG et. al. | arxiv-cs.CV | 2025-11-25 |
| 682 | Chatty-KG: A Multi-Agent AI System for On-Demand Conversational Question Answering Over Knowledge Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Traditional KGQA systems preserve structure but typically support only single-turn QA, incur high latency, and struggle with coreference and context tracking. To address these limitations, we propose Chatty-KG, a modular multi-agent system for conversational QA over KGs. |
REHAM OMAR et. al. | arxiv-cs.CL | 2025-11-25 |
| 683 | SFA: Scan, Focus, and Amplify Toward Guidance-aware Answering for Video TextVQA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, the model must identify question-relevant textual cues and filter out redundant or irrelevant information to ensure answering is guided by the most relevant and informative cues. To address these challenges, we propose SFA, a training-free framework and the first Video-LLM-based method tailored for Video TextVQA, motivated by the human process of answering questions. |
HAIBIN HE et. al. | arxiv-cs.CV | 2025-11-25 |
| 684 | ENACT: Evaluating Embodied Cognition with World Modeling of Egocentric Interaction IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ENACT, a benchmark that casts evaluation of embodied cognition as world modeling from egocentric interaction in a visual question answering (VQA) format. |
QINENG WANG et. al. | arxiv-cs.AI | 2025-11-25 |
| 685 | GHR-VQA: Graph-guided Hierarchical Relational Reasoning for Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose GHR-VQA, Graph-guided Hierarchical Relational Reasoning for Video Question Answering (Video QA), a novel human-centric framework that incorporates scene graphs to capture intricate human-object interactions within video sequences. |
Dionysia Danai Brilli; Dimitrios Mallis; Vassilis Pitsikalis; Petros Maragos; | arxiv-cs.CV | 2025-11-25 |
| 686 | VeriSciQA: An Auto-Verified Dataset for Scientific Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although recent work uses LVLMs to synthesize data at scale, we identify systematic errors in their resulting QA pairs, stemming from LVLMs’ inherent limitations and information asymmetry between figures and text. To address these challenges, we propose a verification-centric Generate-then-Verify framework that first generates QA pairs with figure-associated textual context, then applies cross-modal consistency checks against figures along with auxiliary filters to eliminate erroneous pairs. |
Yuyi Li; Daoyuan Chen; Zhen Wang; Yutong Lu; Yaliang Li; | arxiv-cs.CV | 2025-11-24 |
| 687 | Vidi2: Large Multimodal Models for Video Understanding and Creation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To enable comprehensive evaluation of STG in practical settings, we introduce a new benchmark, VUE-STG, which offers four key improvements over existing STG datasets: 1) Video duration: spans from roughly 10s to 30 mins, enabling long-context reasoning; 2) Query format: queries are mostly converted into noun phrases while preserving sentence-level expressiveness; 3) Annotation quality: all ground-truth time ranges and bounding boxes are manually annotated with high accuracy; 4) Evaluation metric: a refined vIoU/tIoU/vIoU-Intersection scheme. |
VIDI TEAM et. al. | arxiv-cs.CV | 2025-11-24 |
| 688 | EAGER: Edge-Aligned LLM Defense for Robust, Efficient, and Accurate Cybersecurity Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present EAGER, an edge-aligned defense framework that integrates parameter-efficient quantization with domain-specific preference alignment to jointly optimize efficiency, robustness, and accuracy. |
Onat Gungor; Roshan Sood; Jiasheng Zhou; Tajana Rosing; | arxiv-cs.CR | 2025-11-24 |
| 689 | MedPerturbing LLMs: A Comparative Study of Toxicity, Prompt Tuning, and Jailbreaks in Medical QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we evaluate the toxicity of widely used general-purpose LLMs in medical question–answering tasks. |
Arash Asgari; Amirreza Naziri; Laleh Seyyed-Kalantari; | Proceedings of the AAAI Symposium Series | 2025-11-23 |
| 690 | RoadSceneVQA: Benchmarking Visual Question Answering in Roadside Perception Systems for Intelligent Transportation System Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Based on the above, we propose the baseline model RoadMind. |
RUNWEI GUAN et. al. | arxiv-cs.CV | 2025-11-22 |
| 691 | Health-oriented Multimodal Food Question Answering with Implicit and Explicit Knowledge Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Menghao Hu; Y. Song; Xiaoshan Yang; Yaowei Wang; Changsheng Xu; | ACM Transactions on Multimedia Computing, Communications … | 2025-11-22 |
| 692 | Measuring The Impact of Lexical Training Data Coverage on Hallucination Detection in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a complementary question: Does lexical training-data coverage of the question and/or generated answer provide additional signal for hallucination detection? |
Shuo Zhang; Fabrizio Gotti; Fengran Mo; Jian-Yun Nie; | arxiv-cs.CL | 2025-11-22 |
| 693 | SMILE: A Composite Lexical-Semantic Metric for Question-Answering Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Language Model (LLM) based evaluators, though powerful, come with drawbacks like high costs, bias, inconsistency, and hallucinations. To address these issues, we introduce SMILE: Semantic Metric Integrating Lexical Exactness, a novel approach that combines sentence-level semantic understanding with keyword-level semantic understanding and easy keyword matching. |
SHRIKANT KENDRE et. al. | arxiv-cs.CL | 2025-11-21 |
| 694 | HMA: A Hierarchical Multi-Agent System for Document Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Document Question Answering (DocQA) stands as a core task within the domain of natural language processing. Nevertheless, when confronted with complex documents that incorporate … |
Qihang Hou; Guowei Liu; Pinpin Zhu; | Proceedings of the 2025 4th International Conference on … | 2025-11-21 |
| 695 | EduMod-LLM: A Modular Approach for Designing Flexible and Transparent Educational Assistants Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce {\model}, a modular function-calling LLM pipeline, and present a comprehensive evaluation along three key axes: function calling strategies, retrieval methods, and generative language models. |
Meenakshi Mittal; Rishi Khare; Mihran Miroyan; Chancharik Mitra; Narges Norouzi; | arxiv-cs.CL | 2025-11-21 |
| 696 | MGA-VQA: Secure and Interpretable Graph-Augmented Visual Question Answering with Memory-Guided Protection Against Unauthorized Knowledge Use Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MGA-VQA, a multi-modal framework that integrates token-level encoding, spatial graph reasoning, memory-augmented inference, and question-guided compression. |
Ahmad Mohammadshirazi; Pinaki Prasad Guha Neogi; Dheeraj Kulshrestha; Rajiv Ramnath; | arxiv-cs.CV | 2025-11-21 |
| 697 | ESGBench: A Benchmark for Explainable ESG Question Answering in Corporate Sustainability Reports Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ESGBench, a benchmark dataset and evaluation framework designed to assess explainable ESG question answering systems using corporate sustainability reports. |
Sherine George; Nithish Saji; | arxiv-cs.CL | 2025-11-20 |
| 698 | FlipVQA-Miner: Cross-Page Visual Question-Answer Mining from Textbooks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an automated pipeline that extracts well-formed QA and visual-QA (VQA) pairs from educational documents by combining layout-aware OCR with LLM-based semantic parsing. |
ZHEN HAO WONG et. al. | arxiv-cs.AI | 2025-11-20 |
| 699 | Question Answering Models for Information Extraction from Perovskite Materials Science Literature Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we developed and tested a Question Answering (QA) approach to extract material-property relationships from scientific publications. |
Matilda Sipilä; Farrokh Mehryary; Sampo Pyysalo; Filip Ginter; Milica Todorović; | Communications Materials | 2025-11-20 |
| 700 | AVATAAR: Agentic Video Answering Via Temporal Adaptive Alignment and Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although large vision language models (LVLMs) have enhanced performance, they often face challenges with nuanced queries that demand both a comprehensive understanding and detailed analysis. To overcome these obstacles, we introduce AVATAAR, a modular and interpretable framework that combines global and local video context, along with a Pre Retrieval Thinking Agent and a Rethink Module. |
Urjitkumar Patel; Fang-Chun Yeh; Chinmay Gondhalekar; | arxiv-cs.CV | 2025-11-19 |
| 701 | HinTel-AlignBench: A Framework and Benchmark for Hindi-Telugu with English-Aligned Samples Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current multilingual VLM evaluations suffer from four major limitations: reliance on unverified auto-translations, narrow task/domain coverage, limited sample sizes, and lack of cultural and natively sourced Question-Answering (QA). To address these gaps, we present a scalable framework to evaluate VLMs in Indian languages and compare it with performance in English. |
RISHIKANT CHIGRUPAATII et. al. | arxiv-cs.CL | 2025-11-19 |
| 702 | Development of An Intelligent Information Retrieval System Based on Ontology, Linguistic Algorithms and Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study proposes a hybrid semantic question-answering (QA) system for the Kazakh language that integrates ontological modeling, linguistic processing, and large language models (LLMs). |
Assel Mukanova; Aizhan Nazyrova; Altanbek Zulkhazhav; Zhanar Lamasheva; Assem Dauletkaliyeva; | Applied Sciences | 2025-11-19 |
| 703 | Foundational Question Generation for Video Question Answering Via An Embedding-Integrated Approach Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Foundational Question Generation for Video Question Answering via an Embedding-Integrated Approach (FIQ), a framework designed to enhance the reasoning capability of VQA models by improving their foundational comprehension of video content. |
Ju-Young Oh; | arxiv-cs.CV | 2025-11-18 |
| 704 | BBox DocVQA: A Large Scale Bounding Box Grounded Dataset for Enhancing Reasoning in Document Visual Question Answer Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing DocVQA datasets are limited to the page level and lack fine grained spatial grounding, constraining the interpretability and reasoning capability of Vision Language Models (VLMs). To address this gap, we introduce BBox DocVQA a large scale, bounding box grounded dataset designed to enhance spatial reasoning and evidence localization in visual documents. |
WENHAN YU et. al. | arxiv-cs.DB | 2025-11-18 |
| 705 | Audio Question Answering with GRPO-Based Fine-Tuning and Calibrated Segment-Level Predictions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this report, we describe our submission to Track 5 of the DCASE 2025 Challenge for the task of Audio Question Answering(AQA). |
MARCEL GIBIER et. al. | arxiv-cs.SD | 2025-11-18 |
| 706 | Multi Table QA:Evaluating Modern LLM Strategies on Table0 Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce tool-augmented reasoning as a paradigm for multi-table QA, and systematically study two complementary strategies: (1) free-form tool interaction, where models iteratively call exploration and computation tools, and(2) structured agent workflows, which stage tool-use into exploration, preparation, and analysis phases. |
Letian Li; | Computers and Artificial Intelligence | 2025-11-18 |
| 707 | Beyond GeneGPT: A Multi-Agent Architecture with Open-Source LLMs for Enhanced Genomic Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, its reliance on a proprietary model limits scalability, increases operational costs, and raises concerns about data privacy and generalization. In this work, we revisit and reproduce GeneGPT in a pilot study using open source models, including Llama 3.1, Qwen2.5, and Qwen2.5 Coder, within a monolithic architecture; this allows us to identify the limitations of this approach. |
Haodong Chen; Guido Zuccon; Teerapong Leelanupab; | arxiv-cs.AI | 2025-11-18 |
| 708 | Descriptor: Distance-Annotated Traffic Perception Question Answering (DTPQA) Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this article, we provide the dataset itself along with the Python scripts used to create it, which can be used to generate additional data of the same kind. |
NIKOS THEODORIDIS et. al. | arxiv-cs.CV | 2025-11-17 |
| 709 | Collaborative QA Using Interacting LLMs. Impact of Network Structure, Node Capability and Distributed Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we model and analyze how a network of interacting LLMs performs collaborative question-answering (CQA) in order to estimate a ground truth given a distributed set of documents. |
Adit Jain; Vikram Krishnamurthy; Yiming Zhang; | arxiv-cs.AI | 2025-11-17 |
| 710 | QA-Noun: Representing Nominal Semantics Via Natural Language Question-Answer Pairs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce QA-Noun, a QA-based framework for capturing noun-centered semantic relations. |
Maria Tseytlin; Paul Roit; Omri Abend; Ido Dagan; Ayal Klein; | arxiv-cs.CL | 2025-11-16 |
| 711 | A Role-Aware Multi-Agent Framework for Financial Education QA Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Question answering (QA) plays a central role in financial education, yet existing large language model (LLM) approaches often fail to capture the nuanced and specialized reasoning … |
Andy Zhu; Yingjun Du; | Proceedings of the 6th ACM International Conference on AI … | 2025-11-14 |
| 712 | Local Hybrid Retrieval-Augmented Document QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Organizations handling sensitive documents face a critical dilemma: adopt cloud-based AI systems that offer powerful question-answering capabilities but compromise data privacy, or maintain local processing that ensures security but delivers poor accuracy. We present a question-answering system that resolves this trade-off by combining semantic understanding with keyword precision, operating entirely on local infrastructure without internet access. |
Paolo Astrino; | arxiv-cs.CL | 2025-11-13 |
| 713 | SCARE: A Benchmark for SQL Correction and Question Answerability Classification for Reliable EHR Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While prior work has mainly focused on improving SQL generation accuracy or filtering questions before execution, there is a lack of a unified benchmark for evaluating independent post-hoc verification mechanisms (i.e., a component that inspects and validates the generated SQL before execution), which is crucial for safe deployment. To fill this gap, we introduce SCARE, a benchmark for evaluating methods that function as a post-hoc safety layer in EHR QA systems. |
Gyubok Lee; Woosog Chay; Edward Choi; | arxiv-cs.CL | 2025-11-13 |
| 714 | PustakAI: Curriculum-Aligned and Interactive Textbooks Using Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we present the framework PustakAI\footnote{Pustak means `book’ in many Indian languages.} |
Shivam Sharma; Riya Naik; Tejas Gawas; Heramb Patil; Kunal Korgaonkar; | arxiv-cs.CL | 2025-11-13 |
| 715 | Answering Students’ Questions on Course Forums Using Multiple Chain-of-Thought Reasoning and Finetuning RAG-Enabled LLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we experiment fine-tuned LLM with RAG method on the HotpotQA dataset. |
Neo Wang; Sonit Singh; | arxiv-cs.CL | 2025-11-12 |
| 716 | Limitations of Large Language Models in Clinical Problem-solving Arising from Inflexible Reasoning IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To probe potential LLM failure modes in clinical problem-solving, we present the medical abstraction and reasoning corpus (mARC-QA). |
JONATHAN KIM et. al. | Scientific Reports | 2025-11-11 |
| 717 | Testing Question Answering Software with Context-Driven Question Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce CQ^2A, a context-driven question generation approach for testing question-answering systems. |
SHUANG LIU et. al. | arxiv-cs.SE | 2025-11-11 |
| 718 | C-FAITH: A Chinese Fine-Grained Benchmark for Automated Hallucination Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing hallucination benchmarks (especially in Chinese language) rely on human annotations, making automatical and cost-effective hallucination evaluation challenging. To address this, we introduce HaluAgent, an agentic framework that automatically constructs fine-grained question-answering (QA) dataset based on some knowledge documents. |
XU ZHANG et. al. | cikm | 2025-11-10 |
| 719 | FinSage: A Multi-aspect RAG System for Financial Filings Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose the FinSage framework as a solution, utilizing a multi-aspect RAG framework tailored for data retrieval and summarization in multi-modal financial documents. |
XINYU WANG et. al. | cikm | 2025-11-10 |
| 720 | VQA-Induct: Instruction Induction for Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, current approaches for enhancing VQA reasoning performance often assume access to extensive resources such as large annotated datasets, external tools, or numerous demonstrations, which are impractical for real-world users who typically possess only a few demonstrations. We present VQA-Induct, a framework for data-scarce scenarios that leverages MLLMs’ instruction induction capabilities to induce reusable, purely textual task-level instructions from as few as three demonstrations of the same task, then applies these instructions to new instances using only their image-question pairs. |
Po-Chun Chen; Hen-Hsen Huang; Hsin-Hsi Chen; | cikm | 2025-11-10 |
| 721 | Disentangling Complex Questions in LLMs Via Multi-Hop Dependency Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce a novel prompt approach for multi-hop QA viz., MoDeGraph (Multi-Hop Dependency Graphs), that is designed to steer LLMs to extract and model entity relationships in complex questions. |
ROLAND ORUCHE et. al. | cikm | 2025-11-10 |
| 722 | SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SQuAI (https://squai.scads.ai/), a scalable and trustworthy multi-agent retrieval-augmented generation (RAG) framework for scientific question answering (QA) with large language models (LLMs). |
Ines Besrour; Jingbo He; Tobias Schreieder; Michael F\{a}rber; | cikm | 2025-11-10 |
| 723 | A Pivot-Enhanced Question Answering Framework: Using Iterative Sub-Question Decomposition and Answer-to-Question Verification Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Question and Answering(QA) in low-resource languages remains a significant challenge due to the scarcity of high-quality training data. To address this, we propose a robust framework for low-resource QA. |
Seyeon Park; Beakcheol Jang; | cikm | 2025-11-10 |
| 724 | Evaluating Robustness of LLMs in Question Answering on Multilingual Noisy OCR Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we conduct a comprehensive analysis of how OCR-induced noise affects the performance of Multilingual QA Systems. |
Bhawna Piryani; Jamshid Mozafari; Abdelrahman Abdallah; Antoine Doucet; Adam Jatowt; | cikm | 2025-11-10 |
| 725 | PEQQS: A Dataset for Probing Extractive Quantity-focused Question Answering from Scientific Literature Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present how our dataset can be used both for the evaluation of extractive quantity-focused QA from science literature and for exploring the impact of search on the downstream results, specifically focusing on hallucinations resulting from processing non-relevant documents with LLMs. |
Maciej Rybinski; Necva B\{o}l\{u}c\{u}; Huichen Yang; Stephen Wan; | cikm | 2025-11-10 |
| 726 | NLP-QA: A Large-scale Benchmark for Informative Question Answering Over Natural Language Processing Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, scholarly QA development is hindered by the scarcity of large-scale, expertly-annotated datasets, that are needed for modern deep learning models. To address this gap and advance scholarly QA, we introduce NLP-QA, a new dataset of question-answer pairs derived from NLP research documents. |
Avishek Lahiri; Debarshi Kumar Sanyal; Imon Mukherjee; | cikm | 2025-11-10 |
| 727 | Revisiting NLI: Towards Cost-Effective and Human-Aligned Metrics for Evaluating LLMs in Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We re-evaluate a lightweight alternative — off-the-shelf Natural Language Inference (NLI) scoring augmented by a simple lexical-match flag and find that this decades-old technique matches GPT-4o’s accuracy (89.9%) on long-form QA, while requiring orders-of-magnitude fewer parameters. To test human alignment of these metrics rigorously, we introduce DIVER-QA, a new 3000-sample human-annotated benchmark spanning five QA datasets and five candidate LLMs. |
Sai Shridhar Balamurali; Lu Cheng; | arxiv-cs.CL | 2025-11-10 |
| 728 | Structuring Video Semantics with Temporal Triplets for Zero-Shot Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes a structured representation based on temporal triplets to address two major challenges in traditional approaches: temporal fragmentation and entity reference ambiguity. |
LINLIN ZONG et. al. | cikm | 2025-11-10 |
| 729 | Reference-Aligned Retrieval-Augmented Question Answering Over Heterogeneous Proprietary Documents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address these, we propose a RAG-QA framework for internal enterprise use, consisting of: (1) a data pipeline that converts raw multi-modal documents into a structured corpus and QA pairs, (2) a fully on-premise, privacy-preserving architecture, and (3) a lightweight reference matcher that links answer segments to supporting content. |
NAYOUNG CHOI et. al. | cikm | 2025-11-10 |
| 730 | Company-Specific Knowledge Matters: Retrieval-Augmented Generation for Earnings Call Answer Rehearsal Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper explores how to better support corporate executives in answering questions from professional analysts during earnings calls. |
Yung-Yu Shih; Yun-Nung Chen; Chung-Chi Chen; | cikm | 2025-11-10 |
| 731 | QueryBridge: One Million Annotated Questions with SPARQL Queries – Dataset for Question Answering Over Knowledge Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmark datasets (e.g., QALD, LC-QuAD) are limited in size and annotation, hindering QAKG model generalization. To address this, we present QueryBridge, a dataset with over one million annotated questions paired with SPARQL queries. |
Abdelghny Orogat; Ahmed El-Roby; | cikm | 2025-11-10 |
| 732 | BookAsSumQA: An Evaluation Framework for Aspect-Based Book Summarization Via Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: BookAsSumQA automatically generates aspect-specific QA pairsfrom a narrative knowledge graph to evaluate summary quality based on itsquestion-answering performance. Our experiments using BookAsSumQA revealed thatwhile LLM-based approaches showed higher accuracy on shorter texts, RAG-basedmethods become more effective as document length increases, making them moreefficient and practical for aspect-based book summarization. |
Ryuhei Miyazato; Ting-Ruen Wei; Xuyang Wu; Hsin-Tai Wu; Kei Harada; | arxiv-cs.CL | 2025-11-08 |
| 733 | Efficient Semantic Uncertainty Quantification in Language Models Via Diversity-steered Sampling Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a **diversity-steered sampler** that discourages semantically redundant outputs during decoding, covers both autoregressive and masked diffusion paradigms, and yields substantial sample-efficiency gains. |
Ji Won Park; Kyunghyun Cho; | nips | 2025-11-07 |
| 734 | MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles By Reasoning Over Multi-Video Haystacks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarks for video question answering remain limited in scope, typically involving one clip per query, which falls short of representing the challenges of large-scale, audiovisual retrieval and reasoning encountered in practical applications. To bridge this gap, we introduce a novel task named AVHaystacksQA, where the goal is to identify salient segments across different videos in response to a query and link them together to generate the most informative answer. |
SANJOY CHOWDHURY et. al. | nips | 2025-11-07 |
| 735 | Thinker: Learning to Think Fast and Slow Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, this search behavior is often imprecise and lacks confidence, resulting in long, redundant responses and highlighting deficiencies in intuition and verification. Inspired by the Dual Process Theory in psychology, we introduce a simple modification to the QA task that includes four stages: Fast Thinking, where the LLM must answer within a strict token budget; Verification, where the model evaluates its initial response; Slow Thinking, where it refines the initial response with more deliberation; and Summarization, where it distills the refinement from the previous stage into precise steps. |
Stephen Chung; Wenyu Du; Jie Fu; | nips | 2025-11-07 |
| 736 | PT-MoE: An Efficient Finetuning Framework for Integrating Mixture-of-Experts Into Prompt Tuning Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Parameter-efficient fine-tuning (PEFT) methods have shown promise in adapting large language models, yet existing approaches exhibit counter-intuitive phenomena: integrating either matrix decomposition or mixture-of-experts (MoE) individually decreases performance across tasks, though decomposition improves results on specific domains despite reducing parameters, while MoE increases parameter count without corresponding decrease in training efficiency. Motivated by these observations and the modular nature of PT, we propose PT-MoE, a novel framework that integrates matrix decomposition with MoE routing for efficient PT. |
Zongqian Li; Yixuan Su; Nigel Collier; | nips | 2025-11-07 |
| 737 | Cooperative Retrieval-Augmented Generation for Question Answering: Mutual Information Exchange and Ranking By Contrasting Layers Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing RAG methods for simple and multi-hop question answering (QA) are still prone to incorrect retrievals and hallucinations. To address these limitations, we propose CoopRAG, a novel RAG framework for the question answering task in which a retriever and an LLM work cooperatively with each other by exchanging informative knowledge, and the earlier and later layers of the retriever model work cooperatively with each other to accurately rank the retrieved documents relevant to a given query. |
Youmin Ko; Sung Jong Seo; Hyunjoon Kim; | nips | 2025-11-07 |
| 738 | DSAS: A Universal Plug-and-Play Framework for Attention Optimization in Multi-Document Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current solutions either truncate global dependencies or demand costly finetuning, ultimately lacking a universal and simple solution for these challenges. To resolve these limitations, we propose Dual-Stage Adaptive Sharpening (DSAS) containing two modules. |
JIAKAI LI et. al. | nips | 2025-11-07 |
| 739 | QuAnTS: Question Answering on Time Series Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We verify that the large-scale QuAnTS dataset iswell-formed and comprehensive through extensive experiments. |
FELIX DIVO et. al. | arxiv-cs.LG | 2025-11-07 |
| 740 | Attack Via Overfitting: 10-shot Benign Fine-tuning to Jailbreak LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we demonstrate that LLMs can be jailbroken by fine-tuning with only 10 benign QA pairs; our attack exploits the increased sensitivity of LLMs to fine-tuning data after being overfitted. |
Zhixin Xie; Xurui Song; Jun Luo; | nips | 2025-11-07 |
| 741 | Improving Retrieval-Augmented Generation Through Multi-Agent Reinforcement Learning IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Although recent efforts have explored using reinforcement learning (RL) to optimize specific RAG components, these approaches often focus on simple pipelines with only two components or do not adequately address the complex interdependencies and collaborative interactions among the modules. To overcome these limitations, we propose treating the complex RAG pipeline with multiple components as a multi-agent cooperative task, in which each component can be regarded as an RL agent. |
YIQUN CHEN et. al. | nips | 2025-11-07 |
| 742 | IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Vision-language models (VLMs) have demonstrated impressive generalizationacross multimodal tasks, yet most evaluation benchmarks remain Western-centric,leaving open questions about their performance in culturally diverse andmultilingual settings. To address this gap, we introduce IndicVisionBench, thefirst large-scale benchmark centered on the Indian subcontinent. |
ALI FARAZ et. al. | arxiv-cs.CV | 2025-11-06 |
| 743 | Ground-Truth Subgraphs for Better Training and Evaluation of Knowledge Graph Augmented LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We apply SynthKGQA to Wikidata to generate GTSQA, a new datasetdesigned to test zero-shot generalization abilities of KG retrievers withrespect to unseen graph structures and relation types, and benchmark popularsolutions for KG-augmented LLMs on it. |
Alberto Cattaneo; Carlo Luschi; Daniel Justus; | arxiv-cs.LG | 2025-11-06 |
| 744 | BanglaMedQA and BanglaMMedBench: Evaluating Retrieval-Augmented Generation Strategies for Bangla Biomedical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces BanglaMedQA andBanglaMMedBench, the first large-scale Bangla biomedical Multiple ChoiceQuestion (MCQ) datasets designed to evaluate reasoning and retrieval in medicalartificial intelligence (AI). |
Sadia Sultana; Saiyma Sittul Muna; Mosammat Zannatul Samarukh; Ajwad Abrar; Tareque Mohmud Chowdhury; | arxiv-cs.CL | 2025-11-06 |
| 745 | The Illusion of Certainty: Uncertainty Quantification for LLMs Fails Under Ambiguity IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While real-world language is inherentlyambiguous, reflecting aleatoric uncertainty, existing UQ methods are typicallybenchmarked against tasks with no ambiguity. In this work, we demonstrate thatwhile current uncertainty estimators perform well under the restrictiveassumption of no ambiguity, they degrade to close-to-random performance onambiguous data. |
Tim Tomov; Dominik Fuchsgruber; Tom Wollschläger; Stephan Günnemann; | arxiv-cs.LG | 2025-11-06 |
| 746 | Knowledge-Augmented Question Error Correction for Chinese Question Answer System with QuestionRAG Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large language models (LLMs) struggle with this task, frequentlyfailing to interpret user intent (misinterpretation) or unnecessarily alteringthe original question’s structure (over-correction). We propose QuestionRAG, aframework that tackles these problems. |
Longpeng Qiu; Ting Li; Shuai Mao; Nan Yang; Xiaohui Yan; | arxiv-cs.CL | 2025-11-05 |
| 747 | Comparing The Performance of LLMs in RAG-based Question-Answering: A Case Study in Computer Science Literature Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Evaluationmetrics employed in the study include accuracy and precision for binaryquestions and ranking by a human expert, ranking by Google’s AI model Gemini,alongside cosine similarity for long-answer questions. |
Ranul Dayarathne; Uvini Ranaweera; Upeksha Ganegoda; | arxiv-cs.CL | 2025-11-05 |
| 748 | ChiMDQA: Towards Comprehensive Chinese Document QA with Fine-grained Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: With the rapid advancement of natural language processing (NLP) technologies,the demand for high-quality Chinese document question-answering datasets issteadily growing. To address this issue, we present the Chinese Multi-DocumentQuestion Answering Dataset(ChiMDQA), specifically designed for downstreambusiness scenarios across prevalent domains including academic, education,finance, law, medical treatment, and news. |
Jing Gao; Shutiao Luo; Yumeng Liu; Yuanming Li; Hongji Zeng; | arxiv-cs.CL | 2025-11-05 |
| 749 | Understanding New-Knowledge-Induced Factual Hallucinations in LLMs: Analysis, Solution, and Interpretation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Through attention analysis, we find that learning new knowledgereduces the model’s attention to key entities in the question, thus causingexcessive focus on the surrounding context, which may increase the risk ofhallucination. |
Renfei Dang; Peng Hu; Changjiang Gao; Shujian Huang; | arxiv-cs.CL | 2025-11-04 |
| 750 | A Graph-based RAG for Energy Efficiency Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we investigate the use of Large Language Models (LLMs) within agraph-based Retrieval Augmented Generation (RAG) architecture for EnergyEfficiency (EE) Question Answering. |
RICCARDO CAMPI et. al. | arxiv-cs.CL | 2025-11-03 |
| 751 | Benchmarking Geospatial Question Answering with MapQA Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Geospatial question answering (QA) is a fundamental task in navigation and point of interest (POI) searches, yet existing datasets are limited in scale, diversity, and they rely … |
ZEKUN LI et. al. | Proceedings of the 33rd ACM International Conference on … | 2025-11-03 |
| 752 | DEEPAMBIGQA: Ambiguous Multi-hop Questions for Benchmarking LLM Answer Completeness Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing QA benchmarks rarely evaluateboth challenges jointly. To address this, we introduce DeepAmbigQAGen, anautomatic data generation pipeline that constructs QA tasks grounded in textcorpora and linked knowledge graph, generating natural and verifiable questionsthat systematically embed name ambiguity and multi-step reasoning. |
Jiabao Ji; Min Li; Priyanshu Kumar; Shiyu Chang; Saloni Potdar; | arxiv-cs.CL | 2025-11-03 |
| 753 | Pinpointing Trigger Moment for Grounded Video QA: Enhancing Spatio-temporal Grounding in Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this technical report, we introduce a framework to address Grounded VideoQuestion Answering (GVQA) task for the ICCV 2025 Perception Test Challenge. |
Jinhwan Seo; Yoonki Cho; Junhyug Noh; Sung-eui Yoon; | arxiv-cs.CV | 2025-11-03 |
| 754 | Text-VQA Aug: Pipelined Harnessing of Large Multimodal Models for Automated Synthesis Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a pipeline for automated synthesis for text-VQA dataset thatcan produce faithful QA pairs, and which scales up with the availability ofscene text data. |
Soham Joshi; Shwet Kamal Mishra; Viswanath Gopalakrishnan; | arxiv-cs.CV | 2025-11-03 |
| 755 | TrafficNetQA: Question Answering Datasets for Evaluating LLM Performance on Traffic Network Files Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: We propose TrafficNetQA, a benchmark to quantitatively evaluate how well Large Language Models (LLMs) understand transportation-related traffic networks. While there is increasing … |
Donghoon Kwon; Seungmo Kang; Seongjin Choi; | Proceedings of the 33rd ACM International Conference on … | 2025-11-03 |
| 756 | Hybrid Retrieval-Augmented Generation Agent for Trustworthy Legal Question Answering in Judicial Forensics Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: As artificial intelligence permeates judicial forensics, ensuring theveracity and traceability of legal question answering (QA) has become critical.Conventional large language … |
YUEQING XI et. al. | arxiv-cs.AI | 2025-11-03 |
| 757 | When to Trust The Answer: Question-Aligned Semantic Nearest Neighbor Entropy for Safer Surgical VQA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Question Aligned Semantic Nearest NeighborEntropy (QA-SNNE), a black box uncertainty estimator that incorporates questionsemantics into prediction confidence. |
DENNIS PIERANTOZZI et. al. | arxiv-cs.CV | 2025-11-03 |
| 758 | Improving Construction Contract Question Answering Through Embedding Optimization and Semantic Chunking in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Lanqian Zhang; Yan Ning; | Advanced Engineering Informatics | 2025-11-03 |
| 759 | Domain Adaptive Document Reranking for Retrieval Augmented Generation Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Current AI-driven Question-Answering (QA) systems face significant challenges in delivering accurate, domainspecific responses across diverse fields. While Large Language Models … |
Yassine Rachidy; Youssef Hmamouche; Faissal Sehbaoui; A. Seghrouchni; | 2025 IEEE 37th International Conference on Tools with … | 2025-11-03 |
| 760 | CityEQA: A Hierarchical LLM Agent on Embodied Question Answering Benchmark in City Space IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Embodied Question Answering (EQA) has primarily focused on indoor environments, leaving the complexities of urban settings—spanning environment, action, and perception—largely unexplored. To bridge this gap, we introduce CityEQA, a new task where an embodied agent answers open-vocabulary questions through active exploration in dynamic city spaces. |
YONG ZHAO et. al. | emnlp | 2025-11-02 |
| 761 | CAFE: Retrieval Head-based Coarse-to-Fine Information Seeking to Enhance Multi-Document QA Capability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, they still face challenges in balancing retrieval precision and recall, impacting their efficacy in answering questions. To address this, we introduce **CAFE**, a two-stage coarse-to-fine method to enhance multi-document question-answering capacities. |
Han Peng; Jinhao Jiang; Zican Dong; Xin Zhao; Lei Fang; | emnlp | 2025-11-02 |
| 762 | TVQACML: Benchmarking Text-Centric Visual Question Answering in Multilingual Chinese Minority Languages Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, the open-source nature of these benchmarks and the broad sources of training data for MLLMs have inevitably led to benchmark contamination, resulting in unreliable evaluation results. To alleviate this issue, we propose a contamination-free and more challenging TEC-VQA benchmark called Text-Centric Visual Question Answering in Multilingual Chinese Minority Languages(TVQACML), which involves eight languages, including Standard Chinese, Korean, and six minority languages. |
Sha Jiu; Yu Weng; Mengxiao Zhu; Chong Feng; Zheng Liu; | emnlp | 2025-11-02 |
| 763 | Static or Dynamic: Towards Query-Adaptive Token Selection for Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a novel token selection strategy, explore-then-select, that adaptively adjusts static and dynamic information based on question requirements. |
Yumeng Shi; Quanyu Long; Wenya Wang; | emnlp | 2025-11-02 |
| 764 | ComplexTempQA: A 100m Dataset for Complex Temporal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ComplexTempQA,a large-scale dataset consisting of over 100 million question-answer pairs designed to tackle the challenges in temporal question answering. |
Raphael Gruber; Abdelrahman Abdallah; Michael Färber; Adam Jatowt; | emnlp | 2025-11-02 |
| 765 | PrismRAG: Boosting RAG Factuality with Distractor Resilience and Strategized Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose an efficient fine-tuning framework, called PrismRAG, that (i) trains the model with distractor-aware QA pairs mixing gold evidence with subtle distractor passages, and (ii) instills reasoning-centric habits that make the LLM plan, rationalize, and synthesize without relying on extensive human engineered instructions. |
MOHAMMAD KACHUEE et. al. | emnlp | 2025-11-02 |
| 766 | Trustworthy Medical Question Answering: An Evaluation-Centric Survey Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this survey, we systematically examine six key dimensions of trustworthiness in medical QA, i. e. , Factuality, Robustness, Fairness, Safety, Explainability, and Calibration. |
YINUO WANG et. al. | emnlp | 2025-11-02 |
| 767 | Large Language Models Meet Knowledge Graphs for Question Answering: Synthesis and Opportunities IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: In this survey, we propose a new structured taxonomy that categorizes the methodology of synthesizing LLMs and KGs for QA according to the categories of QA and the KG’s role when integrating with LLMs. |
Chuangtao Ma; Yongrui Chen; Tianxing Wu; Arijit Khan; Haofen Wang; | emnlp | 2025-11-02 |
| 768 | ESGenius: Benchmarking LLMs on Environmental, Social, and Governance (ESG) and Sustainability Knowledge IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce ESGenius, a comprehensive benchmark for evaluating and enhancing the proficiency of Large Language Models (LLMs) in Environmental, Social, and Governance (ESG) and sustainability-focused question answering. |
CHAOYUE HE et. al. | emnlp | 2025-11-02 |
| 769 | LiteraryQA: Towards Effective Evaluation of Long-document Narrative QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce LiteraryQA, a high-quality subset of NarrativeQA focused on literary works. |
Tommaso Bonomo; Luca Gioffré; Roberto Navigli; | emnlp | 2025-11-02 |
| 770 | STREAQ: Selective Tiered Routing for Effective and Affordable Contact Center Quality Assurance Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce STREAQ, a two-tier selective routing framework to intelligently route queries between cost-efficient and high-capability models. |
Prajwal Sood; Rajdeep Agrawal; Mayank Sati; Digvijay Anil Ingle; Cijo George; | emnlp | 2025-11-02 |
| 771 | TALON: A Multi-Agent Framework for Long-Table Exploration and Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose TALON, a multi-agent framework designed for question answering over long tables. |
RUOCHUN JIN et. al. | emnlp | 2025-11-02 |
| 772 | MEBench: Benchmarking Large Language Models for Cross-Document Multi-Entity Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Although Large Language Models (LLMs) and Retrieval-augmented Generation (RAG) systems show promise, their performance on cross-document MEQA remains underexplored due to the absence of tailored benchmarks. To address this gap, we introduce MEBench, a scalable multi-document, multi-entity benchmark designed to systematically evaluate LLMs’ capacity to retrieve, consolidate, and reason over scattered and dense information. |
TENG LIN et. al. | emnlp | 2025-11-02 |
| 773 | RTQA : Recursive Thinking for Complex Temporal Knowledge Graph Question Answering with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Current temporal knowledge graph question answering (TKGQA) methods primarily focus on implicit temporal constraints, lacking the capability to handle more complex temporal queries, and struggle with limited reasoning abilities and error propagation in decomposition frameworks. We propose RTQA, a novel framework to address these challenges by enhancing reasoning over TKGs without requiring training. |
ZHAOYAN GONG et. al. | emnlp | 2025-11-02 |
| 774 | Memory-QA: Answering Recall Questions Based on Multimodal Memories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This task poses unique challenges, including the creation of task-oriented memories, the effective utilization of temporal and location information within memories, and the ability to draw upon multiple memories to answer a recall question. To address these challenges, we propose a comprehensive pipeline, Pensieve, integrating memory-specific augmentation, time- and location-aware multi-signal retrieval, and multi-memory QA fine-tuning. |
HONGDA JIANG et. al. | emnlp | 2025-11-02 |
| 775 | FacLens: Transferable Probe for Foreseeing Non-Factuality in Fact-Seeking Question Answering of Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose a lightweight model named Factuality Lens (FacLens), which effectively probes hidden representations of fact-seeking questions for the NFP task. |
YANLING WANG et. al. | emnlp | 2025-11-02 |
| 776 | ProtoVQA: An Adaptable Prototypical Framework for Explainable Fine-Grained Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present ProtoVQA, a unified prototypical framework that (i) learns question-aware prototypes that serve as reasoning anchors, connecting answers to discriminative image regions, (ii) applies spatially constrained matching to ensure that the selected evidence is coherent and semantically relevant, and (iii) supports both answering and grounding tasks through a shared prototype backbone. |
XINGJIAN DIAO et. al. | emnlp | 2025-11-02 |
| 777 | CompKBQA: Component-wise Task Decomposition for Knowledge Base Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the challenge of generating error-free logical forms remains, as skeleton, topic Entity, and relation Errors still frequently occur. To address these challenges, we propose CompKBQA(Component-wise Task Decomposition for Knowledge Base Question Answering), a novel framework that optimizes the process of fine-tuning a LLM for generating logical forms by enabling the LLM to progressively learn relevant sub-tasks like skeleton generation, topic entity generation, and relevant relations generation. |
YUHANG TIAN et. al. | emnlp | 2025-11-02 |
| 778 | BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce BYOKG-RAG, a framework that enhances KGQA by synergistically combining LLMs with specialized graph retrieval tools. |
COSTAS MAVROMATIS et. al. | emnlp | 2025-11-02 |
| 779 | SportReason: Evaluating Retrieval-Augmented Reasoning Across Tables and Text for Sports Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SportReason, a benchmark for retrieval-augmented reasoning on numerical sports questions. |
Kaiyue Feng; Siyue Zhang; Bingsen Chen; Yilun Zhao; Chen Zhao; | emnlp | 2025-11-02 |
| 780 | XLQA: A Benchmark for Locale-Aware Multilingual Open-Domain Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This assumption neglects the cultural and regional variations that affect question understanding and answer, leading to biased evaluation in multilingual benchmarks. To address these limitations, we introduce XLQA, a novel benchmark explicitly designed for locale-sensitive multilingual ODQA. |
Keonwoo Roh; Yeong-Joon Ju; Seong-Whan Lee; | emnlp | 2025-11-02 |
| 781 | T2: An Adaptive Test-Time Scaling Strategy for Contextual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: But they add human bias to the reasoning process and fail to leverage models’ inherent reasoning capabilities. To address these limitations, we present T2: Think-to-Think, a novel framework that dynamically adapts reasoning depth based on question complexity. |
ZHENGYI ZHAO et. al. | emnlp | 2025-11-02 |
| 782 | UNCLE: Benchmarking Uncertainty Expressions in Long-Form Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Along with UNCLE, we propose a suite of new metrics to assess the models’ capabilities to selectively express uncertainty. |
RUIHAN YANG et. al. | emnlp | 2025-11-02 |
| 783 | RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the lack of publicly available RAG-centric preference datasets and specialised RMs, we introduce RAGferee, a methodology that repurposes question-answering (QA) datasets into preference pairs that prioritise groundedness over stylistic features, enabling the training of contextual RMs better suited to judging RAG responses. |
ANDREI CATALIN COMAN et. al. | emnlp | 2025-11-02 |
| 784 | CoCoA: Confidence- and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce CoCoA (Confidence- and Context-Aware Adaptive Decoding), a novel token-level algorithm for principled conflict resolution and enhanced faithfulness. |
Anant Khandelwal; Manish Gupta; Puneet Agrawal; | emnlp | 2025-11-02 |
| 785 | Weaver: Interweaving SQL and LLM for Table Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing approaches that combine SQL and LLM typically rely on rigid, predefined workflows, limiting their adaptability to complex queries. To address these issues, we introduce Weaver, a modular pipeline that dynamically integrates SQL and LLM for table-based question answering (Table QA). |
Rohit Khoja; Devanshu Gupta; Yanjie Fu; Dan Roth; Vivek Gupta; | emnlp | 2025-11-02 |
| 786 | Don’t Forget The Base Retriever! A Low-Resource Graph-based Retriever for Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose GR\small{IEVER}, a lightweight, low-resource, multi-step graph-based retriever for multi-hop QA. |
ANDRE MELO et. al. | emnlp | 2025-11-02 |
| 787 | Answering Narrative-Driven Recommendation Queries Via A Retrieve–Rank Paradigm and The OCG-Agent Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This work formally introduces narrative recommendation as a distinct task and contends that the RAG paradigm is inherently ill-suited for it, owing to information loss in LLMs when retrieving information from from multiple long and fragmented contexts, and limitations in ranking effectiveness. |
YUNXIAO SHI et. al. | emnlp | 2025-11-02 |
| 788 | PakBBQ: A Culturally Adapted Bias Benchmark for QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most LLMs are trained and evaluated on Western centric data, with little attention paid to low-resource languages and regional contexts. To address this gap, we introduce PakBBQ, a culturally and regionally adapted extension of the original Bias Benchmark for Question Answering (BBQ) dataset. |
Abdullah Hashmat; Muhammad Arham Mirza; Agha Ali Raza; | emnlp | 2025-11-02 |
| 789 | LaMP-QA: A Benchmark for Personalized Long-form Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This is mainly due to lack of resources for training and evaluating personalized question answering systems. We address this gap by introducing LaMP-QA—a benchmark designed for evaluating personalized long-form answer generation. |
Alireza Salemi; Hamed Zamani; | emnlp | 2025-11-02 |
| 790 | Let’s Reason Formally: Natural-Formal Hybrid Reasoning Enhances LLM’s Math Capability Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, this integration is challenging due to inherent disparities in problem structure and reasoning format between NL and FL. To address these challenges, we introduce **NL-FL HybridReasoning (NFL-HR)**, an end-to-end framework designed to incorporate the FL expert into NL math problem-solving. |
Ruida Wang; Yuxin Li; Yi R. Fung; Tong Zhang; | emnlp | 2025-11-02 |
| 791 | Confidence-guided Refinement Reasoning for Zero-shot Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Confidence-guided Refinement Reasoning (C2R), a novel training-free framework applicable to question-answering (QA) tasks across text, image, and video domains. |
Youwon Jang; Woo Suk Choi; Minjoon Jung; Minsu Lee; Byoung-Tak Zhang; | emnlp | 2025-11-02 |
| 792 | Faster In-Context Learning for LLMs Via N-Gram Trie Speculative Decoding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the lengthy retrieved contexts and limited token throughput in autoregressive models significantly constrain reasoning speed. To address this challenge, we propose N-Gram Trie Speculative Decoding, a novel approach that leverages the overlap between context and model output. |
JINGLIN CHEN et. al. | emnlp | 2025-11-02 |
| 793 | ReAgent: Reversible Multi-Agent Reasoning for Knowledge-Enhanced Multi-Hop QA IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While large language models (LLMs) have achieved substantial improvements via chain-of-thought (CoT) prompting and retrieval-augmented generation, these methods typically adopt a forward-only workflow—early mistakes persist throughout inference, and contradictions discovered later cannot systematically trigger re-evaluation. To address this limitation, we present ReAgent, a reversible multi-agent reasoning framework. |
ZHAO XINJIE et. al. | emnlp | 2025-11-02 |
| 794 | RAVEN: Query-Guided Representation Alignment for Question Answering Over Audio, Video, Embedded Sensors, and Natural Language Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: We present RAVEN, a unified QA architecture whose core is QuART, a query-conditioned cross-modal gating module that assigns scalar relevance scores to each token across modalities, enabling the model to amplify informative signals and suppress distractors before fusion. |
Subrata Biswas; Mohammad Nur Hossain Khan; Bashima Islam; | emnlp | 2025-11-02 |
| 795 | Truth, Trust, and Trouble: Medical AI on The Edge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a rigorous benchmarking framework via a dataset of over 1,000 health questions. |
MOHAMMAD ANAS AZEEZ et. al. | emnlp | 2025-11-02 |
| 796 | Discrepancy Detection at The Data Level: Toward Consistent Multilingual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MIND, a user-in-the-loop fact-checking pipeline to detect factual and cultural discrepancies in multilingual QA knowledge bases. |
LORENA CALVO-BARTOLOMÉ et. al. | emnlp | 2025-11-02 |
| 797 | Retrieving Support to Rank Answers in Open-Domain Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a novel Question Answering (QA) architecture that enhances answer selection by retrieving targeted supporting evidence. |
Zeyu Zhang; Alessandro Moschitti; Thuy Vu; | emnlp | 2025-11-02 |
| 798 | Tagging-Augmented Generation: Assisting Language Models in Finding Intricate Knowledge In Long Contexts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose Tagging-Augmented Generation (TAG), a lightweight data augmentation strategy that boosts LLM performance in long-context scenarios, without degrading and altering the integrity and composition of retrieved documents. |
ANWESAN PAL et. al. | emnlp | 2025-11-02 |
| 799 | What Are Foundation Models Cooking in The Post-Soviet World? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we investigate the Post-Soviet cultural food knowledge of foundation models by constructing BORSch, a multi-modal dataset encompassing 1147 and 823 dishes in the Russian and Ukrainian languages, centered around the Post-Soviet region. |
Anton Lavrouk; Tarek Naous; Alan Ritter; Wei Xu; | emnlp | 2025-11-02 |
| 800 | SilVar: Speech-Driven Multimodal Model for Reasoning Visual Question Answering and Object Localization Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Moreover, the quality of language models primarily depends on reasoning and prompting techniques, such as chain-of-thought, which remain underexplored when using speech instructions. To address these challenges, we propose SilVar, an end-to-end multimodal model that leverages speech instructions for reasoning-based visual question answering. |
Tan-Hanh Pham; Le Hoang Nam; Phu-Vinh Nguyen; Chris Ngo; Truong-Son Hy; | emnlp | 2025-11-02 |
| 801 | How Accurate Are LLMs at Multi-Question Answering on Conversational Transcripts? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore the capabilities of LLMs to answer multiple questions based on the same conversational context. |
Xiliang Zhu; Shi Zong; David Rossouw; | emnlp | 2025-11-02 |
| 802 | CondAmbigQA: A Benchmark and Dataset for Conditional Ambiguous Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Therefore, identifying possible implicit assumptions is crucial in QA. To address this fundamental challenge, we propose Conditional Ambiguous Question-Answering (CondAmbigQA), a benchmark comprising 2,000 ambiguous queries and condition-aware evaluation metrics. |
Zongxi Li; Yang Li; Haoran Xie; S. Joe Qin; | emnlp | 2025-11-02 |
| 803 | FinRAGBench-V: A Benchmark for Multimodal RAG with Visual Citation in The Financial Domain IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, most existing RAG research in finance focuses predominantly on textual data, overlooking the rich visual content in financial documents, resulting in the loss of key analytical insights. To bridge this gap, we present FinRAGBench-V, a comprehensive visual RAG benchmark tailored for finance. |
Suifeng Zhao; Zhuoran Jin; Sujian Li; Jun Gao; | emnlp | 2025-11-02 |
| 804 | KoBLEX: Open Legal Question Answering with Multi-hop Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, these benchmarks fail to evaluate open-ended and provision-grounded Question Answering (QA). To address this, we introduce a Korean Benchmark for Legal EXplainable QA (KoBLEX), designed to evaluate provision-grounded, multi-hop legal reasoning. |
Jihyung Lee; Daehui Kim; Seonjeong Hwang; Hyounghun Kim; Gary Lee; | emnlp | 2025-11-02 |
| 805 | NitiBench: Benchmarking LLM Frameworks on Thai Legal Question Answering Capabilities Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce NitiBench, a novel benchmark featuring two datasets: (1) NitiBench-CCL, covering Thai financial laws, and (2) NitiBench-Tax, containing Thailand’s official tax rulings. |
PAWITSAPAK AKARAJARADWONG et. al. | emnlp | 2025-11-02 |
| 806 | StepER: Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing knowledge distillation methods overlook the need for different reasoning abilities at different steps, hindering transfer in multi-step retrieval-augmented frameworks. To address this, we propose Step-wise Knowledge Distillation for Enhancing Reasoning Ability in Multi-Step Retrieval-Augmented Language Models (StepER). |
Kyumin Lee; Minjin Jeon; Sanghwan Jang; Hwanjo Yu; | emnlp | 2025-11-02 |
| 807 | Generating Spatial Knowledge Graphs from Automotive Diagrams for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluate three distinct generation pipelines (Per-Attribute, Per-Component, and a Single-Shot baseline) to create the SKG using Large Vision-Language Models (LVLMs). |
Steve Bakos; Chen Xing; Heidar Davoudi; Aijun An; Ron DiCarlantonio; | emnlp | 2025-11-02 |
| 808 | StepSearch: Igniting LLMs Search Ability Via Step-Wise Proximal Policy Optimization IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Previous work has explored reinforcement learning (RL) to train LLMs to perform search-based document retrieval, achieving notable improvements in QA performance, but underperform on complex, multi-hop QA resulting from the sparse rewards from global signal only. To address this gap in existing research, we introduce StepSearch, a framework for search LLMs that trained with step-wise proximal policy optimization method. |
Xuhui Zheng; Kang An; Ziliang Wang; Yuhang Wang; Yichao Wu; | emnlp | 2025-11-02 |
| 809 | TreeRare: Syntax Tree-Guided Retrieval and Reasoning for Knowledge-Intensive Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the performance of such retrieval frameworks is limited by the accumulation of reasoning errors and misaligned retrieval results. To overcome these limitations, we propose TreeRare (Syntax Tree-Guided Retrieval and Reasoning, a framework that utilizes syntax trees to guide information retrieval and reasoning for question answering. |
Boyi Zhang; Zhuo Liu; Hangfeng He; | emnlp | 2025-11-02 |
| 810 | Will It Still Be True Tomorrow? Multilingual Evergreen Question Classification to Improve Trustworthy QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce EverGreenQA, the first multilingual QA dataset with evergreen labels, supporting both evaluation and training. |
SERGEY PLETENEV et. al. | emnlp | 2025-11-02 |
| 811 | FLARE: Faithful Logic-Aided Reasoning and Exploration Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Faithful Logic-Aided Reasoning and Exploration (FLARE), which uses LLMs to plan solutions, formalize queries into logic programs, and simulate code execution through multi-hop search without external solvers. |
Erik Arakelyan; Pasquale Minervini; Patrick Lewis; Pat Verga; Isabelle Augenstein; | emnlp | 2025-11-02 |
| 812 | Audio Query Handling System with Integrated Expert Models and Contextual Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents an audio chatbot system designed to handle a wide range of audio-related queries by integrating multiple specialized audio processing models. |
Naveen Vakada; Arvind Krishna Sridhar; Yinyi Guo; Erik Visser; | emnlp | 2025-11-02 |
| 813 | TempQA: An LLM-based Framework for Temporal Knowledge Graph Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Qianyi Hu; Xinhui Tu; Ao Li; Biao Yao; | Knowl. Based Syst. | 2025-11-01 |
| 814 | Enhancing Long-form Question Answering Via Reflection with Question Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Junjie Xiao; Wei Wu; Jiaxu Zhao; Meng Fang; Jianxin Wang; | Inf. Process. Manag. | 2025-11-01 |
| 815 | DairyGoatQA: A Knowledge Graph Enhanced Large Language Models Approach for Question Answering in The Dairy Goat Domain Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Xiaojin Chen; Sai Zhang; Xinxing Li; Ruiqin Ma; Qinan Zhao; | Expert Syst. Appl. | 2025-11-01 |
| 816 | MDIF: A Multimodal Dynamic Inference Framework for Traffic Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Wenhao Guo; Lingling Zi; Xin Cong; | Applied Intelligence | 2025-11-01 |
| 817 | VinDr-CXR-VQA: A Visual Question Answering Dataset for Explainable Chest X-Ray Analysis with Multi-Task Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present VinDr-CXR-VQA, a large-scale chest X-ray dataset for explainableMedical Visual Question Answering (Med-VQA) with spatial grounding. |
Hai-Dang Nguyen; Ha-Hieu Pham; Hao T. Nguyen; Huy-Hieu Pham; | arxiv-cs.CV | 2025-11-01 |
| 818 | ICSThreatQA: A Knowledge-graph Enhanced Question Answering Model for Industrial Control System Threat Intelligence Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
R. Rani; Mahender Kumar; Gregory Epiphaniou; Carsten Maple; | Expert Syst. Appl. | 2025-11-01 |
| 819 | A Systematic Evaluation of Large Language Models and Retrieval-Augmented Generation for The Task of Kazakh Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: This paper presents a systematic evaluation of large language models (LLMs) and retrieval-augmented generation (RAG) approaches for question answering (QA) in the low-resource … |
Aigerim Mansurova; A. Tleubayeva; A. Nugumanova; A. Shomanov; Sadi Evren Şeker; | Inf. | 2025-10-30 |
| 820 | FARSIQA: Faithful and Advanced RAG System for Islamic Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing Retrieval-Augmented Generation (RAG)systems, relying on simplistic single-pass pipelines, fall short on complex,multi-hop queries requiring multi-step reasoning and evidence aggregation. Toaddress this gap, we introduce FARSIQA, a novel, end-to-end system for FaithfulAdvanced Question Answering in the Persian Islamic domain. |
Mohammad Aghajani Asl; Behrooz Minaei Bidgoli; | arxiv-cs.CL | 2025-10-29 |
| 821 | Beyond Long Context: When Semantics Matter More Than Tokens Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The Clinical Entity Augmented Retrieval (CLEAR)method, introduced by Lopez et al. 2025, uses entity aware retrieval andachieved improved performance with an F1 score of 0.90 versus 0.86 forembedding based retrieval, while using over 70 percent fewer tokens. |
Tarun Kumar Chawdhury; Jon D. Duke; | arxiv-cs.CL | 2025-10-29 |
| 822 | A Multimodal and Dynamically Updatable Benchmark for Aviation Question Answering with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper proposes a multimodal, multi-level benchmark dataset tailored to aviation QA tasks, alongside an automated updating mechanism and a multi-dimensional evaluation framework. |
Liu He; Shuyan Liu; Xiaorui Qin; Ran An; Jianghui Zeng; | International Journal of Robotics and Automation Technology | 2025-10-29 |
| 823 | Adapting Small Language Models to Low-Resource Domains: A Case Study in Hindi Tourism QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present a multi-stagefinetuning strategy to adapt lightweight language models to the Hindi tourismdomain by leveraging both original and synthetic training data. |
Sandipan Majhi; Paheli Bhattacharya; | arxiv-cs.CL | 2025-10-29 |
| 824 | A Knowledge Graph Enhancement Technique for HIPAA Compliant Health Question Answering in Personal Health Libraries Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In contrast, Knowledge Graphs (KGs) have proven effective for QA tasks, especially in extracting structured insights from text, but transforming free text into KGs often leads to information or context loss that can compromise answer accuracy. To overcome this challenge, we present a novel iterative and monotonic KG refinement technique that enriches knowledge representation without sacrificing contextual integrity. |
Hasan Jamil; | ACM Transactions on Computing for Healthcare | 2025-10-28 |
| 825 | BMGQ: A Bottom-up Method for Generating Complex Multi-hop Reasoning Questions from Semi-structured Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Meanwhile, manually curating non-trivially retrievable questions — whereanswers cannot be found through a single direct query but instead requiremulti-hop reasoning over oblique and loosely connected evidence — incursprohibitive human costs and fails to scale, creating a critical data bottleneckfor training high-capability retrieval-and-reasoning agents. To address this, we present an automated framework for generatinghigh-difficulty, training-ready multi-hop questions from semi-structuredknowledge sources. |
BINGSEN QIU et. al. | arxiv-cs.AI | 2025-10-28 |
| 826 | Emotion-Qwen-VL: A Fully Fine-Tuned Multimodal Large Language Model for Micro-Expression Visual Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: This paper presents our solution for the Micro-Expression Visual Question Answering (ME-VQA) task in the 2025 Facial Micro-Expression Grand Challenge (MEGC). To address the … |
YUJING WANG et. al. | Proceedings of the 33rd ACM International Conference on … | 2025-10-27 |
| 827 | Towards Complex Table Question Answering Over Tabular Data Lakes (Extended Version) Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we systematically analyze how LLMs paired with table retrievers can answer queries over private tabular data lakes. |
Daniela Risis; Jan-Micha Bodensohn; Matthias Urban; Carsten Binnig; | Datenbank-Spektrum | 2025-10-27 |
| 828 | MMESGBench: Pioneering Multimodal Understanding and Complex Reasoning Benchmark for ESG Tasks IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To fill the gap, we introduce MMESGBench, a first-of-its-kind benchmark dataset targeted to evaluate multimodal understanding and reasoning across multi-source ESG documents. |
LEI ZHANG et. al. | mm | 2025-10-27 |
| 829 | FineQuest: Adaptive Knowledge-Assisted Sports Video Understanding Via Agent-of-Thoughts Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose FineQuest, the first training-free framework that leverages dual-mode reasoning inspired by cognitive science: i) Reactive Reasoning for straightforward sports queries and ii) Deliberative Reasoning for more complex ones. |
Haodong Chen; Haojian Huang; Xinxiang Yin; Dian Shao; | mm | 2025-10-27 |
| 830 | Valor32k-AVQA V2.0: Open-Ended Audio-Visual Question Answering Dataset and Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despite growing interest in Audio-Visual Question Answering (AVQA), existing datasets often suffer from limited diversity, rigid formats, and insufficient integration of audio and visual modalities. To address these limitations, we introduce Valor32k-AVQA v2.0, a large-scale dataset containing 28,863 real-world videos and over 225,000 QA pairs, designed to support diverse and realistic multimodal understanding. |
Ines Riahi; Abduljalil Radman; Zixin Guo; Rachid Hedjam; Jorma Laaksonen; | mm | 2025-10-27 |
| 831 | VQA2: Visual Question Answering for Video Quality Assessment IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Nevertheless, related work has not been explored in the video domain, leaving substantial room for improvement. To address this gap, we introduce the VQA² Instruction Dataset-the first visual question answering instruction dataset that focuses on video quality assessment. |
ZIHENG JIA et. al. | mm | 2025-10-27 |
| 832 | RSVLM-QA: A Benchmark Dataset for Remote Sensing Vision Language Model-based Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces Remote Sensing Vision Language Model Question Answering (RSVLM-QA) dataset, a new large-scale, content-rich VQA dataset for the RS domain. |
XING ZI et. al. | mm | 2025-10-27 |
| 833 | HierMEQA: A Relationship-Aware Hierarchical Framework for Consistent Micro-Expression Visual Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: The rise of Multimodal Large Language Models (MLLMs) offers new opportunities for Micro-Expression (ME) analysis. This paper introduces Micro-Expression Visual Question Answering … |
LINGSI ZHU et. al. | Proceedings of the 33rd ACM International Conference on … | 2025-10-27 |
| 834 | LMM4Edit: Benchmarking and Evaluating Multimodal Image Editing with LMMs IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end, we introduce EBench-18K, the first large-scale image Editing Benchmark including 18K edited images with fine-grained human preference annotations for evaluating TIE. |
ZITONG XU et. al. | mm | 2025-10-27 |
| 835 | MM-GRADE: A Multi-Modal EDA Tool Documentation QA Framework Leveraging Retrieval Augmented Generation Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: The complexity of EDA tools necessitates the development of advanced documentation query answering systems to enhance user efficiency and reduce the associated learning curve. … |
YUAN PU et. al. | 2025 IEEE/ACM International Conference On Computer Aided … | 2025-10-26 |
| 836 | Hierarchical Sequence Iteration for Heterogeneous Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Retrieval-augmented generation (RAG) remains brittle on multi-step questionsand heterogeneous evidence sources, trading accuracy against latency andtoken/tool budgets. This paper … |
Ruiyi Yang; Hao Xue; Imran Razzak; Hakim Hacid; Flora D. Salim; | arxiv-cs.CL | 2025-10-23 |
| 837 | VLSP 2025 MLQA-TSR Challenge: Vietnamese Multimodal Legal Question Answering on Traffic Sign Regulation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents the VLSP 2025 MLQA-TSR – the multimodal legal questionanswering on traffic sign regulation shared task at VLSP 2025. |
SON T. LUU et. al. | arxiv-cs.CL | 2025-10-23 |
| 838 | Task-guided Dynamic Visual Reasoning for Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we propose a task-guided dynamic visual reasoning method for visual question answering, which models the spatiotemporal states of objects in dynamic scenes, decomposes the questions into task steps, and finally deduces reasoning on the established spatiotemporal dynamic scene graph neural network. |
Yao Cong; Hongwei Mo; | International Journal of Humanoid Robotics | 2025-10-23 |
| 839 | Multimedia-Aware Question Answering: A Review of Retrieval and Cross-Modal Reasoning Architectures Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this survey, we review recent advancements in QAsystems that integrate multimedia retrieval pipelines, focusing onarchitectures that align vision, language, and audio modalities with userqueries. |
Rahul Raja; Arpita Vats; | arxiv-cs.IR | 2025-10-23 |
| 840 | GlobalRAG: Enhancing Global Reasoning in Multi-hop Question Answering Via Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose GlobalRAG, a reinforcementlearning framework designed to enhance global reasoning in multi-hop QA.GlobalRAG decomposes questions into subgoals, coordinates retrieval withreasoning, and refines evidence iteratively. |
JINCHANG LUO et. al. | arxiv-cs.CL | 2025-10-23 |
| 841 | Bridging Language Gaps with Adaptive RAG: Improving Indonesian Language Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To overcome the limitedavailability of Indonesian language dataset, our study employs machinetranslation as data augmentation approach. |
William Christian; Daniel Adamlu; Adrian Yu; Derwin Suhartono; | arxiv-cs.CL | 2025-10-23 |
| 842 | Investigating LLM Capabilities on Long Context Comprehension for Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We shed light into some of the evaluation aspects using amulti-faceted approach. |
Feras AlMannaa; Talia Tseriotou; Jenny Chim; Maria Liakata; | arxiv-cs.CL | 2025-10-21 |
| 843 | IMB: An Italian Medical Benchmark for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present twocomprehensive Italian medical benchmarks: \textbf{IMB-QA}, containing 782,644patient-doctor conversations from 77 medical categories, and \textbf{IMB-MCQA},comprising 25,862 multiple-choice questions from medical specialtyexaminations. |
Antonio Romano; Giuseppe Riccio; Mariano Barone; Marco Postiglione; Vincenzo Moscato; | arxiv-cs.CL | 2025-10-21 |
| 844 | Interpretable Question Answering with Knowledge Graphs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents a question answering system that operates exclusively ona knowledge graph retrieval without relying on retrieval augmented generation(RAG) with large language models (LLMs). |
Kartikeya Aneja; Manasvi Srivastava; Subhayan Das; Nagender Aneja; | arxiv-cs.CL | 2025-10-21 |
| 845 | From Retrieval to Generation: Unifying External and Parametric Knowledge for Medical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Both issues can mislead reasoning and undermineanswer reliability. To address these challenges, we propose MedRGAG, a unifiedretrieval-generation augmented framework that seamlessly integrates externaland parametric knowledge for medical QA. |
Lei Li; Xiao Zhou; Yingying Zhang; Xian Wu; | arxiv-cs.CL | 2025-10-21 |
| 846 | That’s Deprecated! Understanding, Detecting, and Steering Knowledge Conflicts in Language Models for Code Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Building on prior question-answering (QA)research, we extend the investigation of knowledge conflicts to the realm ofcode generation. We propose a domain-agnostic framework for constructing andinterpreting such conflicts, along with a novel evaluation method and datasettailored to code conflict scenarios. |
JAESUNG BAE et. al. | arxiv-cs.CL | 2025-10-21 |
| 847 | Robust Driving QA Through Metadata-Grounded Context and Task-Specific Prompts Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a two-phase vision-language QA system for autonomous driving thatanswers high-level perception, prediction, and planning questions. |
Seungjun Yu; Junsung Park; Youngsun Lim; Hyunjung Shim; | arxiv-cs.CV | 2025-10-21 |
| 848 | Overcoming Dual Drift for Continual Long-Tailed Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we introduce Continual Long-Tailed Visual Question Answering (CLT-VQA) and identify two critical challenges: inner-task prototype drift, where classifier prototypes become biased toward majority classes due to imbalanced data, and inter-task feature drift, where learned features shift over time, causing forgetting of previously learned knowledge. To address these challenges, we propose a unified dual-balance approach that integrates a Balanced Classifier Prototype (BCP) learning module and a Multi-modal Feature Alignment (MFA) module. |
Feifei Zhang; Zhihao Wang; Xi Zhang; Changsheng Xu; | iccv | 2025-10-20 |
| 849 | AVAM: A Universal Training-free Adaptive Visual Anchoring Embedded Into Multimodal Large Language Model for Multi-image Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose a straightforward yet universal Adaptive Visual Anchoring strategy, which can be seamlessly integrated into existing MLLMs, offering significant accuracy improvements through adaptive compression. |
Kang Zeng; Guojin Zhong; Jintao Cheng; Jin Yuan; Zhiyong Li; | iccv | 2025-10-20 |
| 850 | HIS-GPT: Towards 3D Human-In-Scene Multimodal Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a new task to benchmark human-in-scene understanding for embodied agents: Human-In-Scene Question Answering (HIS-QA). |
Jiahe Zhao; Ruibing Hou; Zejie Tian; Hong Chang; Shiguang Shan; | iccv | 2025-10-20 |
| 851 | Ask and Remember: A Questions-Only Replay Strategy for Continual Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present QUestion-only replay with Attention Distillation (QUAD), a novel approach for VQACL that leverages only past task questions for regularization. |
Imad Eddine Marouf; Enzo Tartaglione; Stéphane Lathuilière; Joost Van De Weijer; | iccv | 2025-10-20 |
| 852 | Passing The Driving Knowledge Test IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present DriveQA, an extensive open-source text and vision-based benchmark that exhaustively covers traffic regulations and scenarios. |
Maolin Wei; Wanzhou Liu; Eshed Ohn-Bar; | iccv | 2025-10-20 |
| 853 | GraspCoT: Integrating Physical Property Reasoning for 6-DoF Grasping Under Flexible Language Instructions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose GraspCoT, a 6-DoF grasp detection framework that integrates a Chain-of-Thought (CoT) reasoning mechanism oriented to physical properties, guided by auxiliary question-answering (QA) tasks. |
XIAOMENG CHU et. al. | iccv | 2025-10-20 |
| 854 | Beyond The Destination: A Novel Benchmark for Exploration-Aware Embodied Question Answering IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To improve exploration efficiency, we propose Fine-EQA, a hybrid exploration model that integrates frontier-based and goal-oriented navigation to guide agents toward task-relevant regions more effectively. |
KAIXUAN JIANG et. al. | iccv | 2025-10-20 |
| 855 | ETVA: Evaluation of Text-to-Video Alignment Via Fine-grained Question Generation and Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing text-to-video alignment metrics like CLIPScore only generate coarse-grained scores without fine-grained alignment details, failing to align with human preference. To address this limitation, we propose ETVA, a novel Evaluation method of Text-to-Video Alignment via fine-grained question generation and answering. |
KAISI GUAN et. al. | iccv | 2025-10-20 |
| 856 | TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Given a video and a question, we generate an open-ended answer grounded with the start and end time. For this task, we propose TOGA: a vision-language model for Temporally Grounded Open-Ended Video QA with Weak Supervision. |
AYUSH GUPTA et. al. | iccv | 2025-10-20 |
| 857 | PVChat: Personalized Video Chat with One-Shot Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, we introduce an automated augmentation pipeline that synthesizes identity-preserving positive samples and retrieves hard negatives from existing video corpora, generating a diverse training dataset with four QA types: existence, appearance, action, and location inquiries. |
YUFEI SHI et. al. | iccv | 2025-10-20 |
| 858 | Object-centric Video Question Answering with Visual Grounding and Referring IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existing models primarily focus on high-level comprehension and are limited to text-only responses, restricting the flexibility for object-centric, multi-round interactions. In this paper, we make three contributions:(i) we address these limitations by introducing a VideoLLM, termed as **RGA3**, capable of performing both object referring and grounding for video reasoning tasks in a multi-round conversational manner, i.e., allowing users to iteratively interact with videos using both textual and visual queries; (ii) we propose **STOM** (Spatial-Temporal Overlay Module), a novel approach that allows arbitrary visual prompts to be processed at any timestamp within a video;(iii) we present **VideoInfer**, a manually curated object-centric video instruction dataset featuring question-answering pairs that require reasoning. |
HAOCHEN WANG et. al. | iccv | 2025-10-20 |
| 859 | 4D-Bench: Benchmarking Multi-modal Large Language Models for 4D Object Understanding Related Papers Related Patents Related Grants Related Venues Related Experts Related Code View Save Highlight: Multimodal Large Language Models (MLLMs) have demonstrated impressive 2D image/video understanding capabilities.However, there are no publicly standardized benchmarks to assess the abilities of MLLMs in understanding the 4D objects.In this paper, we introduce 4D-Bench, the first benchmark to evaluate the capabilities of MLLMs in 4D object understanding, featuring tasks in 4D object Question Answering (4D object QA) and 4D object captioning.4D-Bench provides 4D objects with diverse categories, high-quality annotations, and tasks necessitating multi-view spatial-temporal understanding, different from existing 2D image/video-based benchmarks.With 4D-Bench, we evaluate a wide range of open-source and closed-source MLLMs.The results from the 4D object captioning experiment indicate that MLLMs generally exhibit weaker temporal understanding compared to their appearance understanding, notably, while open-source models approach closed-source performance in appearance understanding, they show larger performance gaps in temporal understanding.4D object QA yields surprising findings: even with simple single-object videos, MLLMs perform poorly, with state-of-the-art GPT-4o achieving only 63% accuracy compared to the human baseline of 91%. |
WENXUAN ZHU et. al. | iccv | 2025-10-20 |
| 860 | SMR-agents: Synergistic Medical Reasoning Agents for Zero-shot Medical Visual Question Answering with MLLMs IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Dujuan Wang; Tao Cheng; Sutong Wang; Y. Chen; Yunqiang Yin; | Inf. Process. Manag. | |
| 861 | Multi-Agent Cooperation for Traffic Safety Description and Analysis Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Understanding complex traffic scenes from video remains a core challenge in building intelligent transportation systems, especially under varied viewpoints and semantic demands. … |
R. Kachhadiya; Dhanishtha Patil; David C. Anastasiu; | 2025 IEEE/CVF International Conference on Computer Vision … | 2025-10-19 |
| 862 | Prompt Design for Medical Question Answering with Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
LEONID KULIGIN et. al. | Machine Learning with Applications | 2025-10-17 |
| 863 | DTKG: Dual-Track Knowledge Graph-Verified Reasoning Framework for Multi-Hop QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: These limitations deteriorate the efficiency andaccuracy for multi-hop QA tasks. To address this challenge, we propose a noveldual-track KG verification and reasoning framework DTKG, which is inspired bythe Dual Process Theory in cognitive science. |
CHANGHAO WANG et. al. | arxiv-cs.AI | 2025-10-17 |
| 864 | SQuAI: Scientific Question-Answering with Multi-Agent Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present SQuAI (https://squai.scads.ai/), a scalable and trustworthymulti-agent retrieval-augmented generation (RAG) framework for scientificquestion answering (QA) with large language models (LLMs). |
Ines Besrour; Jingbo He; Tobias Schreieder; Michael Färber; | arxiv-cs.IR | 2025-10-17 |
| 865 | MedTrust-RAG: Evidence Verification and Trust Alignment for Biomedical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose MedTrust-Guided Iterative RAG, aframework designed to enhance factual consistency and mitigate hallucinationsin medical QA. |
YINGPENG NING et. al. | arxiv-cs.CL | 2025-10-16 |
| 866 | PRISM: Agentic Retrieval with LLMs for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Weintroduce an Agentic Retrieval System that leverages large language models(LLMs) in a structured loop to retrieve relevant evidence with high precisionand recall. |
Md Mahadi Hasan Nahid; Davood Rafiei; | arxiv-cs.CL | 2025-10-16 |
| 867 | Applications and Challenges of Retrieval-Augmented Generation (RAG) in Maternal Health: A Multi-Axial Review of The State of The Art in Biomedical QA with LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this context, retrieval-augmented generation (RAG) systems provide a promising approach to enhance traceability, timeliness, and accuracy in tasks such as biomedical question answering (QA). This article presents a narrative and thematic review of the evolution of these technologies in maternal health, structured across five axes: technical foundations of RAG, advancements in biomedical LLMs, conversational agents in healthcare, clinical validation frameworks, and specific applications in obstetric telehealth. |
ADRIANA NOGUERA et. al. | Sci | 2025-10-16 |
| 868 | Interactive Environment-Aware Planning System and Dialogue for Social Robots in Early Childhood Education Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this study, we propose an interactive environment-aware dialog and planning system for social robots in early childhood education, aimed at supporting the learning and social interaction of young children. |
Jiyoun Moon; Seung Min Song; | Applied Sciences | 2025-10-16 |
| 869 | PluriHop: Exhaustive, Recall-Sensitive QA Over Distractor-Rich Corpora Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To study thissetting, we introduce PluriHopWIND, a diagnostic multilingual dataset of 48pluri-hop questions built from 191 real-world wind industry reports in Germanand English. |
Mykolas Sveistrys; Richard Kunert; | arxiv-cs.CL | 2025-10-16 |
| 870 | BioMedSearch: A Multi-Source Biomedical Retrieval Framework Based on LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To evaluate the accuracy ofquestion answering, we constructed a multi-level dataset, BioMedMCQs,consisting of 3,000 questions. |
CONGYING LIU et. al. | arxiv-cs.CL | 2025-10-15 |
| 871 | Who’s Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we present the first systematic evaluation of LLM robustness to inquirypersonas, i.e. user profiles that convey attributes like identity, expertise,or belief. |
Nil-Jana Akpinar; Chia-Jung Lee; Vanessa Murdock; Pietro Perona; | arxiv-cs.CL | 2025-10-14 |
| 872 | Teaching Language Models to Faithfully Express Their Uncertainty Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We construct training data by augmenting modelsamples with uncertainty hedges (i.e. verbal cues such as ‘possibly’ or’likely’) aligned with sample consistency, requiring no supervision beyond themodel and a set of prompts. |
Bryan Eikema; Evgenia Ilia; José G. C. de Souza; Chrysoula Zerva; Wilker Aziz; | arxiv-cs.CL | 2025-10-14 |
| 873 | An Empirical Study for Representations of Videos in Video Question Answering Via MLLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present acomprehensive empirical study of video representation methods for VideoQA withMLLMs. |
Zhi Li; Yanan Wang; Hao Niu; Julio Vizcarra; Masato Taya; | arxiv-cs.IR | 2025-10-14 |
| 874 | ESI: Epistemic Uncertainty Quantification Via Semantic-preserving Intervention for Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we establish a connection between the uncertainty ofLLMs and their invariance under semantic-preserving intervention from a causalperspective. |
Mingda Li; Xinyu Li; Weinan Zhang; Longxuan Ma; | arxiv-cs.CL | 2025-10-14 |
| 875 | Discrepancy Detection at The Data Level: Toward Consistent Multilingual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluate MIND on a bilingualQA system in the maternal and infant health domain and release a dataset ofbilingual questions annotated for factual and cultural inconsistencies. |
LORENA CALVO-BARTOLOMÉ et. al. | arxiv-cs.CL | 2025-10-13 |
| 876 | VeritasFi: An Adaptable, Multi-tiered RAG Framework for Multi-modal Financial Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Retrieval-Augmented Generation (RAG) is becoming increasingly essential forQuestion Answering (QA) in the financial sector, where accurate andcontextually grounded insights from complex public disclosures are crucial.However, existing financial RAG systems face two significant challenges: (1)they struggle to process heterogeneous data formats, such as text, tables, andfigures; and (2) they encounter difficulties in balancing general-domainapplicability with company-specific adaptation. |
ZHENGHAN TAI et. al. | arxiv-cs.IR | 2025-10-12 |
| 877 | RIPRAG: Hack A Black-box Retrieval-Augmented Generation Question-Answering System with Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, weinvestigate a more complex and realistic scenario: the attacker lacks knowledgeof the RAG system’s internal composition and implementation details, and theRAG system comprises components beyond a mere retriever. |
MENG XI et. al. | arxiv-cs.AI | 2025-10-11 |
| 878 | Robust Clinical Querying with Local LLMs: Lexical Challenges in NL2SQL and Retrieval-Augmented QA on EHRs Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Electronic health records (EHRs) are typically stored in relational databases, making them difficult to query for nontechnical users, especially under privacy constraints. We … |
Luka Blašković; Nikola Tanković; I. Lorencin; Sandi Baressi Šegota; | Big Data Cogn. Comput. | 2025-10-11 |
| 879 | LONGQAEVAL: Designing Reliable Evaluations of Long-Form Clinical QA Under Resource Constraints Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LongQAEval, an evaluation framework and set ofevaluation recommendations for limited-resource and high-expertise settings.Based on physician annotations of 300 real patient questions answered byphysicians and LLMs, we compare coarse answer-level versus fine-grainedsentence-level evaluation over the dimensions of correctness, relevance, andsafety. |
Federica Bologna; Tiffany Pan; Matthew Wilkens; Yue Guo; Lucy Lu Wang; | arxiv-cs.CL | 2025-10-11 |
| 880 | AssoMem: Scalable Memory QA with Multi-Signal Associative Retrieval Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Inspired by how humans link informationassociatively, we propose AssoMem, a novel framework constructing anassociative memory graph that anchors dialogue utterances to automaticallyextracted clues. |
KAI ZHANG et. al. | arxiv-cs.CL | 2025-10-11 |
| 881 | Closing The Data-Efficiency Gap Between Autoregressive and Masked Diffusion LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thisproposed method successfully and drastically improves the data efficiency ofarLLM fine-tuning, effectively closing the performance gap with dLLMs. |
Xu Pan; Ely Hahami; Jingxuan Fan; Ziqian Xie; Haim Sompolinsky; | arxiv-cs.CL | 2025-10-10 |
| 882 | NG-Router: Graph-Supervised Multi-Agent Collaboration for Nutrition Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To further address contextual overload, we propose agradient-based subgraph retrieval mechanism that identifies salient evidenceduring training, thereby enhancing multi-hop and relational reasoning.Extensive experiments across multiple benchmarks and backbone modelsdemonstrate that NG-Router consistently outperforms both single-agent andensemble baselines, offering a principled approach to domain-aware multi-agentreasoning for complex nutritional health tasks. |
KAIWEN SHI et. al. | arxiv-cs.CL | 2025-10-10 |
| 883 | AI Knowledge Assist: An Automated Approach for The Creation of Knowledge Bases for Conversational AI Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To this end,we introduce AI Knowledge Assist, a system that extracts knowledge in the formof question-answer (QA) pairs from historical customer-agent conversations toautomatically build a knowledge base. |
Md Tahmid Rahman Laskar; Julien Bouvier Tremblay; Xue-Yong Fu; Cheng Chen; Shashi Bhushan TN; | arxiv-cs.CL | 2025-10-09 |
| 884 | A2Search: Ambiguity-Aware Question Answering with Reinforcement Learning Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent advances in Large Language Models (LLMs) and Reinforcement Learning (RL) have led to strong performance in open-domain question answering (QA). However, existing models … |
FENGJI ZHANG et. al. | ArXiv | 2025-10-09 |
| 885 | IDQuAD: Infectious Disease Question and Answering Dataset Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In conclusion, this study introduces IDQuAD as a foundational dataset for infectious disease research, demonstrating the effectiveness of fine-tuning LLMs and paving the way for future advances in dataset development and LLM refinement for infectious disease tasks. |
Soonchan Kwon; Sujeong Hur; Beakcheol Jang; | PLOS One | 2025-10-09 |
| 886 | LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Thisoften results in incomplete evidence retrieval and degraded answer quality formulti-page reasoning tasks. To address these limitations, we propose LAD-RAG, anovel Layout-Aware Dynamic RAG framework. |
ZHIVAR SOURATI et. al. | arxiv-cs.CL | 2025-10-08 |
| 887 | SUBQRAG: Sub-question Driven Dynamic Graph Rag Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Graph Retrieval-Augmented Generation (Graph RAG) effectively builds aknowledge graph (KG) to connect disparate facts across a large document corpus.However, this broad-view … |
JIAOYANG LI et. al. | arxiv-cs.CL | 2025-10-08 |
| 888 | Asking For It: Question-Answering for Predicting Rule Infractions in Online Content Moderation Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Online communities rely on a mix of platform policies and community-authoredrules to define acceptable behavior and maintain order. However, these rulesvary widely across … |
Mattia Samory; Diana Pamfile; Andrew To; Shruti Phadke; | arxiv-cs.CY | 2025-10-07 |
| 889 | EverydayMMQA: A Multilingual and Multimodal Framework for Culturally Grounded Spoken Visual QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large-scale multimodal models achieve strong results on tasks like VisualQuestion Answering (VQA), but they often fail when queries require culturallygrounded, everyday knowledge, particularly in low-resource and underrepresentedlanguages. To bridge this gap, we introduce Everyday Multimodal andMultilingual QA (EverydayMMQA), a framework for creating large-scale,culturally-grounded datasets for spoken and visual question answering (SVQA). |
FIROJ ALAM et. al. | arxiv-cs.CL | 2025-10-07 |
| 890 | Multi-Hop Question Answering: When Can Humans Help, and Where Do They Struggle? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To better understand how humans might collaborate effectively withAI, we evaluate the performance of crowd workers on these individual reasoningsubtasks. We find that while humans excel at knowledge integration (97\%accuracy), they often fail to recognize when a question requires multi-hopreasoning (67\% accuracy). |
Jinyan Su; Claire Cardie; Jennifer Healey; | arxiv-cs.HC | 2025-10-06 |
| 891 | Knowledge Graph-Guided Multi-Agent Distillation for Reliable Industrial Question Answering with Datasets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Industrial question-answering (QA) systems require higher safety andreliability than general-purpose dialogue models, as errors in high-riskscenarios such as equipment fault diagnosis can have severe consequences.Although multi-agent large language models enhance reasoning depth, they sufferfrom uncontrolled iterations and unverifiable outputs, and conventionaldistillation methods struggle to transfer collaborative reasoning capabilitiesto lightweight, deployable student models. |
Jiqun Pan; Zhenke Duan; Jiani Tu; Anzhi Cheng; Yanqing Wang; | arxiv-cs.CL | 2025-10-03 |
| 892 | StepChain GraphRAG: Reasoning Over Knowledge Graphs for Multi-Hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Yet, challengespersist in integrating iterative reasoning steps with external knowledgeretrieval. To address this, we introduce StepChain GraphRAG, a framework thatunites question decomposition with a Breadth-First Search (BFS) Reasoning Flowfor enhanced multi-hop QA. |
TENGJUN NI et. al. | arxiv-cs.CL | 2025-10-03 |
| 893 | LEAML: Label-Efficient Adaptation to Out-of-Distribution Visual Tasks for Multimodal Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce LEAML, a label-efficient adaptation framework thatleverages both scarce labeled VQA samples and abundant unlabeled images. |
Ci-Siang Lin; Min-Hung Chen; Yu-Yang Sheng; Yu-Chiang Frank Wang; | arxiv-cs.CV | 2025-10-03 |
| 894 | Triplet-Structured Knowledge Integration for Multi-Turn Medical Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper introduces TriMediQ, a triplet-structured approach thatenhances the reasoning reliability of LLMs through explicit knowledgeintegration. |
Zhaohan Meng; Zaiqiao Meng; Siwei Liu; Iadh Ounis; | arxiv-cs.CL | 2025-10-03 |
| 895 | Uncertainty As Feature Gaps: Epistemic Uncertainty Quantification of LLMs in Contextual Question-Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we focus on UQfor the contextual QA task and propose a theoretically grounded approach toquantify epistemic uncertainty. |
YAVUZ BAKMAN et. al. | arxiv-cs.CL | 2025-10-02 |
| 896 | LLM Guided Counterfactual Reasoning for Zero-shot Knowledge Based Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Zhuhan Zhang; Min Jiang; Jun Kong; Jiayi Li; | Neurocomputing | 2025-10-01 |
| 897 | One More Question Is Enough, Expert Question Decomposition (EQD) Model for Domain Quantitative Reasoning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose Expert QuestionDecomposition (EQD), an approach designed to balance the use of domainknowledge with computational efficiency. |
Mengyu Wang; Sotirios Sabanis; Miguel de Carvalho; Shay B. Cohen; Tiejun Ma; | arxiv-cs.CL | 2025-10-01 |
| 898 | TAG-EQA: Text-And-Graph for Event Question Answering Via Structured Prompting Strategies Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TAG-EQA (Text-And-Graph for Event QuestionAnswering), a prompting framework that injects causal event graphs into LLMinputs by converting structured relations into natural-language statements.TAG-EQA spans nine prompting configurations, combining three strategies(zero-shot, few-shot, chain-of-thought) with three input modalities (text-only,graph-only, text+graph), enabling a systematic analysis of when and howstructured knowledge aids inference. |
Maithili Kadam; Francis Ferraro; | arxiv-cs.CL | 2025-10-01 |
| 899 | Question Answering System Based on The Combination of Large Language Model and Knowledge Graph Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Jihong Wang; Yichen Zhang; Wei Liu; | Applied Intelligence | 2025-10-01 |
| 900 | CODNet: Context-based Object Detection Network for Multimodal Image Captioning and Virtual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Chhaya Gupta; N. S. Gill; Preeti Gulia; Giovanni Pau; | Image Vis. Comput. | 2025-10-01 |
| 901 | Learning to Route: A Rule-Driven Agent Framework for Hybrid-Source Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large Language Models (LLMs) have shown remarkable performance on generalQuestion Answering (QA), yet they often struggle in domain-specific scenarioswhere accurate and up-to-date information is required. Retrieval-AugmentedGeneration (RAG) addresses this limitation by enriching LLMs with externalknowledge, but existing systems primarily rely on unstructured documents, whilelargely overlooking relational databases, which provide precise, timely, andefficiently queryable factual information, serving as indispensableinfrastructure in domains such as finance, healthcare, and scientific research.Motivated by this gap, we conduct a systematic analysis that reveals threecentral observations: (i) databases and documents offer complementary strengthsacross queries, (ii) naively combining both sources introduces noise and costwithout consistent accuracy gains, and (iii) selecting the most suitable sourcefor each query is crucial to balance effectiveness and efficiency. |
HAOYUE BAI et. al. | arxiv-cs.CL | 2025-09-30 |
| 902 | A Multimodal LLM Approach for Visual Question Answering on Multiparametric 3D Brain MRI Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce mpLLM, a prompt-conditioned hierarchical mixture-of-experts(MoE) architecture for visual question answering over multi-parametric 3D brainMRI (mpMRI). |
ARVIND MURARI VEPA et. al. | arxiv-cs.CV | 2025-09-30 |
| 903 | RAGferee: Building Contextual Reward Models for Retrieval-Augmented Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address the lack of publicly availableRAG-centric preference datasets and specialised RMs, we introduce RAGferee, amethodology that repurposes question-answering (QA) datasets into preferencepairs that prioritise groundedness over stylistic features, enabling thetraining of contextual RMs better suited to judging RAG responses. |
ANDREI C. COMAN et. al. | arxiv-cs.CL | 2025-09-30 |
| 904 | Boosting Process-Correct CoT Reasoning By Modeling Solvability of Multiple-Choice QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study thisthrough multiple-choice question answering (MCQA), which provides a controlledsetting with fixed answer options. |
Raphael Schumann; Stefan Riezler; | arxiv-cs.AI | 2025-09-30 |
| 905 | SimulRAG: Simulator-based RAG for Grounding LLMs in Long-form Scientific QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this framework, we propose a generalized simulatorretrieval interface to transform between textual and numerical modalities. |
HAOZHOU XU et. al. | arxiv-cs.CL | 2025-09-29 |
| 906 | Can VLM Pseudo-Labels Train A Time-Series QA Model That Outperforms The VLM? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Alternatively, with recent advancements inlarge-scale models, vision-language models (VLMs) have demonstrated thepotential to analyze time-series signals in a zero-shot manner. In this paper,we propose a training approach that uses pseudo labels generated by a VLM.Although VLMs can produce incorrect labels, TSQA models can still beeffectively trained based on the property that deep neural networks areinherently robust to such noisy labels. |
Takuya Fujimura; Kota Dohi; Natsuo Yamashita; Yohei Kawaguchi; | arxiv-cs.LG | 2025-09-29 |
| 907 | Knowledge Extraction on Semi-Structured Content: Does It Remain Relevant for Question Answering in The Era of LLMs? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: The advent of Large Language Models (LLMs) has significantly advancedweb-based Question Answering (QA) systems over semi-structured content, raisingquestions about the continued utility of knowledge extraction for questionanswering. This paper investigates the value of triple extraction in this newparadigm by extending an existing benchmark with knowledge extractionannotations and evaluating commercial and open-source LLMs of varying sizes.Our results show that web-scale knowledge extraction remains a challenging taskfor LLMs. |
KAI SUN et. al. | arxiv-cs.CL | 2025-09-29 |
| 908 | Evaluating RAG-based QA Systems: A Comparative Analysis of LLM As A Judge, Traditional Metrics, and Human Alignment Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Evaluating RAG based Question Answering systems presents ongoing challenges, as traditional NLP metrics often inadequately capture nuanced answer quality and the reliability of … |
Renato Miyaji; Renato Moulin; Samuel Monção; Leonardo Machado; | Brazilian Symposium in Information and Human Language … | 2025-09-29 |
| 909 | Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To move beyond thisparadigm, we introduce a framework to synthesize richer supervisory signals. |
JIANXIN LIANG et. al. | arxiv-cs.CV | 2025-09-29 |
| 910 | Do LLMs Understand Romanian Driving Laws? A Study on Multimodal and Fine-Tuned Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper evaluates Large Language Models (LLMs)on Romanian driving-law QA with explanation generation. |
Eduard Barbu; Adrian Marius Dumitran; | arxiv-cs.CL | 2025-09-28 |
| 911 | MIRAGE: Multi-hop Reasoning with Ambiguity Evaluation for Illusory Questions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Toestablish a robust baseline, we propose CLarifying Ambiguity with a Reasoningand InstructiON (CLARION), a multi-agent framework that significantlyoutperforms existing approaches on MIRAGE, paving the way for more adaptive androbust reasoning systems. |
JEONGHYUN PARK et. al. | arxiv-cs.CL | 2025-09-26 |
| 912 | Do LLM Agents Know How to Ground, Recover, and Assess? A Benchmark for Epistemic Competence in Information-Seeking Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Recent work has explored training Large Language Model (LLM) search agentswith reinforcement learning (RL) for open-domain question answering (QA). |
Jiaqi Shao; Yuxiang Lin; Munish Prasad Lohani; Yufeng Miao; Bing Luo; | arxiv-cs.AI | 2025-09-26 |
| 913 | Detecting (Un)answerability in Large Language Models with Linear Directions Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we study the problem of (un)answerability detection, focusing onextractive question answering (QA) where the model should determine if apassage contains sufficient information to answer a given question. |
Maor Juliet Lavi; Tova Milo; Mor Geva; | arxiv-cs.CL | 2025-09-26 |
| 914 | From Evidence to Trajectory: Abductive Reasoning Path Synthesis for Training Retrieval-Augmented Generation Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose EviPath, an evidence-anchored reasoning path synthesis paradigm for RAGagent development. |
MUZHI LI et. al. | arxiv-cs.CL | 2025-09-26 |
| 915 | JGU Mainz’s Submission to The WMT25 Shared Task on LLMs with Limited Resources for Slavic Languages: MT and QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents the JGU Mainz submission to the WMT25 Shared Task on LLMswith Limited Resources for Slavic Languages: Machine Translation and QuestionAnswering, focusing on Ukrainian, Upper Sorbian, and Lower Sorbian. |
Hossain Shaikh Saadi; Minh Duc Bui; Mario Sanz-Guerrero; Katharina von der Wense; | arxiv-cs.CL | 2025-09-26 |
| 916 | A Comprehensive Evaluation of Transformer-Based Question Answering Models and RAG-Enhanced Design Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents a comprehensive evaluation ofretrieval strategies for multi-hop question answering within aretrieval-augmented generation framework. |
Zichen Zhang; Kunlong Zhang; Hongwei Ruan; Yiming Luo; | arxiv-cs.CV | 2025-09-26 |
| 917 | Beyond Stars: Bridging The Gap Between Ratings and Review Sentiment with LLM Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present an advanced approach to mobile app review analysis aimed ataddressing limitations inherent in traditional star-rating systems. |
Najla Zuhir; Amna Mohammad Salim; Parvathy Premkumar; Moshiur Farazi; | arxiv-cs.AI | 2025-09-25 |
| 918 | RJE: A Retrieval-Judgment-Exploration Framework for Efficient Knowledge Graph Question Answering with LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address theselimitations, we propose Retrieval-Judgment-Exploration (RJE), a framework thatretrieves refined reasoning paths, evaluates their sufficiency, andconditionally explores additional evidence. |
CAN LIN et. al. | arxiv-cs.CL | 2025-09-24 |
| 919 | LOCA: Logical Chain Augmentation for Scientific Corpus Cleaning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, existingscientific question-answering (QA) datasets suffer from high error rates,frequently resulting from logical leaps and implicit reasoning within theanswers. To address this issue, we introduce LOCA (Logical Chain Augmentation),a novel framework for automatically cleaning scientific corpora, implementedthrough an augment-and-review loop. |
YOU-LE FANG et. al. | arxiv-cs.CL | 2025-09-24 |
| 920 | CON-QA: Privacy-Preserving QA Using Cloud LLMs in Contract Domain Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose CON-QA, a hybridprivacy-preserving framework designed specifically for secure questionanswering over enterprise contracts, effectively combining local andcloud-hosted LLMs. |
Ajeet Kumar Singh; Rajsabi Surya; Anurag Tripathi; Santanu Choudhury; Sudhir Bisane; | arxiv-cs.AI | 2025-09-24 |
| 921 | Pathways of Thoughts: Multi-Directional Thinking for Long-form Personalized Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, personalized QA remains relatively underexplored due tochallenges such as inferring preferences from long, noisy, and implicitcontexts, and generating responses that are simultaneously correct,contextually appropriate, and aligned with user expectations and backgroundknowledge. To address these challenges, we propose Pathways of Thoughts (PoT),an inference-stage method that applies to any large language model (LLM)without requiring task-specific fine-tuning. |
ALIREZA SALEMI et. al. | arxiv-cs.CL | 2025-09-23 |
| 922 | Are Smaller Open-Weight LLMs Closing The Gap to Proprietary Models for Biomedical Question Answering? Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we compare several open-weight modelsagainst top-performing systems such as GPT-4o, GPT-4.1, Claude 3.5 Sonnet, andClaude 3.7 Sonnet. |
Damian Stachura; Joanna Konieczna; Artur Nowak; | arxiv-cs.CL | 2025-09-23 |
| 923 | Semantic Reformulation Entropy for Robust Hallucination Detection in QA Tasks Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose Semantic Reformulation Entropy (SRE), whichimproves uncertainty estimation in two ways. |
CHAODONG TONG et. al. | arxiv-cs.CL | 2025-09-22 |
| 924 | MSCoRe: A Benchmark for Multi-Stage Collaborative Reasoning in LLM Agents Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing benchmarkstypically focus on isolated tasks or narrow domains, overlooking models’abilities for multi-stage collaboration and optimization without explicitexternal guidance. To bridge this gap, we propose \textbf{MSCoRe}, a novelbenchmark comprising 126696 domain-specific QA instances spanning scenarios inautomotive, pharmaceutical, electronics, and energy sectors. |
Yuzhen Lei; Hongbin Xie; Jiaxing Zhao; Shuangxue Liu; Xuan Song; | arxiv-cs.CL | 2025-09-22 |
| 925 | Memory-QA: Answering Recall Questions Based on Multimodal Memories Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This task poses unique challenges, including the creation oftask-oriented memories, the effective utilization of temporal and locationinformation within memories, and the ability to draw upon multiple memories toanswer a recall question. To address these challenges, we propose acomprehensive pipeline, Pensieve, integrating memory-specific augmentation,time- and location-aware multi-signal retrieval, and multi-memory QAfine-tuning. |
HONGDA JIANG et. al. | arxiv-cs.AI | 2025-09-22 |
| 926 | LLaVul: A Multimodal LLM for Interpretable Vulnerability Reasoning About Source Code Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our model is trained to integrate paired codeand natural queries into a unified space, enhancing reasoning andcontext-dependent insights about code vulnerability. |
Ala Jararweh; Michael Adams; Avinash Sahu; Abdullah Mueen; Afsah Anwar; | arxiv-cs.AI | 2025-09-21 |
| 927 | AirQA: A Comprehensive QA Dataset for AI Research with Instance-Level Evaluation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we propose AirQA, a human-annotatedcomprehensive paper QA dataset in the field of artificial intelligence (AI),with 13,948 papers and 1,246 questions, that encompasses multi-task,multi-modal and instance-level evaluation. |
TIANCHENG HUANG et. al. | arxiv-cs.CL | 2025-09-21 |
| 928 | Comparing RAG and GraphRAG for Page-Level Retrieval Question Answering on Math Textbook Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Overall, this study highlights both the promises andchallenges of page-level retrieval systems in educational contexts, emphasizingthe need for more refined retrieval methods to build reliable AI tutoringsolutions in providing reference page numbers. |
EASON CHEN et. al. | arxiv-cs.IR | 2025-09-20 |
| 929 | Time to Revist Exact Match Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce TempAnswerQA, abenchmark distilled from Test of Time and TempTabQA, where all questionsrequire a numerical, temporal answer, allowing us to evaluate models beyond EM.We use the forecasting metrics symmetric mean absolute percentage error (sMAPE)and mean absolute scaled error (MASE). |
Auss Abbood; Zaiqiao Meng; Nigel Collier; | arxiv-cs.CL | 2025-09-20 |
| 930 | Benchmarking Contextual and Paralinguistic Reasoning in Speech-LLMs: A Case Study with In-the-Wild Data Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Recent speech-LLMs have shown impressive performance in tasks liketranscription and translation, yet they remain limited in understanding theparalinguistic aspects of speech … |
QIONGQIONG WANG et. al. | arxiv-cs.CL | 2025-09-20 |
| 931 | Question Answering with LLMs and Learning from Answer Sets Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce LLM2LAS, ahybrid system that effectively combines the natural language understandingcapabilities of LLMs, the rule induction power of the Learning from Answer Sets(LAS) system ILASP, and the formal reasoning strengths of Answer SetProgramming (ASP). |
MANUEL BORROTO et. al. | arxiv-cs.AI | 2025-09-20 |
| 932 | RephQA: Evaluating Readability of Large Language Models in Public Health Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce RephQA, abenchmark for evaluating the readability of LLMs in public health questionanswering (QA). |
WEIKANG QIU et. al. | arxiv-cs.CL | 2025-09-19 |
| 933 | Jamendo-QA: A Large-Scale Music Question Answering Dataset Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce Jamendo-QA, a large-scale dataset for Music Question Answering(Music-QA). |
Junyoung Koh; Soo Yong Kim; Yongwon Choi; Gyu Hyeong Choi; | arxiv-cs.MM | 2025-09-19 |
| 934 | SWE-QA: Can Language Models Answer Repository-level Code Questions? IF:3 Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thispaper, we present SWE-QA, a repository-level code question answering (QA)benchmark designed to facilitate research on automated QA systems in realisticcode environments. |
WEIHAN PENG et. al. | arxiv-cs.CL | 2025-09-18 |
| 935 | Quantifying Uncertainty in Natural Language Explanations of Large Language Models for Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Notably, generatingvalid uncertainty estimates for natural language explanations is particularlychallenging due to the auto-regressive generation process of LLMs and thepresence of noise in medical inquiries. To bridge this gap, in this work, wefirst propose a novel uncertainty estimation framework for these generatednatural language explanations, which provides valid uncertainty guarantees in apost-hoc and model-agnostic manner. |
Yangyi Li; Mengdi Huai; | arxiv-cs.CL | 2025-09-18 |
| 936 | Findings of The Third Automatic Minuting (AutoMin) Challenge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper presents the third edition of AutoMin, a shared task on automaticmeeting summarization into minutes. |
Kartik Shinde; Laurent Besacier; Ondrej Bojar; Thibaut Thonet; Tirthankar Ghosal; | arxiv-cs.CL | 2025-09-17 |
| 937 | AQUA-LLM: Evaluating Accuracy, Quantization, and Adversarial Robustness Trade-offs in LLMs for Cybersecurity Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose AQUA-LLM, an evaluationframework designed to benchmark several state-of-the-art small LLMs under fourdistinct configurations: base, quantized-only, fine-tuned, and fine-tunedcombined with quantization, specifically for cybersecurity QA. |
Onat Gungor; Roshan Sood; Harold Wang; Tajana Rosing; | arxiv-cs.CR | 2025-09-16 |
| 938 | HistoryBankQA: Multilingual Temporal Question Answering on Historical Events Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing temporal reasoning datasetsare limited in scale, lack multilingual coverage and focus more on contemporaryevents. To address these limitations, we present HistoryBank, a multilingualdatabase of 10M+ historical events extracted from Wikipedia timeline pages andarticle infoboxes. |
Biswadip Mandal; Anant Khandelwal; Manish Gupta; | arxiv-cs.CL | 2025-09-16 |
| 939 | ParaEQsA: Parallel and Asynchronous Embodied Questions Scheduling and Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper formulates the Embodied Questions Answering (EQsA) problem,introduces a corresponding benchmark, and proposes a system to tackle theproblem. |
Haisheng Wang; Weiming Zhi; | arxiv-cs.RO | 2025-09-15 |
| 940 | Graph-Enhanced Retrieval-Augmented Question Answering for E-Commerce Customer Support Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paper develops a novelretrieval-augmented generation (RAG) framework that uses knowledge graphs (KGs)to improve the relevance of the answer and the factual grounding. |
Piyushkumar Patel; | arxiv-cs.CL | 2025-09-15 |
| 941 | Bridging Vision Language Models and Symbolic Grounding for Video Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study symbolic scene graphs(SGs) as intermediate grounding signals for VQA. |
Haodi Ma; Vyom Pathak; Daisy Zhe Wang; | arxiv-cs.CV | 2025-09-15 |
| 942 | MORQA: Benchmarking Evaluation Metrics for Medical Open-Ended Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thiswork, we introduce MORQA (Medical Open-Response QA), a new multilingualbenchmark designed to assess the effectiveness of NLG evaluation metrics acrossthree medical visual and text-based QA datasets in English and Chinese. |
WEN-WAI YIM et. al. | arxiv-cs.CL | 2025-09-15 |
| 943 | AgenticIE: An Adaptive Agent for Information Extraction from Complex Regulatory Documents Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Declaration of Performance (DoP) documents, mandated by EU regulation,certify the performance of construction products. There are two challenges tomake DoPs machine and human … |
Gaye Colakoglu; Gürkan Solmaz; Jonathan Fürst; | arxiv-cs.CL | 2025-09-15 |
| 944 | Improving LLMs’ Learning for Coreference Resolution Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, weinvestigate the limitations of existing LLM-based approaches to CR-specificallythe Question-Answering (QA) Template and Document Template methods and proposetwo novel techniques: Reversed Training with Joint Inference and IterativeDocument Generation. |
Yujian Gan; Yuan Liang; Yanni Lin; Juntao Yu; Massimo Poesio; | arxiv-cs.CL | 2025-09-14 |
| 945 | !MSA at AraHealthQA 2025 Shared Task: Enhancing LLM Performance for Arabic Clinical Question Answering Through Prompt Engineering and Ensemble Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present our systems for Track 2 (General Arabic Health QA, MedArabiQ) ofthe AraHealthQA-2025 shared task, where our methodology secured 2nd place inboth Sub-Task 1 (multiple-choice question answering) and Sub-Task 2 (open-endedquestion answering) in Arabic clinical contexts. |
Mohamed Tarek; Seif Ahmed; Mohamed Basem; | arxiv-cs.CL | 2025-09-14 |
| 946 | A Knowledge Noise Mitigation Framework for Knowledge-based Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Knowledge-based visual question answering (KB-VQA) requires a model tounderstand images and utilize external knowledge to provide accurate answers.Existing approaches often directly augment models with retrieved informationfrom knowledge sources while ignoring substantial knowledge redundancy, whichintroduces noise into the answering process. To address this, we propose atraining-free framework with knowledge focusing for KB-VQA, that mitigates theimpact of noise by enhancing knowledge relevance and reducing redundancy.First, for knowledge retrieval, our framework concludes essential parts fromthe image-question pairs, creating low-noise queries that enhance the retrievalof highly relevant knowledge. |
Zhiyue Liu; Sihang Liu; Jinyuan Liu; Xinru Zhang; | arxiv-cs.CV | 2025-09-11 |
| 947 | Constructing A Question-Answering Simulator Through The Distillation of LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, wepropose a method named LLM Distillation based Simulator (LDSim), which distillsdomain knowledge and reasoning capability from an LLM to better assistprediction, thereby improving simulation performance. |
Haipeng Liu; Ting Long; Jing Fu; | arxiv-cs.LG | 2025-09-11 |
| 948 | Agentic LLMs for Question Answering Over Tabular Data Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We propose a Natural Language to SQL (NL-to-SQL) approachleveraging large language models (LLMs) such as GPT-4o, GPT-4o-mini, andDeepSeek v2:16b to generate SQL queries dynamically. |
Rishit Tyagi; Mohit Gupta; Rahul Bouri; | arxiv-cs.CL | 2025-09-11 |
| 949 | Enhancing Ancient Ceramic Knowledge Services: A Question Answering System Using Fine-Tuned Models and GraphRAG Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: To address the challenges of extensive domain expertise and deficient semantic comprehension in the digital preservation of ancient ceramics, this paper proposes a knowledge … |
Zhi Chen; Bingxiang Liu; | Inf. | 2025-09-11 |
| 950 | Fusing Knowledge and Language: A Comparative Study of Knowledge Graph-Based Question Answering with LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: While traditional Retrieval Augmented Generation(RAG) approaches are proficient in fact-based and local context-basedextraction from concise texts, they encounter limitations when addressing thethematic and holistic understanding of complex, extensive texts, requiring adeeper analysis of both text and context. This paper presents a comprehensivetechnical comparative study of three different methodologies for constructingknowledge graph triplets and integrating them with Large Language Models (LLMs)for question answering: spaCy, Stanford CoreNLP-OpenIE, and GraphRAG, allleveraging open source technologies. |
Vaibhav Chaudhary; Neha Soni; Narotam Singh; Amita Kapoor; | arxiv-cs.AI | 2025-09-11 |
| 951 | Retrieval-Augmented Generation for Reliable Interpretation of Radio Regulations Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We study question answering in the domain of radio regulations, a legallysensitive and high-stakes area. We propose a telecom-specificRetrieval-Augmented Generation (RAG) pipeline and introduce, to our knowledge,the first multiple-choice evaluation set for this domain, constructed fromauthoritative sources using automated filtering and human validation. |
Zakaria El Kassimi; Fares Fourati; Mohamed-Slim Alouini; | arxiv-cs.IR | 2025-09-11 |
| 952 | LLM Ensemble for RAG: Role of Context Length in Zero-Shot Question Answering for BioASQ Challenge Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we explore how largelanguage models (LLMs) can be used for information retrieval (IR), and anensemble of zero-shot models can accomplish state-of-the-art performance on adomain-specific Yes/No QA task. |
Dima Galat; Diego Molla-Aliod; | arxiv-cs.CL | 2025-09-10 |
| 953 | Towards Knowledge-Aware Document Systems: Modeling Semantic Coverage Relations Via Answerability Detection Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we introduce anovel framework for modelling Semantic Coverage Relations (SCR), whichclassifies document pairs based on how their informational content aligns. |
Yehudit Aperstein; Alon Gottlib; Gal Benita; Alexander Apartsin; | arxiv-cs.CL | 2025-09-10 |
| 954 | A Role-Aware Multi-Agent Framework for Financial Education Question Answering with LLMs Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We evaluated our framework on a set of3,532 expert-designed finance education questions from Study.com, an onlinelearning platform. |
Andy Zhu; Yingjun Du; | arxiv-cs.CL | 2025-09-10 |
| 955 | Avoiding Knowledge Edit Skipping in Multi-hop Question Answering with Guided Decomposition Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To address thisissue, we propose a novel Iterative Retrieval-Augmented Knowledge Editingmethod with guided decomposition (IRAKE) through the guidance from singleedited facts and entire edited cases. |
Yi Liu; Xiangrong Zhu; Xiangyu Liu; Wei Wei; Wei Hu; | arxiv-cs.CL | 2025-09-09 |
| 956 | Aligning LLMs for The Classroom with Knowledge-Based Retrieval — A Comparative RAG Study Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Large language models like ChatGPT are increasingly used in classrooms, butthey often provide outdated or fabricated information that can misleadstudents. |
Amay Jain; Liu Cui; Si Chen; | arxiv-cs.AI | 2025-09-09 |
| 957 | The Role of Exploration Modules in Small Language Models for Knowledge Graph Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thisstudy, we investigate the capabilities of existing integration methods forsmall language models (SLMs) in KG-based question answering and observe thattheir performance is often constrained by their limited ability to traverse andreason over knowledge graphs. To address this limitation, we propose leveragingsimple and efficient exploration modules to handle knowledge graph traversal inplace of the language model itself. |
Yi-Jie Cheng; Oscar Chew; Yun-Nung Chen; | arxiv-cs.CL | 2025-09-09 |
| 958 | Aligning LLMs for The Classroom with Knowledge-Based Retrieval: A Comparative RAG Study Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Large language models like ChatGPT are increasingly used in classrooms, but they often provide outdated or fabricated information that can mislead students. Retrieval Augmented … |
Amay Jain; Liu Cui; Si Chen; | 2025 IEEE International Conference on Teaching, Assessment, … | 2025-09-09 |
| 959 | RTLExplain: A Structured Approach to RTL Code Summarization and Question Answering for Medium-to-Large Designs Using LLMs Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Large Language Models (LLMs) show promise in assisting with Register Transfer Level (RTL) design tasks, including code summarization, documentation, and question answering. … |
TING-HSUN CHI et. al. | 2025 ACM/IEEE 7th Symposium on Machine Learning for CAD … | 2025-09-08 |
| 960 | KERAG: Knowledge-Enhanced Retrieval-Augmented Generation for Advanced Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present KERAG, a novel KG-based RAG pipeline thatenhances QA coverage by retrieving a broader subgraph likely to containrelevant information. |
YUSHI SUN et. al. | arxiv-cs.CL | 2025-09-04 |
| 961 | Facts Fade Fast: Evaluating Memorization of Outdated Medical Knowledge in Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: WhenLLMs memorize outdated medical knowledge, they can provide harmful advice orfail at clinical reasoning tasks. To investigate this problem, we introduce twonovel question-answering (QA) datasets derived from systematic reviews:MedRevQA (16,501 QA pairs covering general biomedical knowledge) andMedChangeQA (a subset of 512 QA pairs where medical consensus has changed overtime). |
Juraj Vladika; Mahdi Dhaini; Florian Matthes; | arxiv-cs.CL | 2025-09-04 |
| 962 | ProMQA-Assembly: Multimodal Procedural QA Dataset on Assembly Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To foster the development, we propose a new multimodal QA dataset onassembly activities. |
KIMIHIRO HASEGAWA et. al. | arxiv-cs.CL | 2025-09-02 |
| 963 | A-SEA3L-QA: A Fully Automated Self-Evolving, Adversarial Workflow for Arabic Long-Context Question-Answer Generation Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present an end-to-end, self-evolving adversarial workflow for long-contextQuestion-Answer (QA) Generation in Arabic. |
Kesen Wang; Daulet Toibazar; Pedro J. Moreno; | arxiv-cs.CL | 2025-09-02 |
| 964 | End-to-End 3D Point Cloud Question Answering Via State Space Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Hatem Qorsham; R. El-Deeb; Ehab Essa; | Appl. Soft Comput. | 2025-09-01 |
| 965 | LOSDF: A Logical Optimization and Semantic Decoupling Framework for Question Answering in Multi-party Conversations Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
SHU ZHOU et. al. | Inf. Process. Manag. | 2025-09-01 |
| 966 | Open-ViTabQA: A Novel Benchmark for Vietnamese Question Answering on Open Domain Wikipedia Table Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Dung Dao; N. Huynh; Khanh Quoc Tran; K. Nguyen; | Knowl. Based Syst. | 2025-09-01 |
| 967 | Temporal-Guided Mixture-of-Experts for Zero-Shot Video Question Answering Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Video Question Answering (VideoQA) is a challenging task in the vision-language field. Due to the time-consuming and labor-intensive labeling process of the question-answer pairs, … |
YIXIN QIN et. al. | IEEE Transactions on Circuits and Systems for Video … | 2025-09-01 |
| 968 | Learned Hallucination Detection in Black-Box LLMs Using Token-level Entropy Production Rate Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This paperintroduces an applied methodology for robust, one-shot hallucination detection,specifically designed for scenarios with limited data access, such asinteracting with black-box LLM APIs that typically expose only a few topcandidate log-probabilities per token. Our approach derives uncertaintyindicators directly from these readily available log-probabilities generatedduring non-greedy decoding. |
Charles Moslonka; Hicham Randrianarivo; Arthur Garnier; Emmanuel Malherbe; | arxiv-cs.CL | 2025-09-01 |
| 969 | TEQA: Temporal Knowledge Graph Enhanced Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Qian Liu; Siling Feng; Mengxing Huang; | Knowl. Based Syst. | 2025-09-01 |
| 970 | Iterative Caption Generation with Heuristic Guidance for Enhancing Knowledge-based Visual Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Fengyuan Liu; Zhongjian Hu; Peng Yang; Xingyu Liu; | Comput. Vis. Image Underst. | 2025-09-01 |
| 971 | Bio-inspired Product Design System Integrating Retrieval-augmented Question Answering and Semantic Fusion Diffusion Model Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Xinhui Kang; Wenjie You; Ying Luo; | Adv. Eng. Informatics | 2025-09-01 |
| 972 | Mimicking Human Attention in Driving Scenarios for Enhanced Visual Question Answering: Insights from Eye-tracking and The Human Attention Filter Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Kaavya Rekanar; Martin Hayes; Ciarán Eising; | Intell. Syst. Appl. | 2025-09-01 |
| 973 | TreeQA: Enhanced LLM-RAG with Logic Tree Reasoning for Reliable and Interpretable Multi-hop Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
XIANGRUI ZHANG et. al. | Knowl. Based Syst. | 2025-09-01 |
| 974 | Decomposing and Revising What Language Models Generate Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, the generatedquestions are often irrelevant and incomplete, resulting in a loss of facts inretrieval.These approaches also fail to aggregate evidence snippets fromdifferent documents and paragraphs. To tackle these problems, we propose a newfact decomposition-based framework called FIDES (\textit{faithful contextenhanced fact decomposition and evidence aggregation}) for attributed QA. |
ZHICHAO YAN et. al. | arxiv-cs.CL | 2025-08-31 |
| 975 | CaresAI at BioCreative IX Track 1 — LLM for Biomedical QA Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce a two-stage inference pipeline for precise short-answerextraction to mitigate verbosity and improve alignment with evaluation metrics.Despite partial improvements, challenges persist in generating strictlyformatted outputs. |
Reem Abdel-Salam; Mary Adewunmi; Modinat A. Abayomi; | arxiv-cs.CL | 2025-08-31 |
| 976 | CaresAI at BioCreative IX Track 1 – LLM for Biomedical QA Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Large language models (LLMs) are increasingly evident for accurate question answering across various domains. However, rigorous evaluation of their performance on complex … |
Reem Abdel-Salam; Mary Adewunmi; M. Abayomi; | ArXiv | 2025-08-31 |
| 977 | Geospatial Question Answering on Historical Maps Using Spatio-Temporal Knowledge Graphs and Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In thisproject, we developed a GeoQA system by integrating a spatio-temporal knowledgegraph (KG) constructed from historical map data with large language models(LLMs). |
Ziyi Liu; Sidi Wu; Lorenz Hurni; | arxiv-cs.IR | 2025-08-29 |
| 978 | DriveQA: Passing The Driving Knowledge Test Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this work, we present DriveQA, anextensive open-source text and vision-based benchmark that exhaustively coverstraffic regulations and scenarios. |
Maolin Wei; Wanzhou Liu; Eshed Ohn-Bar; | arxiv-cs.CV | 2025-08-29 |
| 979 | Benchmarking GPT-5 for Biomedical Natural Language Processing Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: The rapid expansion of biomedical literature has heightened the need for scalable natural language processing (NLP) solutions. While GPT-4 substantially narrowed the gap with … |
Yu Hou; Zaifu Zhan; Rui Zhang; | ArXiv | 2025-08-28 |
| 980 | Multilingual Visual Question Answering for Visually Impaired People Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Ratnabali Pal; Samarjit Kar; Dilip K. Prasad; A. Sekh; | Discover Artificial Intelligence | 2025-08-28 |
| 981 | AI-SearchPlanner: Modular Agentic Search Via Pareto-Optimal Multi-Objective Reinforcement Learning Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper, we propose\textbf{AI-SearchPlanner}, a novel reinforcement learning framework designed toenhance the performance of frozen QA models by focusing on search planning.Specifically, our approach introduces three key innovations: 1) Decoupling theArchitecture of the Search Planner and Generator, 2) Dual-Reward Alignment forSearch Planning, and 3) Pareto Optimization of Planning Utility and Cost, toachieve the objectives. |
Lang Mei; Zhihan Yang; Chong Chen; | arxiv-cs.AI | 2025-08-27 |
| 982 | AraHealthQA 2025: The First Shared Task on Arabic Health Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We introduce AraHealthQA 2025, the Comprehensive Arabic Health QuestionAnswering Shared Task, held in conjunction with ArabicNLP 2025 (co-located withEMNLP 2025). |
HASSAN ALHUZALI et. al. | arxiv-cs.CL | 2025-08-27 |
| 983 | From Search to Reasoning: A Five-Level Rag Capability Framework for Enterprise Data Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Retrieval-Augmented Generation (RAG) has emerged as the standard paradigm for answering questions on enterprise data. Traditionally, RAG has centered on text-based semantic search … |
Gurbinder Gill; Ritvik Gupta; Denis Lusson; Anand Chandrashekar; Donald Nguyen; | 2025 2nd IEEE/ACM International Conference on AI-powered … | 2025-08-27 |
| 984 | Context-Adaptive Synthesis and Compression for Enhanced Retrieval-Augmented Generation in Complex Domains Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: However, in complexdomains involving multiple, lengthy, or conflicting documents, traditional RAGsuffers from information overload and inefficient synthesis, leading toinaccurate and untrustworthy answers. To address this, we propose CASC(Context-Adaptive Synthesis and Compression), a novel framework thatintelligently processes retrieved contexts. |
Peiran Zhou; Junnan Zhu; Yichen Shen; Ruoxi Yu; | arxiv-cs.CL | 2025-08-26 |
| 985 | Chronological Passage Assembling in RAG Framework for Temporal Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Specifically, understanding narrative texts requires morethan isolated segments, as the broader context and sequential relationshipsbetween segments are crucial for comprehension. To address these limitations,we propose ChronoRAG, a novel RAG framework specialized for narrative texts.This approach focuses on two essential aspects: refining dispersed documentinformation into coherent and structured passages and preserving narrative flowby explicitly capturing and maintaining the temporal order among retrievedpassages. |
Byeongjeong Kim; Jeonghyun Park; Joonho Yang; Hwanhee Lee; | arxiv-cs.CL | 2025-08-26 |
| 986 | MQAD: A Large-Scale Question Answering Dataset for Training Music Large Language Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: To compileMQAD, our methodology leverages specialized Music Information Retrieval (MIR)models to extract higher-level musical features and Large Language Models(LLMs) to generate natural language QA pairs. |
ZHIHAO OUYANG et. al. | arxiv-cs.SD | 2025-08-26 |
| 987 | Extracting Information from Scientific Literature Via Visual Table Question Answering Models Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: This study explores three approaches to processing table data in scientificpapers to enhance extractive question answering and develop a software tool forthe systematic review process. |
Dongyoun Kim; Hyung-do Choi; Youngsun Jang; John Kim; | arxiv-cs.IR | 2025-08-26 |
| 988 | AVAM: Universal Training-free Adaptive Visual Anchoring Embedded Into Multimodal Large Language Model for Multi-image Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: In this paper,we propose a straightforward yet universal Adaptive Visual Anchoring strategy,which can be seamlessly integrated into existing MLLMs, offering significantaccuracy improvements through adaptive compression. |
Kang Zeng; Guojin Zhong; Jintao Cheng; Jin Yuan; Zhiyong Li; | arxiv-cs.CV | 2025-08-25 |
| 989 | Agri-Query: A Case Study on RAG Vs. Long-Context LLMs for Cross-Lingual Technical Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: We present a case study evaluating large language models (LLMs) with128K-token context windows on a technical question answering (QA) task. |
Julius Gun; Timo Oksanen; | arxiv-cs.CL | 2025-08-25 |
| 990 | ST-Raptor: LLM-Powered Semi-Structured Table Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Second, methods likeNL2Code and multi-modal LLM QA struggle to understand the complex layouts ofsemi-structured tables and cannot accurately answer corresponding questions. Tothis end, we propose ST-Raptor, a tree-based framework for semi-structuredtable question answering using large language models. |
ZIRUI TANG et. al. | arxiv-cs.AI | 2025-08-25 |
| 991 | CoCoA: Confidence and Context-Aware Adaptive Decoding for Resolving Knowledge Conflicts in Large Language Models Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Faithful generation in large language models (LLMs) is challenged byknowledge conflicts between parametric memory and external context. Existingcontrastive decoding methods tuned … |
Anant Khandelwal; Manish Gupta; Puneet Agrawal; | arxiv-cs.CL | 2025-08-25 |
| 992 | Can Out-of-Distribution Evaluations Uncover Reliance on Shortcuts? A Case Study in Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Despitetheir practicality, such evaluations build upon a strong assumption: that OODevaluations can capture and reflect upon possible failures in a real-worlddeployment. In this work, we challenge this assumption and confront the results obtainedfrom OOD evaluations with a set of specific failure modes documented inexisting question-answering (QA) models, referred to as a reliance on spuriousfeatures or prediction shortcuts. |
Michal Štefánik; Timothee Mickus; Marek Kadlčík; Michal Spiegel; Josef Kuchař; | arxiv-cs.CL | 2025-08-25 |
| 993 | PlantVillageVQA: A Visual Question Answering Dataset for Benchmarking Vision-Language Models in Plant Science Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Our objective remains to provide a publicly available, standardisedand expert-verified database to enhance diagnostic accuracy for plant diseaseidentifications and advance scientific research in the agricultural domain. |
Syed Nazmus Sakib; Nafiul Haque; Mohammad Zabed Hossain; Shifat E. Arman; | arxiv-cs.CV | 2025-08-23 |
| 994 | Distilcyphergpt: Enhancing Large Language Models for Knowledge Graph Question Answering in Cypher Through Knowledge Distillation Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
You Li Chong; Chin-Poo Lee; Ming Kim Lim; | Data Mining and Knowledge Discovery | 2025-08-23 |
| 995 | MedQARo: A Large-Scale Benchmark for Medical Question Answering in Romanian Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Question answering (QA) is an actively studied topic, being a core naturallanguage processing (NLP) task that needs to be addressed before achievingArtificial General Intelligence … |
ANA-CRISTINA ROGOZ et. al. | arxiv-cs.CL | 2025-08-22 |
| 996 | Elimination-based Reasoning with LLM for Multiple-choice Educational Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save |
Qianli Zhao; Meiru Zhang; | Journal of King Saud University Computer and Information … | 2025-08-22 |
| 997 | M3TQA: Massively Multilingual Multitask Table Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Existing multilingualtable benchmarks suffer from geolinguistic imbalance – overrepresenting certainlanguages and lacking sufficient scale for rigorous cross-lingual analysis. Toaddress these limitations, we introduce a comprehensive framework for massivelymultilingual multitask table question answering, featuring m3TQA-Instruct, alarge-scale benchmark spanning 97 languages across diverse language families,including underrepresented and low-resource languages. |
DAIXIN SHU et. al. | arxiv-cs.CL | 2025-08-22 |
| 998 | Hierarchical Vision-Language Reasoning for Multimodal Multiple-Choice Question Answering Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Multimodal Large Language Models (MLLMs) have demonstrated remarkablemultimodal understanding capabilities in Visual Question Answering (VQA) tasksby integrating visual and textual features. |
AO ZHOU et. al. | arxiv-cs.IR | 2025-08-22 |
| 999 | DSRAG: A Domain-Specific Retrieval Framework Based on Document-derived Multimodal Knowledge Graph Summary Related Papers Related Patents Related Grants Related Venues Related Experts View Save Abstract: Current general-purpose large language models (LLMs) commonly exhibit knowledge hallucination and insufficient domain-specific adaptability in domain-specific tasks, limiting … |
MENGZHENG YANG et. al. | ArXiv | 2025-08-22 |
| 1000 | PediatricsMQA: A Multi-modal Pediatrics Question Answering Benchmark Related Papers Related Patents Related Grants Related Venues Related Experts View Save Highlight: Evaluatingstate-of-the-art open models, we find dramatic performance drops in youngercohorts, highlighting the need for age-aware methods to ensure equitable AIsupport in pediatric care. |
Adil Bahaj; Oumaima Fadi; Mohamed Chetouani; Mounir Ghogho; | arxiv-cs.CY | 2025-08-22 |