Question-Guided Evidence Acquisition for Multimodal Visual Question Answering
arXiv:2608. 19739v1 Announce Type: cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably.
Classical and neural NLP: translation, question answering, tokenization and the evaluation of language understanding.
arXiv:2608. 19739v1 Announce Type: cross Abstract: Multimodal LLMs can see a document, but they often can't read it reliably.
arXiv:2608. 20169v1 Announce Type: cross Abstract: We present a novel approach to efficient LLM agent harness optimization through adaptive validation task selection.
arXiv:2608. 19971v1 Announce Type: new Abstract: Multimodal sentiment analysis aims to infer affective states by integrating language, visual, and acoustic cues.
arXiv:2608. 19201v1 Announce Type: cross Abstract: Bioinformatics software and databases are essential components of modern life science research, yet their mentions in the scientific literature are often inconsistent and difficult to systematically identify at scale.
arXiv:2608. 19558v1 Announce Type: new Abstract: Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, while standard F1 scores do not indicate which predictions remain safe to automate when the input distribution changes.
arXiv:2608. 19297v1 Announce Type: new Abstract: While multimodal large language models (MLLMs) excel in medical applications, most of them favor static images or short-term signals.
arXiv:2602. 18400v3 Announce Type: replace-cross Abstract: Missing data problems, such as missing modalities in multi-modal brain MRI and missing slices in cardiac MRI, pose significant challenges in clinical practice.
arXiv:2509. 05441v4 Announce Type: replace-cross Abstract: Latent generative models compress images into learned embeddings prior to synthesis, and the generation quality critically depends on how faithfully these embeddings preserve visual detail.
arXiv:2512. 07847v2 Announce Type: replace Abstract: Benchmarking has been the cornerstone of progress in computer vision, natural language processing, and the broader deep learning domain, driving algorithmic innovation through standardized datasets and reproducible evaluation protocols.
arXiv:2607. 14497v2 Announce Type: replace Abstract: Egocentric Visual Question Answering (VQA) has attracted widespread attention as an important task for enabling Multimodal Large Language Models (MLLMs) to interact with the real world.
arXiv:2608. 19802v1 Announce Type: new Abstract: LLM evaluation pipelines often have many candidate judges: general LLM-as-a-judge prompts, reward models, safety classifiers, confidence variants, and task-specific verifiers.
arXiv:2608. 18752v2 Announce Type: replace-cross Abstract: Statutory retrieval is necessary for citation-grounded legal question answering, but remains underexplored for Greek.
arXiv:2604. 13731v2 Announce Type: replace Abstract: Multi-page Document Visual Question Answering requires reasoning over semantics, layouts, and visual elements in long, visually dense documents.
arXiv:2608. 19662v1 Announce Type: new Abstract: Agentic language models repeatedly encode tool and skill schemas that recur across requests in different combinations and orders, preventing standard prefix caching from reusing their key--value (KV) states.
arXiv:2608. 20281v1 Announce Type: cross Abstract: Large language models often fail to answer questions about a bounded document collection when the source documents are not retrieved at inference time.
arXiv:2608. 20204v1 Announce Type: new Abstract: Legal work, with its heavy reliance on processing large amounts of text, is often considered one of the domains most exposed to the use of LLMs.
arXiv:2501. 16106v2 Announce Type: replace Abstract: Recent advances in multimodal depression recognition for clinical interviews (MDRC) have demonstrated the potential of AI systems by integrating textual, acoustic, and facial cues.
arXiv:2608. 20122v1 Announce Type: new Abstract: Large multimodal models (LMMs) have demonstrated strong OCR recognition capabilities, yet remain vulnerable to adversarial visual text that is readable to humans but challenging for models to localize and recognize.
arXiv:2608. 19526v1 Announce Type: cross Abstract: Stock market analysts and investors face a daily challenge: too much financial news, too little time.
arXiv:2608. 19957v1 Announce Type: new Abstract: Natural language code retrieval is a rapidly evolving task in computer science.