arXiv:2609.13611v1 Announce Type: new
Abstract: The WMT26 General MT task evaluates systems on 10 language pairs that have no human references (neither translated from scratch nor post-edited from MT...
By Diptesh Kanojia, Chi-kiu Lo, Archchana Sindhujan, Samuel Larkin, Greg Hanneman, Alon Lavie
KnowBench is a new benchmark for clinical AI that measures Effort Reduction (ER), the proportion of system-generated clinical work product accepted by clinicians after expert and safety review. The metric is applied uniformly across various administrative tasks—visit notes, billing codes, orders, EHR summarization, patient summaries, and decision support—using the clinician’s review-and-attestation as ground truth. An initial deployment of Knowtex’s models achieved an aggregate ER of 97.99% across more than one million encounters in six months, with specialty-specific ER ranging from 96.8% to 98.9%.
By Jocelyn Kang, Caroline Zhang
The paper proposes new conditions for philosophers to engage with citizen deliberation in the AI era, focusing on how Large Language Models (LLMs) could support democratic processes such as citizen assemblies. It outlines the Democratic Commons project, an interdisciplinary effort that evaluates LLMs against five democratic principles, with a central concern about political bias and the democratic use of AI in experimental participatory settings. The study emphasizes the need for philosophical and political theory foundations to meaningfully assess AI’s role in democratic participation.
By Bernard Reber (CEVIPOF)
arXiv:2608.16344v3 Announce Type: replace
Abstract: Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and...
By Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Or\u{a}san, Chrysoula Zerva, Ricardo Rei, Fr\'ed\'eric Blain, Andr\'e F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
arXiv:2605.27840v2 Announce Type: replace-cross
Abstract: Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generatio...
By Zhisheng Zhang, Xiang Li, Yixuan Zhou, Jing Peng, Guoyang Zeng, Zhiyong Wu
arXiv:2609.13768v1 Announce Type: new
Abstract: Multi-hop question answering often fails when retrieval treats evidence as isolated matches to the original question, since the facts needed to answer...
By An Nguyen Phu, Dung Nguyen Quang, Luu Hieu An, Linh Ngo Van, Trung Le, Thien Huu Nguyen
The study investigates how different tokenization methods affect ECG Transformer models by comparing eight strategies across four backbone architectures on the CPSC2018 classification task. Physiology-aware tokenizations such as Median-beat and HeartLang achieve higher mean macro-AUCs (0.893 and 0.889) than point-wise and patch-wise approaches (0.822 and 0.824), while also reducing sequence length and training memory usage. Combining the two physiology-aware representations further improves macro-AUC by 8.2%.
By Jiawei Li, Fabio Bonassi, Johan Sundstr\"om, Thomas B. Sch\"on, Ant\^onio H. Ribeiro
The paper presents a system for the MedReason 2026 challenge that tackles both multiple‑choice and open‑ended medical visual question answering using offline, containerized inference. Key findings include that comparing answer semantics rather than labels boosts retrieval‑only accuracy from 20.0 % to 57.5 % on a 200‑case holdout, and that varying the number of in‑prompt retrieved examples has minimal impact on final accuracy (93.5 %–94.0 %). The final system achieves 94.0 % MCQ accuracy on the development set and 93.20 % on the official pre‑evaluation, far surpassing the off‑the‑shelf baseline.
"whyItMatters":"The results demonstrate that semantic‑aware retrieval and careful adapter tuning can dramatically improve medical VQA performance, offering a practical approach for high‑accuracy, offline inference in clinical settings."
By Tristan Kirscher (ICube, Institut Strauss), Niklas C. Koser (CAU), Soren Pirk (CAU)
arXiv:2604.27724v2 Announce Type: replace
Abstract: Medical retrieval-augmented generation (RAG) systems typically operate on text chunks extracted from biomedical literature, discarding the rich vis...
By Xupeng Chen, Binbin Shi, Chenqian Le, Jiaqi Zhang, Kewen Wang, Ran Gong, Jinhan Zhang, Chihang Wang
arXiv:2609.13154v1 Announce Type: new
Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et...
By Shamin Chokshi
The paper compares Complement Naive Bayes (NB) with zero‑shot and few‑shot large language models (LLMs) across a wide range of model sizes and text classification tasks. NB outperforms LLMs when labeled data is available, achieving comparable accuracy to large LLMs while running thousands of samples per second on a CPU. In zero‑data sentiment settings, LLMs still dominate, but NB remains the best choice for resource‑constrained HPC practitioners, and the authors provide a Kubernetes Helm operator to automate model selection.
By Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank W\"urthwein
arXiv:2609.14142v1 Announce Type: new
Abstract: Large language models (LLMs) can struggle with time-series question answering (TS-QA), especially when numerical signals are serialized as text and req...
By Ivan Delgado, Himansi Gupta, Bishal Khatri, Niharika Sapre, Lameta Shamoon, Onat Gungor, Tajana Rosing
arXiv:2609.15671v1 Announce Type: cross
Abstract: Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive and bandwidth-constrained settings. Fede...
By Md Khalid Syfullah, Alvi Ataur Khalil
arXiv:2509.19375v2 Announce Type: replace-cross
Abstract: Large language models are increasingly used for clinical text classification, where overconfident misclassifications can directly affect pati...
By Mridul Sharma, Adeetya Patel, Zaneta D' Souza, Samira Abbasgholizadeh Rahimi, Siva Reddy, Sreenath Madathil
arXiv:2508.02312v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine...
By Kang Chen, Xiuze Zhou, Yuanhui Yu, Yuanguo Lin, Hefeng Chen, Congyu Cai, Li Shen
FLoKD is an adaptive knowledge‑distillation framework designed for federated fine‑tuning of low‑rank LLMs over wireless networks. It transmits intermediate LoRA activations instead of full parameters or token‑level logits, and uses transformer block importance scoring plus dataset selection to reduce communication. Experiments on WikiText‑103, PTB, and Dialog show a 50‑65% reduction in communication while maintaining competitive perplexity.
By Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi
arXiv:2609.15224v1 Announce Type: new
Abstract: We present LOVER, a \underline{L}ong to sh\underline{O}rt \underline{V}ideo \underline{E}vidence \underline{R}einforced model for grounded question ans...
By Kaiyan Chen, Junbin Xiao, Xun Yang
arXiv:2603.18908v5 Announce Type: replace
Abstract: Independently trained language models often learn compatible late-stage representations, despite differences in training objectives, architectures,...
By Matt Gorbett, Suman Jana
arXiv:2609.08977v3 Announce Type: replace-cross
Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
By Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao
arXiv:2609.15137v1 Announce Type: cross
Abstract: 3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic f...
By Davit Soselia, Joseph JaJa, Amitabh Varshney