arXiv:2609.14467v1 Announce Type: cross
Abstract: Automating clinical documentation from long-form doctor-patient conversations remains challenging for modern audio-language models. While cascaded AS...
By Ziyu Zhang, Mingchen Shao, Wenjie Tian, Tianlun Zuo, Longhao Li, Lei Xie
arXiv:2609.14770v1 Announce Type: cross
Abstract: Generalisations are common in scientific communication, even though they are semantically ambiguous. An automated method is needed to identify and ca...
By Chenxin Diao, Nataliya Stepanova, Emily Allaway
arXiv:2508.02312v2 Announce Type: replace-cross
Abstract: Large Language Models (LLMs), now a foundation in advancing natural language processing, power applications such as text generation, machine...
By Kang Chen, Xiuze Zhou, Yuanhui Yu, Yuanguo Lin, Hefeng Chen, Congyu Cai, Li Shen
The paper proposes new conditions for philosophers to engage with citizen deliberation in the AI era, focusing on how Large Language Models (LLMs) could support democratic processes such as citizen assemblies. It outlines the Democratic Commons project, an interdisciplinary effort that evaluates LLMs against five democratic principles, with a central concern about political bias and the democratic use of AI in experimental participatory settings. The study emphasizes the need for philosophical and political theory foundations to meaningfully assess AI’s role in democratic participation.
By Bernard Reber (CEVIPOF)
The paper compares Complement Naive Bayes (NB) with zero‑shot and few‑shot large language models (LLMs) across a wide range of model sizes and text classification tasks. NB outperforms LLMs when labeled data is available, achieving comparable accuracy to large LLMs while running thousands of samples per second on a CPU. In zero‑data sentiment settings, LLMs still dominate, but NB remains the best choice for resource‑constrained HPC practitioners, and the authors provide a Kubernetes Helm operator to automate model selection.
By Mohammad Firas Sada, Dmitry Mishin, John Graham, Seungmin Kim, Mahidhar Tatineni, Frank W\"urthwein
KnowBench is a new benchmark for clinical AI that measures Effort Reduction (ER), the proportion of system-generated clinical work product accepted by clinicians after expert and safety review. The metric is applied uniformly across various administrative tasks—visit notes, billing codes, orders, EHR summarization, patient summaries, and decision support—using the clinician’s review-and-attestation as ground truth. An initial deployment of Knowtex’s models achieved an aggregate ER of 97.99% across more than one million encounters in six months, with specialty-specific ER ranging from 96.8% to 98.9%.
By Jocelyn Kang, Caroline Zhang
The paper presents a system for the MedReason 2026 challenge that tackles both multiple‑choice and open‑ended medical visual question answering using offline, containerized inference. Key findings include that comparing answer semantics rather than labels boosts retrieval‑only accuracy from 20.0 % to 57.5 % on a 200‑case holdout, and that varying the number of in‑prompt retrieved examples has minimal impact on final accuracy (93.5 %–94.0 %). The final system achieves 94.0 % MCQ accuracy on the development set and 93.20 % on the official pre‑evaluation, far surpassing the off‑the‑shelf baseline.
"whyItMatters":"The results demonstrate that semantic‑aware retrieval and careful adapter tuning can dramatically improve medical VQA performance, offering a practical approach for high‑accuracy, offline inference in clinical settings."
By Tristan Kirscher (ICube, Institut Strauss), Niklas C. Koser (CAU), Soren Pirk (CAU)
arXiv:2609.14795v1 Announce Type: cross
Abstract: There is a tradeoff in machine translation meta-evaluation between prioritizing alignment with adequacy versus fluency. The balance depends on the co...
By Behzad Shayegh, Niloofar Kazemi
MoVT is a new framework for text‑to‑motion generation that uses a cross‑modal augmented motion tokenizer to project 3D motion tokens into 2D, enriching the motion codebook with real‑world video patterns. The enriched tokens are mapped back to 3D, creating aligned 3D and 2D codebooks that better capture intricate motions. These codebooks feed a generative masked transformer, which predicts masked motion tokens in a modality‑agnostic way, allowing text‑index pairs from the 2D codebook and annotated videos to further improve generation quality. Empirical tests show MoVT outperforms previous state‑of‑the‑art methods on several key metrics.
By Beibei Jing, Tianle Guo, Youjia Zhang, Zikai Song, Yawei Luo, Junqing Yu, Tao Guan, Wei Yang
FLoKD is an adaptive knowledge‑distillation framework designed for federated fine‑tuning of low‑rank LLMs over wireless networks. It transmits intermediate LoRA activations instead of full parameters or token‑level logits, and uses transformer block importance scoring plus dataset selection to reduce communication. Experiments on WikiText‑103, PTB, and Dialog show a 50‑65% reduction in communication while maintaining competitive perplexity.
By Xinlu Zhang, Na Yan, Yang Su, Yansha Deng, Toktam Mahmoodi
FedV-KGQA addresses multi‑hop question answering over vertically partitioned knowledge graphs where each silo holds disjoint relation types. The system trains local embeddings, concatenates silo‑specific entity views, anchors questions at a topic entity, and ranks candidates without sharing raw triples. Experiments show federated fusion nearly matches centralized accuracy, that anchoring and enrichment are more critical than embedding choice, and that the cheapest encoder depends on target accuracy.
By Md Saikat Islam Khan Bappy, Oshani Seneviratne
arXiv:2609.15671v1 Announce Type: cross
Abstract: Visual Question Answering (VQA) with Vision-Language Models (VLMs) is increasingly used in privacy-sensitive and bandwidth-constrained settings. Fede...
By Md Khalid Syfullah, Alvi Ataur Khalil
arXiv:2609.13648v1 Announce Type: new
Abstract: Solar energy decision support is fragmented across dashboards that provide data without explanation, research papers are slow to parse, and general-pur...
By Jyotsna Singh
arXiv:2609.14422v1 Announce Type: new
Abstract: Agentic Network Operations (NetOps) are an emerging paradigm promising to enable workload-aware, self-adjustable, and reliable autonomous networks. Whi...
By Tobias Labarta, Frederik Pahde, Novak Boskov, Maximilian Dreyer, David Birkenberger, Manzoor Ahmed Khan, Sebastian Lapuschkin, Wojciech Samek
arXiv:2609.15224v1 Announce Type: new
Abstract: We present LOVER, a \underline{L}ong to sh\underline{O}rt \underline{V}ideo \underline{E}vidence \underline{R}einforced model for grounded question ans...
By Kaiyan Chen, Junbin Xiao, Xun Yang
arXiv:2609.14790v1 Announce Type: new
Abstract: Detecting video highlights, the most informative or engaging moments in a video, is important for applications such as video summarization and content...
By Michal Byra, Alberto Presta, Grzegorz Stefanski, Krzysztof Arendt
arXiv:2609.15137v1 Announce Type: cross
Abstract: 3D Gaussian language fields provide an explicit, spatially grounded representation for 3D visual question answering (VQA), but their dense semantic f...
By Davit Soselia, Joseph JaJa, Amitabh Varshney
arXiv:2609.08977v3 Announce Type: replace-cross
Abstract: In this work, we present Gander, a native multimodal duplex interaction model that builds on MiniCPM-o 4.5 and is further adapted for realtim...
By Orantqing, Shengpeng Ji, Junlong Tong, Jialong Zuo, Dongjie Fu, Di Cao, Yangzhuo Li, Shangda Wu, Franz, Evan, Theron Veyra, Changhao Pan, Jingyu Lu, Dongchao Yang, Zhifei Xie, Yang Tan, Xiaoyu Shen, Xiaoda Yang, Wenfu Wang, Teddy Sun, Steve Yves, Zhou Zhao
arXiv:2606.20980v2 Announce Type: replace-cross
Abstract: As Self-Driving Cars continue to expand internationally and use multi-modal systems such as VLMs as a cognitive backbone for their Action mod...
By Adrian Cespedes, Marcelo Chincha, Dunant Cusipuma, Victor Flores-Benites, David Ortega, Arturo Deza
arXiv:2609.13814v1 Announce Type: new
Abstract: Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and li...
By Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou, Jincenzi Wu, Xujie Song, Kexin Shi, Zihang Liu, Pengcheng Zhu, Jiayi Zhou, Baoyue Zhang, Changhao Zhang, Zitong Wang, Jinhong Wang, Tong Niu, Jingjing Liu, Junan Lin, Haolin He, Hengshuo Chu, Yuhui Chen, Jian Liu, Yuge Huang, Junliang Xing, Yuntao Wang, Weiqiang Wang, Chun Yu, Yuanchun Shi