Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B mo...
Generative image models can produce highly realistic images, raising concerns about potential misuse in creating deceptive content. Current deepfake approaches face several challenges, including limit...
arXiv:2609.26347v1 Announce Type: cross
Abstract: The scarcity of non-English language data in specialized domains significantly limits the development of effective Natural Language Processing (NLP)...
By Julien Knafou, Luc Mottin, Ana\"is Mottaz, Alexandre Flament, Patrick Ruch
The paper introduces Geometry‑Aware Hyperbolic Residual Quantization, a method that adapts residual vector quantization to hyperbolic space while preserving its telescoping structure. It achieves this by using Hyperbolic Residual Aggregation in the forward pass and a discounted Hyperbolic Straight‑Through Estimator in the backward pass, thereby avoiding geometric inconsistencies and unstable gradients. Experiments on hierarchical prediction, recommendation, image tokenization, and neural audio coding demonstrate improved stability and structural organization of hyperbolic residual codes, with a noted trade‑off between compression efficiency and hierarchical organization.
By Alessio Colombo, Melika Ayoughi
The paper investigates how entropy over chain‑of‑thought tokens influences policy decisions such as gradient application, pruning, and collapse detection. By separating scaffold tokens from substantive content, the authors analyze entropy, Kullback–Leibler divergence, and entropy velocity for each channel, proving differences between raw and content conventions and bounding answer diversity. Empirical results across 23 configurations show that scaffold tokens can account for up to 41% of high‑entropy positions, with entropy share growing through distillation, while content conventions outperform raw surprisal on compression tasks and reveal significant answer leakage in re‑fed chains.
By Marios Papamichalis, Regina Ruane
The paper investigates how different action tokenization methods affect closed‑loop control in autoregressive vision‑language‑action models. It compares analytical, linear, and nonlinear representations, showing that lower reconstruction error does not guarantee better policy performance. The study highlights the need to evaluate tokenization on multiple criteria, including sequence predictability and decoder stability, rather than relying solely on reconstruction fidelity.
By Yuxin Yang, Gaohan He, Changxue Guan, Hangming Liu
The paper introduces a retrieved‑span training approach for query‑focused meeting summarization on the QMSum benchmark. By fine‑tuning a 406 M Fusion‑in‑Decoder model on 2,000‑word retrieved spans, the authors recover a 6.30 ROUGE‑1 loss incurred when moving from capped long input and achieve a test score of 36.33 ROUGE‑1, comparable to a larger 1.2 B system. The smaller model uses roughly one‑third the parameters and less than half the peak inference memory, while span‑regime fine‑tuning adds significant gains over the baseline.
"whyItMatters":"The study demonstrates that efficient, smaller models can match or exceed larger systems on QMSum using span‑based fine‑tuning, offering a practical path for scalable query‑focused meeting summarization."
By Edward Xi Yang (Ertas AI)
A new large-scale Arabic fact‑checking dataset called Arafa has been created using an automated pipeline that generates claims from Arabic Wikipedia, mutates them into counterfactuals, and validates them against supporting or refuting evidence. The dataset contains 181,976 claim‑evidence pairs labeled as supported, refuted, or not enough information, and human evaluation shows high inter‑annotator agreement and strong validation accuracy. Fine‑tuned transformer models on Arafa achieve a Macro F1‑score of 77%, demonstrating its usefulness for Arabic fact‑checking tasks.
By Christophe Khalil, Shady Elbassuoni, Rida Assaf
The paper argues that calibration—how well a language model’s confidence aligns with its actual correctness—should be a standard evaluation metric for large language models (LLMs). It notes that while calibration metrics exist, they are rarely applied outside specialized NLP subfields, leading to unverified confidence scores in new models, datasets, and benchmarks. The authors highlight the risks of miscalibration both at deployment (overconfident errors causing harm) and in research workflows (affecting LLM-as-a-judge, synthetic data generation, and active learning). They call for every NLP subfield to pair its primary performance metric with a calibration score, treating calibration as an essential property of every model.
By Mario Sanz-Guerrero, Katharina von der Wense
The paper introduces Knowledge Pull Requests (KPRs), a framework that enables continual document authoring by making each change interpretable. KPRs extract claims from new knowledge sources, filter and route them to appropriate sections, and flag conflicts with existing content, producing a ChangeLog that separates knowledge changes from textual edits. Experiments on revising Wikipedia and updating query-driven reports show that KPRs integrate more information, better preserve existing content, and improve question answering performance compared to rewriting from scratch or using frontier models with search.
By Alexander Martin, Benjamin Van Durme
Vision‑language models (VLMs) can lose accuracy when images are resized, even with minimal changes. The study shows that such small visual configuration changes—like tiling or token arrangement—cause more correctness flips across multiple checkpoints and benchmarks. Interestingly, in many cases the models still read the correct answer but fail to use it, and attention interventions reveal that configuration shifts weaken the use of readable information. By guiding models with field cues and their own transcriptions, the authors correct 97.2% of these errors.
By Dingyang Lin, Yingfeng Luo, Chenglong Wang, Chenwei Zhu, Anxiang Ma, Jingbo Zhu, Tong Xiao
The paper introduces a video summarization method that fuses facial embeddings, 3D body‑shape features, and visual appearance within a multi‑object tracking framework. By hierarchically assigning identities and using bidirectional anchoring, it robustly recovers trajectories even under heavy occlusion or low visual quality. Keyframes are selected through a multi‑factor weighting scheme that balances biometric clarity, social interaction, and motion dynamics, while Adaptive Non‑Maximum Suppression guarantees temporal diversity, resulting in a compact, identity‑centric summary.
By Milad Mirjalili, Enrique Alegre Guti\'errez, Eduardo Fidalgo Fern\'andez, V\'ictor Gonz\'alez Castro, Roc\'io Alaiz Rodr\'iguez, Manuel Castej\'on Limas
LLaVA‑Assessor is a unified large multi‑modal model (LMM) designed for visual quality assessment, combining image and video inputs. It introduces a two‑task framework—quality interpretation and quality scoring—supported by an adaptive architecture, a rigorous human‑annotated dataset, and a machine‑synthesized data expansion pipeline. The model employs a prompt‑disentanglement strategy to stabilize multi‑task training and achieves strong performance across 11 quality scoring test sets and 4 interpretation benchmarks.
By Ziheng Jia, Zicheng Zhang, Jiaying Qian, Guangtao Zhai, Xiongkuo Min
MultiViewDx is a physician‑validated multimodal instruction dataset that links medical imaging studies with patient context and normalizes heterogeneous reports into an evidence‑linked workflow (evidence → findings → differential discussion → diagnosis). The dataset covers a wide range of imaging modalities and uses a unified image‑text retriever to ensure that instruction synthesis is grounded in source‑supported evidence. Fine‑tuned models on MultiViewDx achieve the highest average accuracy on four MedVQA benchmarks and receive the strongest overall rating on JAMA Clinical Challenge cases, with ablations confirming the importance of case‑level multi‑view organization and evidence‑linked reasoning.
By Junda Wang, Zonghai Yao, Yujan Ting, Eric Z. Chen, Hieu Tran, Hong Yu, Weijing Huang, Terrence Chen
arXiv:2609.23735v2 Announce Type: new
Abstract: Scientific agents support a range of literature-based research tasks, such as retrieval, question answering, evidence-grounded generation, and claim as...
By ScholarSeed AI Team, Caoqinwei Gong, Xue Jiang, Wei Luo, Xiaoyu Qiu, Jiayi Sheng, Yi Wang, Zheng Yu, Ao Zhang, Haifan Zhang, Hanwei Zhang, Jihai Zhang, Yuan Cao, Wei Chen, Liyun Dai, Wenkai Fang, Guanglei Wang, Kai Ying, Tingyu Zhu, Wotao Yin
arXiv:2609.24324v1 Announce Type: new
Abstract: Electroencephalography (EEG) provides a non-invasive window into dynamic brain activity, yet modeling long-horizon EEG sequences remains challenging du...
By Weishan Ye, Yue Pan, Li Zhang, Gan Huang, Zhen Liang
arXiv:2609.24620v1 Announce Type: new
Abstract: Answering epidemiological questions from real-world clinical data requires medical coding, schema-aware SQL, and validation of implicit choices about p...
By Angelo Ziletti, Leonardo D'Ambrosi, Melanie Tuchardt, Tim Kondziella
arXiv:2609.22586v1 Announce Type: cross
Abstract: Modern audio-language models are no longer judged only on what words they can transcribe, but on whether they can reason over what they hear: recover...
By Harshit Rajgarhia, Asif Shaik, Rachuri Lokesh, Sushanta Kumar Pani, Abhishek Mukherji
arXiv:2609.25298v1 Announce Type: new
Abstract: Cultural evaluation coverage and robustness in language models are difficult to diagnose because pretraining corpora and cultural benchmarks are rarely...
By Yusser Al Ghussin, Eva Gavaller, Cristina Espa\~na-Bonet, Josef van Genabith, Simon Ostermann
arXiv:2609.25890v1 Announce Type: new
Abstract: Short-to-long training is a simple curriculum for speech models, but its gains can be difficult to interpret. In speech token language models, length-b...
By Hongjin Song, Runwu Shi, Weiqiao Shan, Jiale Luo, Yujin Wang, Yifei Wu, Chunxiang Jin