The paper benchmarks static embedding models—Word2Vec, FastText, and Doc2Vec—for detecting anomalous HTTP requests using a single‑class classification framework. It introduces HEDA, a modular pipeline that trains both embeddings and detectors solely on benign traffic in an unsupervised setting. Experiments on synthetic and real datasets show that FastText embeddings consistently yield high detection rates with controlled false positives.
By Amanda Riverol, Gustavo Betarte, Rodrigo Mart\'inez, \'Alvaro Pardo
The paper proposes using large language models (LLMs) to identify disagreements among models as a way to focus expert effort on revising codebooks for large‑scale text annotation. Three expert feedback methods are evaluated: editing LLM‑generated revisions (Codebook Verifying), answering questions about disagreements (Question Answering), and labeling disagreement cases with rationales (Rationale Labeling). Experiments on tutoring‑session transcripts show that Rationale Labeling achieves the highest LLM‑labeling accuracy (64.9%) compared to the expert‑revised codebook (57.8%), with Question Answering also outperforming the baseline (60.5%).
By Zeyu He, Zhuqian Zhou, Kirk Vanacore, Rene F. Kizilcec, Ting-Hao 'Kenneth' Huang
LabourCrew is a multi‑agent Retrieval‑Augmented Generation (RAG) framework designed for trustworthy statutory question answering in labour law. It introduces three grounding mechanisms: StatuteGraph, an evidence‑exchange ledger, and a calibrated trust gate that controls false‑accept rates. Evaluated on a Bangla Labour Act QA set, LabourCrew achieves a false‑accept rate of 0.081 and higher answer relevancy than existing RAG methods, demonstrating that calibrated abstention is key to auditable legal QA.
By Fatema Tuj Johora Faria, Mukaffi Bin Moin, Jubayer Al Mahmud, M. F. Mridha, Md. Alam Hossain
The paper introduces CASPER, a Chain-of-Thought Attribute-Specific Prompting framework designed to generate accurate, coherent, and attribute-relevant summaries of interrogative dialogues. It leverages structured prompting, iterative refinement, and a hierarchical evaluation mechanism called RoleEval to improve factual consistency and contextual completeness. The authors also present MINDSum, a new dataset of 6,000 annotated utterance pairs, and show that CASPER outperforms existing summarization models on ROUGE, BERTScore, and human expert evaluations.
By A Aditya Bhardwaj, Arjit Singh Arora, Md Shad Akhtar
LatWeave is a deterministic multi‑hop question‑answering framework that structures knowledge into a multidimensional lattice and reduces QA to three operators—meet, compare, and abstain—while limiting LLM use to extraction and planning. It achieves near‑lossless performance on complete knowledge benchmarks (e.g., MetaQA) and strong results on templated multi‑hop datasets (e.g., 2WikiMultihopQA), while transparently handling incomplete knowledge through abstention. The approach offers reproducible, auditable answer paths with no performance penalty within its operating envelope.
By Yuze Ren, Shaoheng Fan, Tao Wang, Yabo Yan, Han Han
The paper introduces the Large Knowledge Model (LKM), a scientific knowledge infrastructure that converts research literature into shared, computationally accessible reasoning graphs. LKM aligns questions, claims, and reasoning chains across papers, creating a Scientific Reasoning Landscape with Question, Workflow, and Evidence views. The system enhances scientific search, evidence‑grounded QA, and research planning, achieving notable accuracy gains on ChemBench, PubMedQA, and SciBench.
By Yuan Huang, Sihan Hu, Hongyu Gu, Chao Ma, Jiaxing Zhang, Zhiyong Zou, Caiyu Fan, Yan Xiao, Mingjun Xu, Chenyu Xie, Mingzhen Ju, Zhehao Ma, Qi Zhang, Baozong Wang, Yu Li, Zhiyuan Yao, Ruoxue Liao, Xinyu Li, Linfeng Zhang, Kun Chen, Weinan E
StateComp introduces a method for long‑horizon agents to decide when to compress historical interactions based on the current agent state, rather than relying on fixed windows or periodic schedules. The framework uses a two‑stage annotation process to create KEEP and READY labels, trains an imbalance‑aware router on frozen language model representations, and groups adjacent READY interactions into compact summaries. Experiments on WorkBuddyBench show that StateComp cuts agent and summarization tokens by 52.27% and speeds up representation extraction 12.67‑fold while preserving task performance.
By Mingxuan Wang, Hongyue Chen, Yinglong Guo, Fei Luo, Chao Ning, Bo Wang, Guorun Yao, Yanbiao Ma, Jungong Han
The study investigates how data volume, model size, and training duration affect the performance of fMRI foundation models. Using over 200 datasets and 10,000 GPU‑hours, the authors find that larger models benefit more from additional data, and that at a fixed compute budget, increasing data yields greater gains than enlarging the model. By selecting optimal combinations of data, size, and duration, they produce models that outperform existing fMRI foundation models on out‑of‑distribution tasks while requiring less pretraining compute.
By Wenhao Ye, Xuanye Pan, Junfeng Xia, Junxiang Zhang, Mo Wang, Quanying Liu
LexLattice is an extractive summarizer that models a legal act’s hierarchy as a two‑dimensional semantic lattice and consolidates information over it using a masked 2D neural cellular automata before selecting content. The method achieves state‑of‑the‑art ROUGE scores across all 24 languages of EUR‑Lex‑Sum in both multilingual and cross‑lingual settings, outperforming large instruction‑tuned baselines while using only a 1.8 M‑parameter consolidator on a frozen multilingual encoder. A consolidator trained on high‑resource languages transfers almost losslessly to unseen languages, suggesting the model operates on language‑agnostic semantic geometry rather than surface form.
By Sujay Uday Rittikar, Sheela Ramanna
arXiv:2603.09222v2 Announce Type: replace
Abstract: Efficient context compression is critical for retrieval-augmented question answering in resource-constrained settings, where long retrieved context...
By Thao Do, Dinh Phu Tran, An Vo, Seon Kwon Kim, Daeyoung Kim
The paper investigates how small language models (SLMs) perform in knowledge graph question answering (KGQA) when evaluated on the reasoning paths they take, rather than just the final answer. Using the THESEUS navigation and traceability framework, the authors test frozen, off‑the‑shelf SLMs as local action policies that choose graph actions and decide when to stop, without any task‑specific training or free‑form answer generation. By measuring both Hits@1 and Path Edit Distance (PED) across the Kinship and MQuAKE‑ST datasets, the study finds that models vary significantly in both answer accuracy and path fidelity, and that prompting can either help or hurt navigation depending on the model.
"whyItMatters":"The results show that evaluating SLMs solely on endpoint accuracy can be misleading, highlighting the need to assess reasoning path fidelity in KGQA tasks."
By Eduin E. Hernandez, Sergio A. Diaz, Luis F. Garcia, Nurassyl Askar, Stefano Rini
The paper introduces a probabilistic‑circuit framework for fusing opinions from multiple black‑box experts in noisy, conflict‑prone environments. It dynamically assigns context‑specific credibility to each expert, allowing reliable aggregation without needing access to their internal models or retraining. Experiments on multiple‑choice question answering with large language models show that this method outperforms individual models and static ensemble baselines, consistently improving predictive accuracy and decision reliability under disagreement.
By Pranuthi Tenali, Sahil Sidheekh, Saurabh Mathur, Vijayalakshmi Saravanan, Erik Blasch, Kristian Kersting, Sriraam Natarajan
arXiv:2609.28007v1 Announce Type: cross
Abstract: Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain docume...
By Imtiaz Ul Hassan, \"Oyk\"u Akbulut, Onur Kaya, Ardhendu Behera, Swagat Kumar, Peter Matthew, Yonghuai Liu
arXiv:2609.27197v1 Announce Type: new
Abstract: Minimum Risk Training (MRT) enables neural machine translation models to directly optimize sequence-level evaluation metrics instead of relying only on...
By Hung Phan, Waqwoya Abebe, Youssef Hussein, Supriya Chinthavali, Dalton Lunga, Ali Jannesari
arXiv:2609.26820v1 Announce Type: cross
Abstract: Physiological time series such as electrocardiograms (ECG) and electroencephalograms (EEG) exhibit complex temporal structure, substantial acquisitio...
By Naser Mansour, Sidahmed Benabderrahmane, Ameer Rahwan
arXiv:2609.27784v1 Announce Type: cross
Abstract: Semi-structured documents are ubiquitous in scientific reports, financial statements, and technical manuals. Question answering over such documents r...
By Teng Lin, Yuyu Luo, Nan Tang
arXiv:2609.28117v1 Announce Type: cross
Abstract: In this paper, we introduce a gradient-based head attribution strategy where the Token-level Max-Margin loss is backpropagated to the attention maps....
By Pawe{\l} M\k{a}ka, Yusuf Can Semerci, Jan Scholtes, Gerasimos Spanakis
arXiv:2609.28222v1 Announce Type: new
Abstract: Unified 3D vision-language systems must combine complementary geometry, scale, and appearance cues while supporting tasks from instance segmentation to...
By Xueqi Qiu, Xingyu Miao, Jingjing Deng, Haoran Duan, Yang Long, Ling Shao
arXiv:2609.27233v1 Announce Type: new
Abstract: Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adja...
By Zixuan Lan, Jessica Yang, Yanhong Li, Karen Livescu, Jiawei Zhou
The article "From Words to Vectors: What Happens in Between?" explores the process of converting textual data into numerical representations, focusing on techniques such as TF-IDF and vector space models. It discusses how these representations enable text classification tasks and provides a practical overview of the underlying concepts. The piece serves as a guide for readers interested in the mechanics of text preprocessing and feature extraction for machine learning.
By Nikhil Dasari