arXiv:2608.30935v1 Announce Type: cross
Abstract: Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embod...
By Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu, Xiaoyang Wang, Yueyu Wang, Qianli Ma, Fan Yang, Ran Mei, Jia Wei, Jiangpeng Hu, Xuhao Liu, Hongming Chen, Yuanbin Shao, Yiyang Lin, Ziliang Li, Liang Pan, Xinhang Liu, Yuntao Ma, Tingxiang Fan
arXiv:2608.30614v1 Announce Type: new
Abstract: Organizing unstructured feedback text into hierarchical taxonomy is a fundamental challenge in NLP, particularly in domains where feedback arrives at m...
By Sandeep Sricharan Mukku, Albert Aristotle Nanda, Rohit Pyati
arXiv:2608.30163v1 Announce Type: cross
Abstract: Real-world knowledge resides in multimodal documents, necessitating retrieval-augmented generation (RAG) for accurate question answering. However, ex...
By Ruofan Hu, Shengyang Xu, Minjie Hong, Xiaoda Yang, Sashuai Zhou, Ke Lei, Tao Jin, Zhou Zhao
arXiv:2608.31142v1 Announce Type: cross
Abstract: The 2025--2026 AI market has seen a wave of stealth releases: frontier models launched anonymously on developer platforms under codenames. For their...
By Yisen Xi
arXiv:2608.30679v1 Announce Type: cross
Abstract: Large Reasoning Models produce Long Chains-of-Thought (LCoTs) which involve breaking down the problem into smaller reasoning steps before reaching th...
By B\'er\'enice Jaulmes, Mehwish Alam
The paper presents a rigor‑matched audit comparing two periodic‑step, search‑based layer‑skipping methods for efficient large language model inference: a confidence‑gated early‑exit baseline (ConfLayers) and a self‑speculative decoding approach (SWIFT). Across two Qwen2.5 model scales and tasks (GSM8K reasoning and CNN/DailyMail summarization), SWIFT consistently outperforms ConfLayers in accuracy and, after separating search overhead, achieves faster true inference speed in most settings. The study also evaluates two trained‑routing methods (LayerRoute and LayerDrop), finding modest speedups but significantly lower accuracy, especially for LayerRoute on GSM8K at 1.5B.
By Prateek Kumar Sikdar, Arpan Ghosh
Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.
By Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto
arXiv:2608.31038v1 Announce Type: new
Abstract: Incremental Named Entity Recognition (INER) stands as a pivotal task in information extraction, emphasizing the successive identification of new entity...
By Duzhen Zhang, Yahan Yu, Xiuyi Chen, Chenxing Li, Dong Yu
arXiv:2511.22707v2 Announce Type: replace-cross
Abstract: In web environments, user preferences are often refined progressively as users move from browsing broad categories to exploring specific item...
By Tianxin Wei, Xuying Ning, Xuxing Chen, Ruizhong Qiu, Yupeng Hou, Yan Xie, Shuang Yang, Zhigang Hua, Jingrui He
arXiv:2605.06582v4 Announce Type: replace
Abstract: Modern learning systems represent perceptual signals with continuous vectors, but comparison, retrieval, memory, alignment, and reasoning are often...
By Adhiraj Banerjee, Vipul Arora
The paper compares generative and encoder-based neural models for multilingual Named Entity Recognition (NER) across the eleven languages of the Naamapadam benchmark. Five classic model families, four decoder-only large language models fine‑tuned with LoRA and 4‑bit NF4 quantisation, and nine generative models in zero‑to‑5‑shot inference were evaluated under strict CoNLL span‑level metrics. Encoder-based models (mBERT and XLM‑R) achieved substantially higher F1 scores—up to 0.675 on Hindi—than any generative architecture, with gaps of 7.5–40 percentage points; the best few‑shot result reached only 28% of the encoder baseline. The study identifies three language clusters (encoder‑dominant, partial‑coverage, and failure‑zone) and offers deployment guidelines based on transfer learning and low‑resource NLP principles.
By Jakkala Mahesh, Jatavath Shravan Kumar, Komalla Shivani, Sujoy Sarkar
arXiv:2608.30092v1 Announce Type: cross
Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single...
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
arXiv:2606.16206v2 Announce Type: replace
Abstract: Large language models are increasingly proposed as educational tutors, yet stronger task-solving ability does not necessarily imply stronger learni...
By Junyi Yao, Zihao Zheng, Baichuan Li
arXiv:2608.29890v1 Announce Type: new
Abstract: Biomedical Named Entity Recognition (NER) is fundamental to healthcare AI applications, including clinical decision support and medical information ext...
By Nhu Vo, Phuong Nguyen, Nu Uyen Phuong Le, Inigo Jauregi Unanue, Dung D. Le, Massimo Piccardi, Wray Buntine
arXiv:2502.00857v2 Announce Type: replace
Abstract: Large Language Models (LLMs) increasingly provide direct answers to user questions, raising concerns about reduced engagement in critical thinking...
By Jamshid Mozafari, Bhawna Piryani, Abdelrahman Abdallah, Adam Jatowt
arXiv:2608.30425v1 Announce Type: new
Abstract: Cross-lingual aspect-based sentiment analysis (ABSA) transfers knowledge from a source language with annotated data to a target language, enabling fine...
By Jakub \v{S}m\'{i}d, Pavel P\v{r}ib\'{a}\v{n}, Pavel Kr\'{a}l
arXiv:2608.30297v1 Announce Type: new
Abstract: Attributes describing data content and context can induce diverse imbalance patterns that go beyond label imbalance alone. However, existing studies pr...
By Hanshu Rao, Guangzeng Han, Xiaolei Huang
arXiv:2601.12868v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) increasingly operate in high-stakes settings where demographic attributes such as race and ethnicity may be expl...
By Shiyue Hu, Ruizhe Li, Yanjun Gao
IndicQE-APE is a consolidated benchmark that unifies quality estimation (QE) and automatic post‑editing (APE) data for nine Indic language pairs, comprising 126,754 instances with multiple aligned labels such as direct assessment, human post‑edit, word‑level OK/BAD tags, and error explanations. The dataset includes a stratified test set across four difficulty axes and supports training and evaluation of six prompted large language models, three COMET metrics, and three APE systems. Experiments reveal that segments with conflicting holistic and token‑level quality signals are consistently ranked lower, while annotator disagreement shows no effect when controlled for score distribution.
whyItMatters":"The benchmark provides a unified resource for training and evaluating QE and APE across Indic languages, enabling consistent comparison of models and metrics on a shared dataset."
By Diptesh Kanojia, Archchana Sindhujan, Sourabh Deoghare, Daria Sokova, Shenbin Qian, Girish Koushik, Tharindu Ranasinghe, Constantin Or\u{a}san, Chrysoula Zerva, Ricardo Rei, Fr\'ed\'eric Blain, Andr\'e F. T. Martins, Marco Turchi, Matteo Negri, Anoop Kunchukuttan, Mitesh M. Khapra, Pushpak Bhattacharyya
arXiv:2608.30475v1 Announce Type: cross
Abstract: We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, cove...
By Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir, Md Arid Hasan, Basel Mousi, Nadir Durrani, Fahim Dalvi, Zien Sheikh Ali, Erchin Serpedin, Hasan Kurban, Mustafa Jarrar, Shammur Absar Chowdhury, Firoj Alam