QVAC Genesis III is a 191.43 B‑token synthetic STEM corpus covering 19 domains and multiple difficulty levels, created through a dual generation strategy that uses a weak edge‑scale student model to generate corrective explanations and contrastive reasoning. The authors evaluate the corpus with an LLM‑as‑a‑parser protocol and demonstrate that 1.7 B‑parameter models trained on QVAC Genesis III outperform those trained on Cosmopedia‑v2 and the Cosmo‑1B model on ARC, GPQA Diamond, and MMLU STEM benchmarks, achieving up to +28.57% improvement on ARC‑E and a 99.45% valid answer rate.
By Davide Vitabile, N. Ranjan, Akshay Nambiar, Kamal K. Gupta, Amril Nazir
arXiv:2609.13154v1 Announce Type: new
Abstract: Recent advances in large language models (LLMs) have made prompts increasingly large and complex. Techniques such as chain-of-thought reasoning (Wei et...
By Shamin Chokshi
arXiv:2609.36965v1 Announce Type: new
Abstract: System One models such as Jev offer an efficient alternative to generative language models for tasks that require decisions rather than open-ended resp...
By Zexiao Wang, Zihao Zhang, Xudong Wang, Pan Wang, Ziyi Ye, Haoyu Zhao, Zuxuan Wu, Shuicheng Yan
The study evaluates how well pre‑trained models can classify the Bloom level of AI‑generated educational questions, a task that is crucial for ensuring pedagogical quality. Traditional machine‑learning models perform poorly on out‑of‑distribution data, whereas transformer and large‑language models achieve higher accuracy, especially after feature‑engineering techniques such as text splicing and appending learning objectives. Retraining the models yields the most significant performance gains across all datasets.
By Michael Lawrence Castanares, Princess Ventures, Allan Tan
arXiv:2606. 01982v1 Announce Type: new Abstract: Schema-constrained information extraction from diverse educational and labor-market corpora remains an open challenge in natural language processing because existing pipelines rely primarily on lexical-surface methods that cannot recover implicit competencies, lack grounding in shared taxonomies, and provide no formal measures of extraction reliability or document-level completeness.
By Sherzod Turaev, Mary John, Mamoun Awad, Nazar Zaki, Khaled Shuaib
arXiv:2510.13935v3 Announce Type: replace-cross
Abstract: The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which ne...
By Kenan Alkiek, David Jurgens, Vinod Vydiswaran
OptSkills is an archetype‑centric agent that learns and reasons about optimization problems using large language models. It clusters problems by underlying archetypes, explores diverse modeling and solver configurations within each cluster, and distills successful trajectories into reusable workflow‑level skills. The system achieves state‑of‑the‑art accuracy on multiple datasets, outperforming prior methods on challenging benchmarks such as MIPLIB‑NL and OOD NLCO.
By Haochen Yang, Ke Zhao, Mengyuan Ma, Xingyu Lu, Xiangfeng Wang, Hong Qian
The study examines whether domain‑adaptive continued pretraining (DAPT) on a learner‑writing corpus (EFCAMDAT) can enhance transformer‑based automated essay scoring (AES) for English proficiency tests. Researchers applied DAPT to BERT, RoBERTa, and DistilBERT and compared the adapted models with their original checkpoints on the FCE and IELTS datasets, evaluating both in‑domain scoring and few‑shot cross‑dataset transfer. Results show that full‑corpus DAPT yields mixed effects, while proficiency‑specific DAPT often outperforms full‑corpus DAPT and sometimes even the non‑adapted baseline, though benefits vary by proficiency composition and encoder architecture and do not consistently transfer across tests.
By Duy Anh Nguyen
arXiv:2607. 04223v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) reduces but does not eliminate hallucination, and existing detectors return a single answer-level score that does not indicate which sentence is unsupported, or why.
By Mohamed Aly Bouke
arXiv:2606. 11552v1 Announce Type: cross Abstract: Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation.
By Lexington Whalen, Yuki Ito, Ryo Sakamoto
Large language models (LLMs) achieve strong relation extraction (RE), but their computational demands and reliance on proprietary APIs limit deployment in resource-constrained or privacy-sensitive settings. We investigate how far small language models (SLMs) can close this gap across general-domain and literary text.
Large language models (LLMs) achieve remarkable performance across a wide range of tasks, but their autoregressive decoding process incurs substantial inference costs due to inherently sequential token generation. Speculative decoding addresses this bottleneck by employing a lightweight draft model to propose multiple future tokens that are subsequently verified in parallel by a larger target model.