arXiv:2609.28007v1 Announce Type: cross
Abstract: Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain docume...
By Imtiaz Ul Hassan, \"Oyk\"u Akbulut, Onur Kaya, Ardhendu Behera, Swagat Kumar, Peter Matthew, Yonghuai Liu
Most Turkish-capable large language models (LLMs) are evaluated using general-purpose benchmarks rather than long, structurally complex domain documents. This paper evaluates five open-weight 7B-8B mo...
arXiv:2603. 29002v3 Announce Type: replace-cross Abstract: Modern large language models (LLMs) increasingly depends on efficient long-context processing and generation mechanisms, including sparse attention, retrieval-augmented generation (RAG), and compressed contextual memory, to support complex reasoning.
By Zifan He, Rui Ma, Yizhou Sun, Jason Cong
Orthrus is a serving system that performs heterogeneous batching of embedding and generative models within a single inference loop. It uses chunked embedding with incremental pooling and workload‑aware batch composition to unify conflicting computational patterns. Experiments on four A100 GPUs show that Orthrus improves throughput by 1.28×–4.52× and reduces p99 latency by up to 55.8% compared to baseline deployments.
By Dohyun Park, Hubertus Franke, Daniel G. Waddington, Swaminathan Sundararaman, Yongjoo Park
The paper reports on building a retrieval‑augmented legal assistant for Uzbek that operates in both a managed cloud service and an on‑premises deployment. It introduces two new domain benchmarks—one for retrieval and one for end‑to‑end QA—and shows that fine‑tuning an open‑weight text embedder (UTE‑1) can close the performance gap with proprietary models under tight cost and latency constraints. The authors also provide negative results for a QLoRA experiment and release the benchmarks, evaluation code, and the fine‑tuned embedder for future low‑resource legal NLP work.
By Tatul Danielyan, Mariam Avetisyan, Hrant Davtyan
arXiv:2608. 08020v1 Announce Type: new Abstract: Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from \emph{how much} compute to spend, to \emph{where} to allocate it.
By Lijie Yang, Hongyin Luo, Tri Dao, Ravi Netravali