arXiv:2609.14829v1 Announce Type: cross
Abstract: We introduce Enemray, a Hassaniya-centric language model that enables general-purpose interaction in Hassaniya. Enemray is trained around a stability...
By Cheikh Ahmed
arXiv:2605.14322v4 Announce Type: replace
Abstract: Language agents are increasingly deployed in professional workflows, yet tutoring remains a high-stakes capability that existing evaluations only p...
By Zixin Chen, Peng Liu, Rui Sheng, Haobo Li, Jianhong Tu, Xiaodong Deng, Kashun Shum, Dayiheng Liu, Huamin Qu
Bangla-English tutoring requires more than producing a correct translation: learners also need explanations of grammar differences, awareness of their likely errors, and targeted practice. We present...
arXiv:2601.02933v4 Announce Type: replace
Abstract: Human evaluation is the gold standard for multilingual NLP, but is often skipped in practice and substituted with automatic metrics because it is n...
By Vil\'em Zouhar, Tom Kocmi
arXiv:2604. 07341v2 Announce Type: replace-cross Abstract: Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair, owing to the complex engineering effort required to adapt new PL pairs.
By Ali Reza Ibrahimzada, Brandon Paulsen, Daniel Kroening, Reyhaneh Jabbarvand
The paper introduces CourseChat, an on‑premises, multi‑course retrieval‑augmented generation (RAG) tutor designed for undergraduate business education. It runs behind a campus web gateway, with each of six courses identified by a unique course reference number (CRN) sharing dual AI hosts that provide a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. The authors evaluate different model sizes, noting that larger models failed speed requirements while a 12B and 7B model met the classroom speed gate; a mixture‑of‑experts variant improved some corrections but introduced new errors, leading them to retain an 8B model for production pending further improvement. The study highlights that model selection, evidence sourcing, serving compatibility, and product design must be considered together, though it does not demonstrate learning gains and calls for separate evaluation of faculty ratings, peak‑load capacity, and public‑gateway acceptance.
By Sidney Shapiro, Joshua Lindemann