arXiv Computation and Language

From Tool Use to Technological Agency: LoopCAT as a Local-First, Open-Source Tool for Translation Technology Education

arXiv Machine Learning
Aug 6

ReCodeAgent: A Multi-agent Workflow for Language-Agnostic Translation and Validation of Large-Scale Repositories

arXiv:2604. 07341v2 Announce Type: replace-cross Abstract: Most repository-level code translation and validation techniques have been evaluated on a single source-target programming language (PL) pair, owing to the complex engineering effort required to adapt new PL pairs.

By Ali Reza Ibrahimzada, Brandon Paulsen, Daniel Kroening, Reyhaneh Jabbarvand
arXiv AI
2d ago

On-Premises Multi-Course RAG Tutoring for Business Education: Hardware-Software Trade-offs in a Campus AI Tutor

The paper introduces CourseChat, an on‑premises, multi‑course retrieval‑augmented generation (RAG) tutor designed for undergraduate business education. It runs behind a campus web gateway, with each of six courses identified by a unique course reference number (CRN) sharing dual AI hosts that provide a FastAPI service, a local vector database, and a local large language model (LLM) served by Ollama. The authors evaluate different model sizes, noting that larger models failed speed requirements while a 12B and 7B model met the classroom speed gate; a mixture‑of‑experts variant improved some corrections but introduced new errors, leading them to retain an 8B model for production pending further improvement. The study highlights that model selection, evidence sourcing, serving compatibility, and product design must be considered together, though it does not demonstrate learning gains and calls for separate evaluation of faculty ratings, peak‑load capacity, and public‑gateway acceptance.

By Sidney Shapiro, Joshua Lindemann
arXiv Computation and Language
Sep 18

Evaluating Communicative Success in Machine-Translated Conversation

The paper introduces a three‑layer checklist-and-judge framework to evaluate interpreter agents that mediate live conversation across languages. It assesses semantic, pragmatic, and cultural‑social dimensions—naturalness, intent, and social appropriateness—rather than just fidelity, in both single‑turn and multi‑turn settings. Extensive validation shows that conventional MT metrics miss failures in stronger interpreters, and that context, structured instructions, and cultural cues influence communicative success.

By Faiz Ghifari Haznitrama, Alice Oh
arXiv Computation and Language
Sep 3

Expos\'ia: Teaching and Assessment of Academic Writing Skills for Research Project Proposals and Peer Feedback

Exposía is the first public dataset linking academic writing and feedback in higher education, comprising student research project proposals, peer and instructor comments, and free-text reviews collected from a Computer Science course. It includes human assessment scores based on a fine‑grained, pedagogically‑grounded schema for both writing and feedback. The dataset is used to benchmark large language models on automated scoring of proposals and student reviews, revealing that different LLMs excel at each task and that closed‑source models outperform open‑weight ones, while a multi‑aspect prompting strategy proves most effective for classroom deployment.

By Dennis Zyska, Alla Rozovskaya, Ilia Kuznetsov, Iryna Gurevych