arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.
By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
OraclePhys is a fine‑tuning framework for large language models on structural mechanics, comprising a graded benchmark (OraclePhys‑Bench), a 30K supervision dataset (OraclePhys‑30K), and a controlled training study. The study shows that the form of the label’s answer, rather than its length, determines what the model learns, and that certain training objectives can produce models that match or exceed existing LLMs on spatial structural response tasks. The trained 8B model reaches the data‑precision frontier, outperforming zero‑shot and 32‑shot baselines at a specialist level.
By Mingyu Li, Guorui Song, Jing Lin, Haoqian Wang
arXiv:2608. 05151v1 Announce Type: cross Abstract: Wastewater operators need answers grounded in how their plant's variables interact and how fast effects propagate, not in generic pretraining text, when asking causal questions such as "why is N2O rising?
By Gary Simethy, Daniel Ortiz Arroyo, Petar Durdevic
arXiv:2608. 16394v1 Announce Type: new Abstract: Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards.
By Vahid Zolfaghari, Nenad Petrovic, Andr\'E Schamschurko, Alois Knoll
arXiv:2609.16145v1 Announce Type: new
Abstract: We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? W...
By Gautam Kishore
arXiv:2609.01244v1 Announce Type: new
Abstract: Every supervised fine-tuning run forces the same chain of decisions, such as learning rate, batch size, LoRA or full fine-tuning, how many epochs, whic...
By Charles O'Neill, Mudith Jayasekara, Harry Partridge
The paper evaluates how three large mixture‑of‑experts models (Alibaba, OpenAI, NVIDIA) can be fine‑tuned to reason in a low‑resource language, specifically Greek. Accuracy metrics show little change, but the authors uncover significant qualitative improvements: after supervised fine‑tuning, models reason in Greek on ~98% of items, with better grammaticality and retained general ability. Reinforcement learning with pre‑registered rewards further eliminates reasoning‑channel leaks and format skips, while the Greek‑reasoning habit remains robust to an accuracy‑only gradient.
By Ayoub Kirouane, Christos Petrocheilos
arXiv:2609.31688v2 Announce Type: replace
Abstract: In verifiable domains such as math and coding, finding one correct solution among many attempts can matter more than the pass rate of each attempt....
By Eric Fithian, Kirill Skobelev, X. Y. Han
MemToC is a controlled benchmark that tests how large language models resolve conflicts between their internal memory and tool outputs. It contains 6,504 episodes built from 542 factual questions, each paired with a model‑generated closed‑book answer and a tool return whose correctness is known, creating four distinct source‑correctness scenarios. Across five 7‑9B open‑weight models, tool responses overwhelmingly dominate closed‑book answers, and only a minority of instruction‑tuned models correctly retain a verified answer when the tool is wrong, while most follow a correct tool or repeat a wrong tool.
By Arseniy Varlamov, Rishat Zinnatullin, Elisei Rykov, Alexander Panchenko, Ilseyar Alimova
arXiv:2609.25008v1 Announce Type: new
Abstract: I pretrained a language model end-to-end in Rust - alone, with no team, no PyTorch, and no Python in the training path - for $164 in rented GPU time. I...
By Arif Adito
arXiv:2607. 29431v1 Announce Type: new Abstract: Large language models increasingly generate optimization models from natural language, but existing evaluation often reduces a generated model and its ground truth to a single equivalent/not-equivalent verdict or an execution-success rate--labels that are neither independently checkable nor faithful to the multiple distinct senses in which two formulations can agree.
By Penglin Zhu, Jungang Xu
arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.
By Ayushi Agarwal