arXiv AI By Qiming Bao, Juho Leinonen, Paul Denny, Michael J. Witbrock

RLearner-LLM: Balancing Logical Grounding and Fluency in Large Language Models via Hybrid Direct Preference Optimization

Read the original on arXiv AI →

arXiv:2605. 04539v4 Announce Type: replace-cross Abstract: Direct Preference Optimization (DPO), the efficient alternative to PPO-based RLHF, falls short on knowledge-intensive generation: standard preference signals from human annotators or LLM judges exhibit a systematic verbosity bias that rewards fluency over logical correctness.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 26

NebulaExp-8B: An Empirical Post-Training Pipeline via Full-Scale Ablation Research

arXiv:2606. 26671v1 Announce Type: new Abstract: Post-training alignment determines the reasoning and human preference following capabilities of large language models, yet most existing works withhold detailed data construction, filtering rules and training recipes, which hinders community reproducibility and lightweight model optimization.

By Qiaobo Hao, Yangqian Wu, Shunyi Wang, Zhongjian Zhang, Ziqun Li, Yayin He, Muqing Li, Chen Zhong
arXiv Computation and Language
3d ago

Listening to the Wise Few: Query-Key Alignment Unlocks Latent Correct Answers in Large Language Models

arXiv:2410.02343v2 Announce Type: replace Abstract: Large language models (LLMs) routinely fail to output the correct option in multiple-choice question answering (MCQA) while encoding the answer int...

By Eduard Tulchinskii, Kristian Kuznetsov, Laida Kushnareva, Anastasia Voznyuk, Andrei Andriiainen, Irina Piontkovskaya, Evgeny Burnaev, Serguei Barannikov
arXiv Computation and Language
Sep 16

Style-Debiased DPO: Updating LLM Knowledge with Factuality-Aware Synthetic Preference Data

The paper introduces Style‑Debiased DPO (SD‑DPO), a method that refines large language models’ ability to retrieve stored knowledge by using preference optimization that corrects for style differences while preserving factual accuracy. SD‑DPO evaluates on the EntiGraph storing‑side framework and outperforms baseline CPT on the QuALITY reading‑comprehension benchmark, achieving higher accuracy with far fewer training tokens. In a knowledge‑editing setting (AToKE), SD‑DPO attains an overall accuracy of 0.982, correctly answering queries with either new or old facts based on the requested time period.

By Takayuki Yamamoto, Daisuke Kawahara
arXiv Machine Learning
Sep 11

Domain-Specific Hallucination Detection in Large Language Models

The paper introduces a multi‑signal pipeline for detecting hallucinations in large language models, combining fine‑tuned DeBERTa‑v3 classification, Monte Carlo Dropout uncertainty, and temperature‑scaled calibration. On the HaluEval benchmark it achieves high performance (F1 = 0.915, AUROC = 0.977) across QA, summarization, and dialogue, and shows that 25 % of training data yields 77 % of full‑data performance. The authors also demonstrate that applying Direct Preference Optimization to a Qwen2.5‑0.5B generator cuts hallucination rates from 85.5 % to 37.7 %, and that domain‑specific fine‑tuning (PubMedBERT on SciFact) outperforms general‑domain models for biomedical text.

By Varun Teja Chundru, Debasmita Biswas