arXiv AI By Yida Lin

PACT: Pairwise-Anchored Calibrated Tuning for Single-Token Typed Decisions

Read the original on arXiv AI →

The paper introduces PACT, a tuning method for single-token typed-decision models that leverages contrastive pair data to add four training terms—difference-in-differences margin, permutation-consistency, evidence-necessity, and ordinal transport cost—without requiring new annotations. PACT achieves comparable accuracy to existing recipes while reducing position bias and ordinal error, and it improves robustness and stability across seeds. The authors provide code, data splits, and trained adapters for reproducibility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Can a Cacheable Decision Model Follow Rules?

The paper evaluates Certo, a small non‑generative decision model that scores candidate actions based on text. It compares a joint scorer that processes state, rules, and candidates together with a cacheable encoder that pre‑encodes candidates to reduce cost. Experiments show the cacheable approach loses rule sensitivity, while targeted counterfactual supervision can recover performance on synthetic tasks; however, on real rules the joint scorer still outperforms the cacheable version, and cross‑domain mixtures do not improve accuracy.

By Dushyant Rajput (AltSlate Labs LLP), Nirdesh Chauhan (AltSlate Labs LLP), Siddharth Kosaraju (AltSlate Labs LLP)
arXiv Machine Learning
Aug 13

LODESTAR: Trustworthy Entropy Is Navigated, Not Merely Measured -- Reinforced Polarizer Keeps a Frozen LLM from Being Confidently Misled by the Wrong Evidence

arXiv:2608. 11922v1 Announce Type: cross Abstract: Predictive-distribution entropy makes a strong selection rule in retrieval-augmented question answering: across five QA benchmarks, keeping the candidate answer that a frozen respondent LLM produces with the lowest answer-token entropy lifts mean answer $F_1$ from 0.

By Po-Jen Ko, Che-Cheng Wu, Hung-Chun Hsu, Li-Yang Chang, Chuan-Ju Wang
arXiv Computation and Language
Sep 17

English Word Sense Disambiguation in 2026: When the Labels Become the Bottleneck

The paper reports that in English all‑words word sense disambiguation (WSD), the scarcity of high‑quality labels—not the models—has become the limiting factor. The authors introduce lexEN, a human‑adjudicated correction layer over the Maru2022 ALL_NEW benchmark, and SenseBench, a living leaderboard for LLM WSD evaluation. They show that frontier large language models reach about 95 % accuracy on lexEN‑v1, that relabeling corpora with these models improves downstream systems, and that fine‑grained WordNet senses are often ill‑posed, with coarsening improving both annotator agreement and model performance. "whyItMatters":"The study highlights that improving label quality and managing annotation costs are now the critical challenges for advancing WSD performance, as model accuracy is already near its theoretical ceiling."

By Vassili Philippov, Amro Salman, Dmitrii Andreev, Penny Hands, Emil Kaiumov, Pavel Katunin, Anton Nikolaev