arXiv Machine Learning By Oscar Mir\'o L\'opez-Feliu, Daimy van Loo, Xanthos Kekkos, Mikel Blom, Clara Rus

Reproducing FACTER: Fairness via Conformal Thresholding and Prompt Repair

Read the original on arXiv Machine Learning →

arXiv:2606. 28620v1 Announce Type: cross Abstract: Fayyazi et al.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Sep 2

RPCBench: A Benchmark for Proactive Premise Critique in LLM-based Recommendation

RPCBench is a new benchmark designed to evaluate large language models’ ability to critique recommendation requests by detecting, diagnosing, and handling flawed premises. It includes evidence‑grounded test instances across five recommendation domains and ten types of premise failures, and introduces a fine‑grained evaluation framework covering detection, error localization, handling strategy, and evidence faithfulness. Experiments with 11 LLMs reveal that proactive detection is the main bottleneck, with models struggling most on underspecified‑premise errors and showing that optimal critique quality occurs at intermediate reasoning lengths.

By Zhongru Chen, Yuan Wu, Yi Chang
arXiv AI
Sep 2

Retrieval, Scoring, and Decoding Shape Performance and Stability in LLM-based Conversational Recommendation

The study evaluates large language models (LLMs) as rerankers in conversational movie recommendation, comparing proprietary, open-weight, and fine-tuned LLMs against collaborative-filtering and sequential baselines on the ReDial benchmark. Results show that the best proprietary LLM achieves an NDCG@10 of 0.1497 with a shared semantic candidate pool, outperforming non-LLM baselines, while open-weight LLMs do not surpass a tuned shallow autoencoder under the same protocol. The analysis also highlights that reranker performance is highly sensitive to candidate generation, pool size, scoring policy, and decoding temperature, suggesting these factors should be reported as standard evaluation fields.

By Ante Kapetanovic, Tomislav Duricic, Andro Mercep, Emanuel Lacic
arXiv Machine Learning
Jun 30

Diagnosing and Mitigating Retrieval Bottlenecks in LLM-Based Cold-Start Recommendation

arXiv:2606. 29947v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as rerankers in recommender systems, with the expectation that semantic understanding will help in cold-start and long-tail regimes.

By Zhe Dong (University of Maine at Presque Isle), Fang Qin (Stanford University), Manish Shah (Independent Researcher), Yicheng Wang (Independent Researcher)
arXiv AI
3d ago

Reasoning with Evidence, Not Merely Rationales: Verifiable Preference Proofs for LLM-Based Recommendation

The paper introduces PROVE-REC, a two‑pass framework that generates verifiable preference proofs for large language model (LLM) recommendation systems. In Pass A, the model condenses user interaction histories into a compact proof of positive and avoidance claims linked to specific evidence entries. Pass B then uses only this proof and its evidence to predict the next item, ensuring the recommendation follows the reasoning path. Verification steps compare masked evidence and removed claims to confirm grounding and influence, while a ranking‑preservation objective retains useful historical information. Experiments on diverse real‑world datasets show PROVE‑REC outperforms strong baselines by up to 7.45%, producing claims that are both better grounded and more influential to recommendation quality.

By Yu Hou, Nathaniel Kang, Pengkai Wang, Hua Li