arXiv Machine Learning

GenCAR: Generative Counterfactual Alignment with Risk-Controlled Selection for Out-of-Distribution Recommendation

GenCAR introduces a method for out‑of‑distribution recommendation that balances utility and risk by controlling the proxy‑label false discovery rate (FDR). It frames the problem as an α‑Valid Counterfactual Recommendation (α‑VCR) task, coupling counterfactual supervision with calibrated set selection using conformal p‑values and Benjamini–Hochberg filtering. The approach theoretically bounds counterfactual approximation error and guarantees finite‑sample, distribution‑free FDR control under various dependence assumptions, and empirical results show improved OOD candidate recovery across benchmarks.

arXiv AI
4d ago

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

The paper introduces OTROPE, a likelihood‑free method for off‑policy evaluation of large language models (LLMs) that uses optimal transport to align labeled samples from a behavior model with unlabeled samples from a target model in a semantic space. OTROPE corrects human‑labeled residuals with proxy predictors, achieving a doubly robust evaluation without requiring behavior‑policy modeling or density‑ratio estimation. The authors provide theoretical guarantees for consistency and convergence, and demonstrate through synthetic and real LLM tasks that OTROPE outperforms existing baselines and can elevate weaker evaluators to match or exceed stronger ones.

By Liner Xiang, Wenbo Zhang, Hengrui Cai
arXiv AI
Jun 6

Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

arXiv:2605. 11632v2 Announce Type: replace-cross Abstract: Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), offering a causally grounded approach to unraveling black-box LLM behavior.

By Yilong Wang, Qianli Wang, Bohao Chu, Yihong Liu, Jing Yang, Simon Ostermann
arXiv AI
Sep 17

A Zeroth-Order Paradigm for LLM Preference Alignment

The paper introduces Comparison-based Preference Optimization (ComPO), a zeroth-order method that aligns large language models with human preferences using comparison oracles instead of direct differentiable loss optimization. It provides theoretical convergence guarantees for both offline and online variants under smoothness, gradient sparsity, and oracle compatibility assumptions, and establishes performance bounds under local coverage and in-distribution reward accuracy. Experiments on several LLMs (Mistral, Llama, Gemma-2, Qwen3, Gemma-3) show that ComPO outperforms existing direct alignment methods, achieving higher length-controlled win rates and diagnostics that suggest mitigation of likelihood displacement.

By Peter Chen, Xi Chen, Wotao Yin, Tianyi Lin
arXiv Machine Learning
Sep 7

Consensus Group Relative Policy Optimization for Text Generation

Consensus Group Relative Policy Optimization (C‑GRPO) is a new training method that distills Minimum Bayes Risk (MBR) decoding into a group‑relative objective, enabling text generation models to approximate MBR performance without the costly inference‑time sampling and scoring. C‑GRPO only needs a utility function and policy samples, avoiding the need for gold references or curated preference data. Experiments on WMT 2024 machine translation and XSum summarization show that C‑GRPO matches MBR decoding quality while reducing inference overhead and outperforming other reference‑free baselines.

By Yuki Ichihara, Yuu Jinnai, Kaito Ariu, Eiji Uchibe