arXiv AI By Arunkumar V, Manoranjan Gandhudi, Gangadharan G. R., Arun Prakash, S. Senthilkumar

MA-SBI: Misspecification-Aware Simulation-Based Inference via Side-Channel Guidance

Read the original on arXiv AI →

arXiv:2606. 16923v1 Announce Type: new Abstract: Simulation-based inference (SBI) of latent parameters is often hindered by simulator misspecification, the mismatch between simulated and real-world observations caused by inherent modeling simplifications.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
4d ago

OTROPE: Optimal Transport-based Robust Off-policy Evaluation for Large Language Models

The paper introduces OTROPE, a likelihood‑free method for off‑policy evaluation of large language models (LLMs) that uses optimal transport to align labeled samples from a behavior model with unlabeled samples from a target model in a semantic space. OTROPE corrects human‑labeled residuals with proxy predictors, achieving a doubly robust evaluation without requiring behavior‑policy modeling or density‑ratio estimation. The authors provide theoretical guarantees for consistency and convergence, and demonstrate through synthetic and real LLM tasks that OTROPE outperforms existing baselines and can elevate weaker evaluators to match or exceed stronger ones.

By Liner Xiang, Wenbo Zhang, Hengrui Cai
arXiv AI
2d ago

Nous: Learning and Certifying Memory Decisions Before Source Calibration

The paper introduces Nous, a framework that separates learning, calibration, and revision certification for belief‑based agent memory. It shows that learning and certifying useful memory decisions can require quadratically fewer records than source calibration, and provides finite‑sample certificates for policy improvement without needing to recover source reliability. Experiments on MiniGrid environments demonstrate that the new certificate reliably accepts improvements over incumbents, outperforming earlier methods.

By Pranav Singh