arXiv Machine Learning By Jihoon Tack, Philippe Laban, Jennifer Neville

LLMs Get Lost in Evolving User Intent

Read the original on arXiv Machine Learning →

arXiv:2607. 20734v1 Announce Type: new Abstract: As LLMs become more capable, they are increasingly deployed as collaborative agents, taking on user-delegated tasks through iterative interaction.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
6d ago

Mutable Transcripts: Mitigating Context Pollution through Editable Conversation State

The paper introduces mutable transcripts, an interaction paradigm that lets users edit prior turns in a chat, turning the conversation history into an editable state rather than a fixed record. A prototype was built and tested with 17 participants, who preferred mutable transcripts over standard chat for clarity, confidence, and ease of use, and reported less need to restart conversations. Analysis of user study transcripts shows that mutable transcripts can shorten conversations and remove outdated context, suggesting that user-driven revisions improve interaction quality.

By Dan Barry, Andrew Hines
arXiv AI
Aug 20

The Lifecycle of LLM-as-a-Judge for Large-Scale Recommendation Explanations

The paper introduces a lifecycle framework for LLM-as-a-Judge systems used to evaluate recommendation explanations at Netflix. It outlines four phases—Birth, Training, Deployment, and Monitoring—detailing how each stage addresses specific technical and operational challenges. The authors report that after five weeks of A/B testing, judge-aligned explanations increased novel content viewing and successful browse-to-play sessions without quality takedowns.

By Emma Yanyang Kong, JJ Tan, Ishan Gupta, Lars Olds, Claire Campbell, David Fagnan, Veli Balin, Rohan Gosain, Louis Garcia, Minsu Jang