arXiv Machine Learning By Can Wu, Xinrui Chen, Ou Wu, Yi Du

Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning

Read the original on arXiv Machine Learning →

The paper introduces BRIDGE, a method for efficient instruction tuning of large language models that jointly selects prompt context and response supervision. BRIDGE uses a shared validation-directed interaction surrogate to evaluate token utility conditioned on the other side’s retained state, and employs budgeted alternating selection and structure-aware projection to produce coherent supervision spans. Experiments across three model families show BRIDGE outperforms independent selection methods in mathematical reasoning, code generation, and instruction following, with its advantage increasing as compression grows.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 2

PARTREP: Learning What to Repeat for Decoder-only LLMs

While decoder-only LLMs excel at a vast array of natural language tasks, it suffers from an asymmetric information flow induced by causal attention: later tokens are richer in contextual grounding than earlier ones. A simple and effective remedy is prompt repetition -- just appending a second copy of prompt before generation can redistribute grounding across positions and improve reasoning performance.