Token Utility Is Selection-Conditioned: Coupled Selection of Prompt Context and Response Supervision for Efficient Instruction Tuning
Read the original on arXiv Machine Learning →The paper introduces BRIDGE, a method for efficient instruction tuning of large language models that jointly selects prompt context and response supervision. BRIDGE uses a shared validation-directed interaction surrogate to evaluate token utility conditioned on the other side’s retained state, and employs budgeted alternating selection and structure-aware projection to produce coherent supervision spans. Experiments across three model families show BRIDGE outperforms independent selection methods in mathematical reasoning, code generation, and instruction following, with its advantage increasing as compression grows.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.