Probing Memorization of Tabular In-Context Learning
arXiv:2606. 31208v1 Announce Type: new Abstract: Large tabular models (LTMs), i.
PolicyLong introduces a dynamic on‑policy approach to constructing long‑context data for large language models, addressing the off‑policy gap of previous single‑pass methods. By repeatedly re‑screening data using the model’s current entropy landscape, it creates a self‑curriculum that aligns training distribution with evolving model capabilities. Experiments on RULER, HELMET, and LongBench‑v2 demonstrate consistent performance gains, especially at longer contexts, outperforming EntropyLong and NExtLong.
arXiv:2606. 31208v1 Announce Type: new Abstract: Large tabular models (LTMs), i.
Effective machine learning depends not only on how we model data, but also on what data we choose to collect. While large sequence models have revolutionized data modeling, the problem of automated data selection, or "intrinsic curiosity", remains a significant challenge.
arXiv:2608. 07935v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) adapts a language model by distilling guidance from a frozen teacher on trajectories sampled from the student.
arXiv:2607. 07847v1 Announce Type: new Abstract: As large language models (LLMs) become increasingly capable, the next question is how can we enable models to continually learn?
arXiv:2606. 03841v1 Announce Type: new Abstract: Recent progress in Large Language Model (LLM) agents has enabled promising advances in automated data science.
The paper introduces a holistic framework for large language model self‑evolution that uses learnable information gain to assess the novelty of each training round. Information gain is theoretically linked to the Kullback‑Leibler divergence and entropy change between successive data distributions, and practically estimated by fitting a small language model and scoring new data with negative log‑likelihood. The proposed ATRI method reweights samples within a round and stops training across rounds when information gain is low, and experiments on popular datasets show its effectiveness.
arXiv:2608. 01672v1 Announce Type: cross Abstract: Effective long-context modeling is not merely about retaining more of the past, but about preserving the information that may prove relevant later.
arXiv:2606. 18677v1 Announce Type: cross Abstract: Tabular stream learning requires predictions on sequentially arriving examples under distribution shift.
arXiv:2606. 19476v1 Announce Type: cross Abstract: Effective machine learning depends not only on how we model data, but also on what data we choose to collect.
arXiv:2605. 25582v2 Announce Type: replace Abstract: Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy reuse introduces distribution mismatch that existing trust-region techniques mitigate primarily by enforcing conservative optimization, often leaving rich training signals underutilized.
MInTRL (Minimal Intervention Reinforcement Learning) expands exploration in on-policy reinforcement learning by inserting sparse, local corrections into rollouts via a judge-intervention policy. These interventions replace erroneous suffixes and immediately return control to the main policy, allowing the agent to explore beyond its natural trajectory while maintaining on-policy data. The method uses a sequence-level advantage-regression objective, avoiding importance sampling, and demonstrates superior performance on math and code benchmarks compared to standard on-policy and off-policy baselines.
arXiv:2610.02140v1 Announce Type: cross Abstract: Introducing new capabilities to frontier models has long been the goal of posttraining, which predominantly employs supervised finetuning (SFT) and r...