CAVEWOMAN: How Large Language Models Behave Under Linguistic Input and Output Compression
arXiv:2606. 24083v1 Announce Type: cross Abstract: "Talk short.
arXiv:2606. 24083v1 Announce Type: cross Abstract: "Talk short.
arXiv:2609.16454v1 Announce Type: new Abstract: Recent work by Doshi and Hauser (2024), Bisbee et al. (2024), and Xie et al. (2026) raises concerns that outputs from large language models (LLMs) tend...
arXiv:2606. 07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently.
arXiv:2606. 07559v2 Announce Type: replace-cross Abstract: Fine-tuning a language model often fails silently when its correct completion must outrank a near-synonym competitor.
arXiv:2606. 24998v1 Announce Type: new Abstract: Language models are running out of high-quality training data, and even aggressively deduplicated corpora retain some amount of repetition.
The paper introduces Janus, a method for validating error patterns in language models by comparing error rates across predefined yes/no properties and using shuffled decoy labels to set significance thresholds. Janus requires that a pattern’s error difference surpasses the decoy-derived threshold and is replicated on held‑out data before reporting. Experiments on a controlled code‑finding task confirm several meaningful error patterns, while on MuSiQue and LongBench v2 Janus reports no confirmed patterns for the tested properties, contrasting with standard shuffling tests that sometimes confirm patterns.
arXiv:2607. 18292v2 Announce Type: replace Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2607. 18292v3 Announce Type: replace-cross Abstract: Bigger language models are less reliable.
arXiv:2607. 18292v1 Announce Type: cross Abstract: As language models scale, answers start truer but degrade faster: scaling buys capability but erodes reliability.
arXiv:2506.17871v4 Announce Type: replace-cross Abstract: Despite their impressive capabilities, aligned large language models (LLMs) often generate outputs that lack diversity. What drives this cons...
The study evaluates how extractive prompt compressors affect token costs across ten languages, finding that compressors trained on English data widen the token premium gap for non‑English languages, while a multilingual compressor does not. The gap is tied to the supervision data rather than model architecture, and aggressive compression can reduce non‑English contexts to near‑zero utility. A translate‑then‑compress approach can match or outperform native compression at roughly half the token cost in several languages.
arXiv:2606. 19558v1 Announce Type: new Abstract: Fidelity metrics, such as per-token KL divergence (KLD) against a high-precision reference, are often used in practice as low-cost proxies for benchmark quality.