This paper introduces Sentence Splitter, a self-supervised framework built upon a T5-based encoder--decoder architecture for uncovering the latent factual structure of natural language sentences. The proposed method identifies the semantic boundary between a descriptive prefix (head) and its factual completion (tail) by formulating sentence splitting as a discrete segmentation problem, where a sentence of length $N$ admits $N$ possible split points but only one recovers the intended head--tail structure.
arXiv:2608.30627v1 Announce Type: new
Abstract: As language-model compute continues to scale, high-quality training data is becoming an increasingly important bottleneck. Conventional next-token pred...
By Haoran Que, Jiajun Shi, Ting Huang, Renming Pang, Jiaheng Liu, Ge Zhang, Wenhao Huang, Shen Yan, Wei Ye, Shikun Zhang
The paper proposes a self‑supervised framework that trains language models to predict concepts—sets of semantically equivalent tokens—rather than single tokens. This approach improves alignment with human similarity judgments, boosts performance on classification, clustering, and reranking tasks, and yields comparable or stronger downstream reasoning while lowering perplexity on semantically meaningful words and only slightly increasing overall perplexity.
By Christine Zhang, Dan Jurafsky, Chen Shani
arXiv:2609.15338v1 Announce Type: cross
Abstract: Large Language Models (LLMs) primarily perform inference at the token level, resulting in substantial memory overhead and compromised computational e...
By Peipei Li, Dongsen Zhang, Yuchen Liu, Wenjun Xu
arXiv:2607. 19354v1 Announce Type: new Abstract: Spreadsheet applications are used by hundreds of millions worldwide, yet writing formulas remains a significant barrier.
By Cy Xie
arXiv:2608. 19529v1 Announce Type: cross Abstract: Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language.
By Su Yan, Rakesh Iyer