arXiv:2607. 18570v1 Announce Type: cross Abstract: Discourse relations provide document structure, critical to language understanding and enabling language model performance and ethicality.
By Abhidip Bhattacharyya, Shira Wein
arXiv:2606. 00230v1 Announce Type: new Abstract: Grokking, the phenomenon in which neural networks generalize long after fitting their training data, has been studied in supervised settings on many epochs.
By Sherin Muckatira, Namrata Shivagunde, Vijeta Deshpande, Anna Rumshisky
arXiv:2602. 09992v2 Announce Type: replace-cross Abstract: Several recent contributions have evaluated the Poverty of the Stimulus Hypothesis (PoSH) using Artificial Neural Networks (ANNs).
By Xiulin Yang, Arianna Bisazza, Nathan Schneider, Ethan Gotlieb Wilcox
arXiv:2608. 15507v1 Announce Type: cross Abstract: A consistent concept of the current time is important for temporal reasoning, yet how language models represent the current time is not well understood.
By Suze van Adrichem, Aditi Bhaskar, Diyi Yang, Christopher Potts, Jing Huang
arXiv:2608. 06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}.
By Haris Riaz, Hyungji Kim, Mihai Surdeanu
arXiv:2608. 19529v1 Announce Type: cross Abstract: Many real-world AI systems represent entities, behaviors, and structured information using discrete machine-native symbols rather than natural language.
By Su Yan, Rakesh Iyer