arXiv:2609.24932v1 Announce Type: new
Abstract: Despite the success of neural models in natural language processing, their black-box nature limits interpretability and conceals the linguistic phenome...
By David Torres-Moreno, Jorge Hermosillo-Valadez, Asela Reig-Alamillo
arXiv:2608.29529v1 Announce Type: cross
Abstract: Semantic alignment between specialized normative texts is challenging when equivalent requirements use different terms, syntax, and levels of abstrac...
By William Schroeder
arXiv:2606. 08497v1 Announce Type: new Abstract: As deep language models (DLMs) are increasingly deployed in high-stakes domains such as healthcare, understanding their decision rationale becomes paramount for ensuring trust, safety, and accountability.
By Minyoung Hwang, Seokhyun Lee, Changhee Lee
arXiv:2608. 06417v1 Announce Type: new Abstract: The proliferation of misinformation online has driven demand for scalable detection systems.
By Pedro Barcelos, Ot\'avio Parraga, Marcelo M. Mussi, Lucas M. Fraga, Lucas S. Kupssinsk\"u, Rodrigo C. Barros
arXiv:2605. 05103v3 Announce Type: replace-cross Abstract: We introduce the \textbf{Concept Field} of a text corpus: a local drift field with pointwise uncertainty, estimated in sentence-embedding space from the deltas between consecutive sentences.
By Nicholas S. Kersting, Vittorio Castelli, Chieh Ting Yeh, Xinzhu Wang, Saad Taame, Khaoula Allak
arXiv:2607.00171v2 Announce Type: replace
Abstract: Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover on...
By Andrianos Michail, Stylianos Psychias, Michelle Wastl, Simon Clematide, Rico Sennrich, Juri Opitz
arXiv:2608.27813v1 Announce Type: new
Abstract: Structural probes were introduced by Hewitt and Manning to reconstruct syntactic trees from a neural language model's latent representations. They are...
By Juan Pablo Vigneaux, Mary Kennedy, Khalil Iskarous, Robert Frank, Matilde Marcolli
arXiv:2606. 29180v1 Announce Type: new Abstract: A Knowledge Graph (KG) represents facts as structured triples and is widely used to organize relational knowledge across diverse domains.
By Seungryeol Baek, Wooseok Sim, Hogun Park
The paper investigates when compressed vector representations can provide exact linear or affine readouts for a finite lexicon’s truth conditions, establishing a necessary and sufficient row‑space condition. It shows that the augmented truth matrix’s rank determines the minimal dimension needed for exact linear (rank r) and affine (rank r − 1) readouts, and that exact readouts preserve Boolean connectives. Experiments on GloVe and word2vec embeddings reveal that while many predicates are linearly separable, none achieves exact affine recovery from pretrained embeddings, yet supervised transductive training can attain exact affine recovery at dimensions meeting the theoretical bound, preserving most of the original variance.
"whyItMatters":"The results provide a precise mathematical criterion for when vector embeddings can faithfully encode logical truth conditions, informing both theoretical understanding and practical training of language models."
By Daniel Quigley
arXiv:2509. 25045v3 Announce Type: replace-cross Abstract: Despite their capabilities, Large Language Models (LLMs) remain opaque with limited understanding of their internal representations.
By Marco Bronzini, Carlo Nicolini, Bruno Lepri, Jacopo Staiano, Andrea Passerini
arXiv:2607. 22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality.
By Yunting Song, Matthew Watson, Peter Grabowski, Jun Qin
The paper investigates cross‑lingual transfer for sequential sentence classification (SSC) in research papers, focusing on 13 non‑English languages. Experiments show that linguistic proximity does not reliably predict transfer success, whereas structural similarity in rhetorical organization—particularly label distribution similarity—correlates positively with performance. The authors introduce three generative‑model methods that exploit structural cues, achieving parity with strong encoder baselines on‑domain and outperforming them when transferring to unseen languages.
By Kazuhiro Yamauchi, Marie Katsurai