A companion paper showed that a transformer's feed-forward layer can be rebuilt from explicit fuzzy set operations - intersection, set-difference, and a self-forgetting sequence quantifier - so its hidden units read as named logical operators at no cost to language-model quality. That left the other half of the transformer opaque.
arXiv:2607. 04319v1 Announce Type: cross Abstract: A companion paper showed that a transformer's feed-forward layer can be rebuilt from explicit fuzzy set operations - intersection, set-difference, and a self-forgetting sequence quantifier - so its hidden units read as named logical operators at no cost to language-model quality.
By Mark Oskin
arXiv:2610.00526v1 Announce Type: cross
Abstract: In-context learning (ICL) can be amortized into latent objects (task vectors, function vectors, context vectors) that recover few-shot behavior at ze...
By Gunmay Jhingran
Baobab compiles an OWL 2 DL (ΣROIQ) ontology with a finite ABox into a Sentential Decision Diagram (SDD), saturating a propositional core and instantiating remaining DL features over the active domain. The resulting evidence‑conditioned weighted model count trains a perception network to recognize real images under partial ABox supervision, enabling a CNN to recover latent ontology concepts that an independent perception would miss. When supervision allows multiple ontology‑consistent completions, Baobab’s mixture indexed by query justifications represents the calibrated posterior, achieving Bayes‑optimal performance on a real‑image MNIST task where single‑WMC and learned mixtures fail, thereby characterizing and mitigating reasoning shortcuts in a non‑Horn description logic.
By Olga Mashkova, Asaad Mohammedsaleh, Fernando Zhapa-Camacho, Robert Hoehndorf
arXiv:2608. 06111v1 Announce Type: cross Abstract: Positional embeddings (PE) in Transformers encode token distance and order but are largely agnostic to \textit{syntactic structure}.
By Haris Riaz, Hyungji Kim, Mihai Surdeanu
arXiv:2607. 08946v1 Announce Type: new Abstract: A transformer can be built from operators that are legible by construction -- bounded, named units that read as fuzzy set operations rather than dense activations -- but legibility must be pressed for during training, and the pressure has a failure mode.
By Mark Oskin
arXiv:2608.20873v1 Announce Type: new
Abstract: Every way of teaching a deployed language model something new -- full fine-tuning, adapter merging, model editing -- replaces the released checkpoint,...
By Zifeng Liu, Zhiyong Du, Yaxin Lu, Yiming Mao, Zhenhe Wang, Wenqi Shi, Zhengkun Jing
The paper argues that meaning identity—whether two sentences convey the same idea after wording changes—is not encoded in the geometry of independently produced sentence embeddings. Experiments on frozen off‑the‑shelf encoders and language models show that identity can only be reliably computed when both sentences are processed together in a single forward pass, yielding high accuracy (0.90–0.96) on PAWS‑X, whereas independent embeddings or simple fusion methods perform near chance. Even advanced bi‑encoder fine‑tuning improves performance on PAWS but fails to generalize to other similarity tasks, underscoring that identity is a cheap computed operator rather than a property of individual sentence vectors.
By Jiaqi Deng
arXiv:2609.24942v1 Announce Type: new
Abstract: A model generalizes outside its training distribution only when it computes a representation structurally equivalent to the generating mechanism, not a...
By Filipe Marinho Rocha, In\^es Dutra, V\'itor Santos Costa, Lu\'is Paulo Reis
The paper shows that AI‑text detectors, rather than learning a clear AI‑versus‑human boundary, amplify an inherited predictability axis that already exists in language models. This amplification causes detectors to over‑flag fluent, formal human writing while missing high‑temperature AI outputs, and the bias persists across languages, code, and detector architectures. A training‑free operator can relocate the bias but cannot erase it, underscoring that the unfairness is a structural cost of out‑of‑distribution generalization.
By Alexander Smirnov
arXiv:2608.23268v1 Announce Type: new
Abstract: Frontier multimodal large language models (MLLMs) deliver impressive perception yet still falter on scientific and mathematical reasoning. Parameter-le...
By Jieke Wang, Tiancheng Shen, Yibo Yang, Ming-Hsuan Yang
Relational BabyLM is a decoder‑only Transformer that replaces standard self‑attention with a Dual Attention Transformer (DAT) to separate object‑level lexical features from structural/relational information. The model incorporates a Next‑Latent Prediction objective to compress history into a dense belief state and introduces a RoPE‑based symbol‑retrieval mechanism. On the BabyLM 2026 challenge, the best model ranks 6th overall and 3rd on the NLP‑task subset, outperforming GPT‑2 on most benchmarks and achieving the highest EWoK score among strict‑track entries.
By Adrian Brasoveanu, Ece Takmaz, Jakub Dotla\v{c}il