Train and Fine-Tune Sentence Transformers Models
Related stories
Training and Finetuning Sparse Embedding Models with Sentence Transformers
Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers
Train 400x faster Static Embedding Models with Sentence Transformers
Training and Finetuning Reranker Models with Sentence Transformers
Introduction to Transformers: an NLP Perspective
arXiv:2311. 17633v2 Announce Type: replace-cross Abstract: Transformers have dominated empirical machine learning models of natural language processing.
Training and Finetuning Multimodal Embedding & Reranker Models with Sentence Transformers
Diverging Transformer Predictions for Human Sentence Processing: A Comprehensive Analysis of Agreement Attraction Effects
The study evaluates eleven autoregressive transformer models on English agreement attraction scenarios using a surprisal-based approach. Results show that while transformers match human reading times for prepositional phrase configurations, they perform poorly on object‑extracted relative clauses, with predictions diverging across models and failing to capture human interference patterns. The authors argue that current transformers cannot adequately model human morphosyntactic processing and call for more rigorous, comprehensive testing to avoid misleading conclusions from limited syntactic setups.
Your Transformer Can Hold Two Thoughts at Once: Evidence of Linear Superposition in LLMs
The paper demonstrates that Large Language Models, despite their non‑linear components, exhibit a fundamental linearity property: when inputs from two distinct text streams are linearly combined, the model outputs a superposition of the individual next‑token distributions. This "Superposition Linearity Hypothesis" appears to be an intrinsic feature of the Transformer architecture, tends to weaken during pretraining, but can be largely restored with lightweight fine‑tuning. The authors also present a guided decoding method that separates the superposed outputs, allowing two coherent continuations to be generated from a single forward pass.
Train a Sentence Embedding Model with 1B Training Pairs
Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers
TimpaTeks: Automatic In-place Text Sequence Modification via Diffusion Language Model Steering
arXiv:2606. 08408v1 Announce Type: cross Abstract: We extend activation steering to diffusion language models (DLMs) and study a novel problem that arose due to the inference mechanism of DLMs: Modifying a text in-place to manifest a different concept.