arXiv Computation and Language By Marco Simoni, Aleksandar Fontana, Giulio Rossolini, Andrea Saracino

DANTINOX: A Unified Framework for Multi-Paradigm Language Modeling

Read the original on arXiv Computation and Language →

The paper introduces DantinoX, an open‑source JAX/Flax library that unifies autoregressive decoding, discrete masked diffusion, and continuous flow‑matching language modeling under a single modular Transformer backbone. By keeping the backbone, tokenizer, initialization, and training infrastructure consistent, users can switch between generation paradigms, attention mechanisms, or hardware topologies with only a configuration change. This design enables controlled cross‑paradigm comparisons within one API for training, streaming inference, and benchmarking.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 1

Simulstream: Open-Source Toolkit for Evaluation and Demonstration of Streaming Speech-to-Text Translation Systems

Simulstream is an open‑source toolkit designed to evaluate and demonstrate streaming speech‑to‑text translation systems. It supports both incremental and re‑translation decoding on long‑form speech, offers fine‑grained logging for quality and latency metrics, and includes an interactive web interface for real‑time visualization and comparison. The toolkit addresses the fragmented evaluation landscape by providing a unified framework that accommodates different decoding strategies and input formats.

By Marco Gaido, Sara Papi, Mauro Cettolo, Matteo Negri, Luisa Bentivogli
arXiv Computation and Language
Aug 31

CoFrGeNet: Continued Fraction Architectures for Language Generation

CoFrGeNet introduces Continued Fraction Generative Networks, a new function class that replaces Multi-head Attention and Feed-Forward Networks in Transformer blocks with fewer parameters. The architecture includes custom gradient formulations for efficient optimization and can be plugged into existing Transformer workflows with minimal changes. Experiments on GPT2‑xl and Llama3 show competitive or superior performance on downstream tasks while using 1/2 to 2/3 of the original parameters and shorter pre‑training time.

By Amit Dhurandhar, Vijil Chenthamarakshan, Dennis Wei, Tejaswini Pedapati, Karthikeyan Natesan Ramamurthy, Rahul Nair