Fine-Tune MMS Adapter Models for low-resource ASR
Related stories
Fine-Tune W2V2-Bert for low-resource ASR with ๐ค Transformers
Fine-Tune Whisper For Multilingual ASR with ๐ค Transformers
Introducing the FFASR Leaderboard: Benchmarking ASR in the Real World
Fine-tuning LLMs to 1.58bit: extreme quantization made easy
Diving into Kronecker Adapters: Component Design Matters
arXiv:2602. 01267v3 Announce Type: replace Abstract: Kronecker adapters have emerged as a promising approach for fine-tuning large-scale models, enabling high-rank updates through tunable component structures.
Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR
arXiv:2606. 24169v1 Announce Type: new Abstract: Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an English-only (EN) encoder.
MURMUR: An Efficient Inference System for Long-Form ASR
arXiv:2606. 01483v1 Announce Type: cross Abstract: Long-form automatic speech recognition (ASR) requires both high accuracy and low latency, but existing systems force a trade-off between the two.
(LoRA) Fine-Tuning FLUX.1-dev on Consumer Hardware
Operator Fusion for LLM Inference on the Tensix Architecture
arXiv:2606. 09879v1 Announce Type: new Abstract: This study addresses on-device inference bottlenecks of Transformer models on Tenstorrent's Tensix architecture and proposes an operator fusion strategy that enhances data locality.
An 84-Format Numeric Catalog with Bit-Exact Conformance Vectors: A Vendor-Neutral Reference for FP8, BF16, MXFP4, and Microscaling Formats
arXiv:2606. 09686v1 Announce Type: cross Abstract: Numeric format proliferation in machine learning hardware -- FP8 (E4M3 and E5M2), BF16, MXFP4, microscaling block formats, and dozens of research variants -- has outpaced the availability of vendor-neutral, bit-exact reference material.
Speech LLMs in Low-Resource Scenarios: Data Volume Requirements and the Impact of Pretraining on High-Resource Languages
arXiv:2508. 05149v2 Announce Type: replace-cross Abstract: Large language models (LLMs) have demonstrated potential in handling spoken inputs for high-resource languages, reaching state-of-the-art performance in various tasks.