← Back to all news
Hugging Face Blog November 9, 2020

Leveraging Pre-trained Language Model Checkpoints for Encoder-Decoder Models

Read the original on Hugging Face Blog →

The Flow has not summarised this story yet — read it at Hugging Face Blog.

  • llms

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
4d ago

CForce: Boosting Parallel Decoding for dLLMs via Consistency Forcing

arXiv:2608. 13925v1 Announce Type: cross Abstract: Diffusion large language models (dLLMs) accelerate language generation by predicting multiple masks in a single forward pass.

By Yuji Ren, Chenkai Xu, Zhuocheng Gong, Jianguo Li, Zhijie Deng
llmsdiffusionefficiencysafety
More like this →
arXiv AI
Aug 11

One Adapter Pair per Model: A Universal Activation Interface for Language Models

arXiv:2608. 09521v1 Announce Type: new Abstract: Activation-based tools are usually tied to one model's native hidden space, requiring probes, sparse autoencoders, and natural-language interpreters to be rebuilt or rediscovered for each new language model.

By Su-Hyeon Kim, Jiwan Mun, Yo-Sub Han
llms
More like this →
arXiv Machine Learning
Jul 13

Complexity-Guided Component-wise Initialization for Language Model Pretraining

arXiv:2607. 09204v1 Announce Type: cross Abstract: Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization.

By Konstantin Garbers, Nicholas Oh
llmsnlp
More like this →
Hugging Face Trending Papers
Jun 24

Does Translation-Enhanced Speech Encoder Pre-training Affect Speech LLMs?

Connecting a pre-trained speech encoder to a Large Language Model (LLM) is the standard architecture for building Speech LLMs. However, a structural misalignment exists between the encoder and the LLM.

llmsmultimodal
More like this →
arXiv AI
Jun 3

Plan, Verify and Fill: A Structured Parallel Decoding Approach for Diffusion Language Models

arXiv:2601. 12247v3 Announce Type: replace-cross Abstract: Diffusion Language Models (DLMs) present a promising non-sequential paradigm for text generation, distinct from standard autoregressive (AR) approaches.

By Miao Li, Hanyang Jiang, Sikai Cheng, Hengyu Fu, Yuhang Cai, Baihe Huang, Tinghan Ye, Xuanzhou Chen, Pascal Van Hentenryck
llmsdiffusionbenchmarks
More like this →
Hugging Face Trending Papers
Jun 17

Dango: A Strictly L1-Only Large Language Model for Studying Second Language Acquisition

We introduce Dango, a 1. 8B-parameter large language model designed for controlled studies of L1-to-L2 (Japanese-to-English) transfer in second language acquisition (SLA).

llmsfine-tuning
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.0.0 · bb4ee0e