arXiv AI

MimeLens: Position-Agnostic Content-Type Detection for Binary Fragments

arXiv:2606. 04171v1 Announce Type: cross Abstract: File-type classification underlies many workflows like malware triage, forensic carving, packet inspection, and storage indexing.

arXiv Computation and Language
Sep 16

Where Post-Training Quantization Breaks Text Embedders: A Measured Map Across Four Embedder Families

arXiv:2609. 16391v1 Announce Type: cross Abstract: Weight-only post-training quantization is the cheapest way to shrink a retrieval embedder, and the received advice for applying it -- protect the embedding table, allocate bits by module sensitivity, prefer a ranking-aware objective over weight reconstruction -- was carried into LLM quantization largely intact.

By Hyojung Han
arXiv Machine Learning
Sep 14

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models

The paper investigates whether small models distilled from larger ones behave similarly when using byte versus token tokenization. It introduces two methods—Marginalize‑It (approximate) and End‑Of‑Token (exact)—to convert token logits to byte logits, and conducts a large‑scale study on decoder‑only dense transformers ranging from 1 billion to 1 trillion bytes of data. Results show that while token‑based models excel early, byte‑based models eventually surpass them with more compute, achieving higher performance ceilings, greater data efficiency, and lower logit storage costs.

By Kalyani Marathe, Artidoro Pagnoni, Tomasz Limisiewicz, Margaret Li, Mike Lewis, Luke Zettlemoyer, Srinivasan Iyer
arXiv Machine Learning
Jul 8

Multi-Channel Spread-Spectrum Code Watermarking

arXiv:2607. 06009v1 Announce Type: cross Abstract: Attributing code to the large language model that produced it is essential for provenance, licensing, and misuse accountability, yet no deployed watermark meets this need.

By Soohyeon Choi, Debin Gao, Yue Duan
arXiv Machine Learning
Jul 7

Walma: Learning to See Memory Corruption in WebAssembly

arXiv:2603. 24167v2 Announce Type: replace-cross Abstract: WebAssembly's (Wasm) monolithic linear memory turns a single memory-corruption bug into a bidirectional threat: a compromised module can attack its embedding host, and a malicious host can tamper with a trusted module's state.

By Oussama Draissi, Mark G\"unzel, Ahmad-Reza Sadeghi, Lucas Davi