arXiv:2608. 00144v1 Announce Type: new Abstract: Membership inference (MIA) on language models is usually summarised by an aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines separate members from non-members from surface text alone.
By Victor Maricato
arXiv:2608. 00144v2 Announce Type: replace Abstract: Membership inference (MIA) on language models is usually summarised by aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines can separate members from non-members using surface text alone.
By Victor Maricato
arXiv:2607. 24276v1 Announce Type: cross Abstract: Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words.
By Priyansh Srivastava
arXiv:2608. 09093v1 Announce Type: cross Abstract: How a document's arrangement is written down, its notation, is a training variable that no dataset card records.
By E. M. Freeburg
arXiv:2608. 05889v1 Announce Type: cross Abstract: Large language models (LLMs) can leave small stylistic traces in text written with their help.
By Przemys{\l}aw Czuma (Polish Association for Artificial Intelligence in Medicine)
Large language models (LLMs) process text through subword tokenizers rather than directly reading characters or words. Because these tokenizers are trained predominantly on English-centric corpora, they introduce a systematic and often overlooked disadvantage for many non-English languages.
arXiv:2601. 20336v5 Announce Type: replace-cross Abstract: Do the functional narratives in cryptocurrency whitepapers correspond to how their tokens behave in markets?
By Murad Farzulla
arXiv:2608. 06305v1 Announce Type: new Abstract: Retrieval-augmented generation over long documents is dominated by one design: chunk the text, embed the chunks, and surface the top-k nearest neighbours of the query.
By Sagar Tamang, Ayush Vyas, Tabarakul Hazarika
arXiv:2608. 03397v1 Announce Type: new Abstract: MMLongBench-Doc is a long-document QA benchmark of 1,082 questions over 135 PDFs.
By Mingtian Zhang
arXiv:2606. 18192v1 Announce Type: new Abstract: As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs).
By Nick Bettencourt, Xiaowei Ding, Kay Giesecke
arXiv:2608. 10715v1 Announce Type: cross Abstract: Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing.
By Lena Holzwarth, Rita Gonz\'alez-M\'arquez, Dmitry Kobak
arXiv:2607. 00224v1 Announce Type: cross Abstract: Watermarking promises a statistical trace of large language model (LLM) use, but real documents, after editing or paraphrasing, rarely arrive as purely human-written or purely machine-generated.
By Shuwen Chai, Qiaosen Wang