arXiv:2605.10073v3 Announce Type: replace
Abstract: Pre-trained language models advance patent classification and retrieval by encoding claims as flat token sequences, but they overlook the dependenc...
By Yongmin Yoo, Qiongkai Xu, Zhangkai Wu, Longbing Cao
arXiv:2609.36550v1 Announce Type: new
Abstract: Retrieval-augmented generation is widely used in professional writing, yet whether retrieval grounds revision or merely injects templates is rarely tes...
By Josepha Michiko Leo, Hyun-seok Min, Yehoon Jang, Irvan Zidny, Jin-Woo Chung, Sungchul Choi
arXiv:2608.21924v1 Announce Type: new
Abstract: Patent litigation imposes substantial costs on firms and distorts R&D incentives, making early risk identification a practically important task. While...
By Takao Arai, Hiroyasu Inoue
DECSELFMASK is a decoder‑only classification method that uses unlabeled clinical text to improve performance. It creates self‑supervised training examples by masking portions of the text identified as relevant through relevance attribution, then trains the model to reconstruct the masked tokens via next‑token prediction. Experiments on 136 tasks from 1.9 M Italian hospital notes show consistent gains across five models, outperforming base models (+9.1 Macro F1), continual pretraining (+6.3), and synthetic label generation (+12.5).
By Pietro Ferrazzi, Matteo Merler, Giovanni Bonetta, Alberto Lavelli, Bernardo Magnini
arXiv:2608.29899v1 Announce Type: cross
Abstract: Dense retrieval over long documents is expensive. Token-level encoders scale quadratically in sequence length, and most long-context embedding models...
By Devrim \c{C}avu\c{s}o\u{g}lu, Emre Akba\c{s}
Sparse Coverage is an unsupervised semantic retrieval framework designed for patent prior‑art search. It maps local span embeddings to a sparse vocabulary of embedding‑space centers chosen via a coverage‑oriented k‑center objective, allowing spans to activate nearby centers and produce sparse representations that work with inverted‑index retrieval. Experiments on CLEF‑IP 2013 demonstrate that Sparse Coverage matches or surpasses dense patent encoders in document‑level recall while remaining competitive at the passage level, making it an effective first‑stage retrieval approach for patent search.
By You Zuo (ALMAnaCH), Kim Gerdes (LISN, Qatent, STL), \'Eric de la Clergerie (ALMAnaCH), Beno\^it Sagot (ALMAnaCH)