arXiv:2605. 04135v2 Announce Type: replace-cross Abstract: Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do.
By David Gringras, Misha Salahshoor
arXiv:2607. 29539v1 Announce Type: cross Abstract: Standard AI-text detection benchmarks compare human-written text against text generated directly by large language models (LLMs).
By Gaetano Perrone, Simon Pietro Romano
arXiv:2606. 28353v1 Announce Type: cross Abstract: Linking FDA-approved medical devices to their underlying United States Patent and Trademark Office (USPTO) patents enables critical applications such as recall root-cause analysis, M&A-driven IP discovery, and technology trajectory mapping.
By Yang Qingqing, Liu Haijiang, Li Moyan
arXiv:2607. 18360v1 Announce Type: cross Abstract: Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set.
By Patrik Reizinger, Wieland Brendel
arXiv:2606. 10794v3 Announce Type: replace Abstract: Existing black-box LLM provenance methods achieve comparability by querying every candidate model with the same diagnostic prompts.
By Jiaxu Liu, Sunnan Mu, Dong Huang, Liuyin Wang, Jing Shao, Jie Zhang
arXiv:2608. 14551v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used for title-abstract screening in systematic reviews, but their decisions lack calibrated uncertainty.
By Arya Rahgozar, Pouria Mortezaagha
arXiv:2606. 02959v1 Announce Type: new Abstract: Published evaluations of prompt-injection and jailbreak detectors for Large Language Models often suffer from two systematic weaknesses: per-dataset threshold tuning and undisclosed operating points.
By Ryle Goehausen, Marcus Sousa
QuantumNovelty is an open‑source, skill‑orchestrating language agent that both creates quantum‑computing artifacts—such as papers, ansatz candidates, and patent drafts—and evaluates them through simulated referee and patent‑examiner panels. Its core innovation is an audit‑and‑falsify layer of deterministic gates (Pareto domination, numerical recomputation, Wilson intervals, and cross‑vendor consensus) that restricts claims to those that survive rigorous checks, with every model call logged for transparency. In initial tests on a planted adversarial corpus and a small real‑world deployment, the system successfully flagged all overclaims without false positives and produced panels that were more conservative than typical public acceptance rates.
By Shlomo Kashani
arXiv:2604. 19047v2 Announce Type: replace-cross Abstract: Existing QA benchmarks typically assume distinct documents with minimal overlap, yet real-world retrieval-augmented generation (RAG) systems operate on corpora such as financial reports, legal codes, and patents, where information is highly redundant and documents exhibit strong inter-document similarity.
By Hanjun Cho, Jay-Yoon Lee
arXiv:2605. 10246v2 Announce Type: replace Abstract: AI scientist systems are increasingly deployed for autonomous research, yet their academic integrity has never been systematically evaluated.
By Zonglin Yang, Xingtong Liu, Xinyan Xu
arXiv:2606. 09556v1 Announce Type: new Abstract: AI Scientist agents are often evaluated as if capability were mainly a function of model quality, prompting, or reasoning scaffolds.
By Yinan Wang
arXiv:2607. 20494v1 Announce Type: new Abstract: Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a live Gemini backend behind an active regex filter.
By Alexandre Cristov\~ao Maiorano