arXiv Computation and Language

Benchmarking Patent Drafting from Inventor-Style Disclosures

Hugging Face Trending Papers
Jul 11

When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation

Patent claim drafting is a challenging legal drafting task that requires technical expertise, precise linguistic control, strict adherence to formal conventions, and the preservation of complex logical relationships among claim elements. While Chain-of-Thought (CoT) prompting has been widely used to improve the reasoning capabilities of large language models (LLMs), recent evidence suggests that its benefits may be limited, or even negative, in highly structured or pattern-sensitive tasks.

arXiv AI
Aug 19

QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents

QuantumNovelty is an open‑source, skill‑orchestrating language agent that both creates quantum‑computing artifacts—such as papers, ansatz candidates, and patent drafts—and evaluates them through simulated referee and patent‑examiner panels. Its core innovation is an audit‑and‑falsify layer of deterministic gates (Pareto domination, numerical recomputation, Wilson intervals, and cross‑vendor consensus) that restricts claims to those that survive rigorous checks, with every model call logged for transparency. In initial tests on a planted adversarial corpus and a small real‑world deployment, the system successfully flagged all overclaims without false positives and produced panels that were more conservative than typical public acceptance rates.

By Shlomo Kashani
arXiv AI
Aug 20

Redakto - The Incognito Tab for LLMs

Redakto is a new tool designed to anonymize text before it is processed by large language models (LLMs). It offers state‑of‑the‑art redaction of personally identifiable information (PII) and pseudonymization, accessible via a web interface, REST APIs, and model context protocol hooks. The authors evaluate its performance on legal and medical datasets, showing that anonymized texts retain utility comparable to the originals, enabling LLM tasks without significant loss of effectiveness.

By Saurav Kumar Saha, Tom R\"ohr, Felix Bie{\ss}mann
arXiv AI
Aug 19

SGHA: Evidence-Grounded Research Problem Discovery with Local Language Models

The paper introduces SGHA, a fully automated system that discovers research problems by structuring scientific literature into evidence-linked objects and a typed evidence graph. SGHA operates entirely on a local 9B open‑weight language model, avoiding proprietary frontier‑model APIs, and outputs traceable research‑problem families with assumptions, objectives, success criteria, and ambiguities. Comparative experiments in five machine‑learning domains show that SGHA’s corpus‑first, evidence‑constrained approach yields inspectable research‑problem formulation without relying on external models.

By Sarvesh Gharat, Junpei Komiyama
arXiv Computation and Language
2d ago

Expectations and Practices around AI Disclosure in CS Research

The paper examines AI disclosure policies in top computer science venues, finding them to be highly under‑specified. A survey of 109 researchers shows that disclosures are deemed most necessary for research design tasks and when human involvement is low, and it compiles researchers’ expectations for disclosure content. Analysis of 13,867 disclosure statements from EMNLP 2025 and ICLR 2026 reveals a significant mismatch between these expectations and actual practice, such as frequent disclosure of writing assistance despite it being considered less necessary.

By Arati Mohapatra, Danish Pruthi
arXiv AI
Jun 11

Can AI Agents Synthesize Scientific Conclusions?

arXiv:2606. 11337v1 Announce Type: new Abstract: Scientific AI agents increasingly retrieve evidence, reason across sources, and synthesize conclusions used in consequential decisions.

By Hayoung Jung, Pedro Viana Diniz, Jos\'e Reinaldo Corr\^ea Roveda, Abner Fernandes da Silva, Haeun Jung, Enoch Tsai, Aleksandra Korolova, Manoel Horta Ribeiro