arXiv AI

Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting Agents

Hugging Face Trending Papers
Jul 11

When Reasoning Hurts Legal Drafting: The Verbalization Bottleneck in Patent Claim Generation

Patent claim drafting is a challenging legal drafting task that requires technical expertise, precise linguistic control, strict adherence to formal conventions, and the preservation of complex logical relationships among claim elements. While Chain-of-Thought (CoT) prompting has been widely used to improve the reasoning capabilities of large language models (LLMs), recent evidence suggests that its benefits may be limited, or even negative, in highly structured or pattern-sensitive tasks.

arXiv Machine Learning
Aug 24

Designing a Robust LLM-Based Evaluation System for Agentic AI in Drug Discovery Through Human Alignment

The paper introduces an LLM-as-a-Judge framework for evaluating the outputs of an agentic drug discovery assistant, ChatInvent, deployed at AstraZeneca. It defines four quality dimensions—Completeness, Relevancy, Structural Clarity, and Scope Adherence—alongside deterministic Tool Call Correctness checks, and validates the judge against five expert annotators. After optimizing the best-performing judge with few-shot demonstrations, alignment with human majority votes improves from 0.80 to 0.86, and the framework reveals that informal question phrasing does not degrade output quality.

By Emma Granqvist, Roc\'io Mercado, Samuel Genheden
arXiv AI
2d ago

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

VeriHarness is a method that enhances verification for large language model agents tackling long‑horizon tasks without needing reference answers at test time. It transforms the base LLM into an agentic verifier by providing a workspace, evidence tools, and reusable verification skills, using disagreement resolution and consensus challenge to evaluate competing claims. Across five benchmarks and two frontier models, VeriHarness outperforms baselines, achieving significant performance gains and demonstrating self‑improvement of verification skills from failure feedback.

By Caiqi Zhang, Rujun Han, Zifeng Wang, Zoey CuiZhu, Nigel Collier, Tomas Pfister, Chen-Yu Lee
Hugging Face Trending Papers
Aug 6

SkillTV-Bench: Benchmarking How Well Judges Perform on Skill-Augmented Agentic Execution

LLM agents increasingly execute long-horizon tasks through tool use and environment interaction, shifting evaluation from final-response scoring to verification of complete executions. For skill-augmented agents, verification additionally requires the procedural knowledge encoded in task-time skills, because this knowledge indicates what evidence to inspect and which failures are task-critical.

arXiv AI
Jun 12

Fantastic Scientific Agents and How to Build Them: AgentBuild for Rietveld Refinement

arXiv:2606. 12834v1 Announce Type: new Abstract: As scientific workflows shift from deterministic executables to LLM-based agents, the development practices on offer, such as fine-tuning, reinforcement learning, and prompt-and-go, bury the scientist's judgment.

By Woong Shin, Craig A. Bridges, Marshall T. McDonnell, Rafael Ferreira da Silva
arXiv AI
Jun 2

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.

By Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma, Nathan Miller
arXiv AI
Aug 19

QuantumNovelty: A Skill-Orchestrating Language Agent for Referee-Style Review and Patentability Screening of Quantum Papers and Patents

QuantumNovelty is an open‑source, skill‑orchestrating language agent that both creates quantum‑computing artifacts—such as papers, ansatz candidates, and patent drafts—and evaluates them through simulated referee and patent‑examiner panels. Its core innovation is an audit‑and‑falsify layer of deterministic gates (Pareto domination, numerical recomputation, Wilson intervals, and cross‑vendor consensus) that restricts claims to those that survive rigorous checks, with every model call logged for transparency. In initial tests on a planted adversarial corpus and a small real‑world deployment, the system successfully flagged all overclaims without false positives and produced panels that were more conservative than typical public acceptance rates.

By Shlomo Kashani