arXiv Computation and Language By Sahil Mishra, Niranjan Rajeev, Tanmoy Chakraborty

Tree-of-Concerns: Hierarchical Multi-Agent Debate for Unstated-Limitation Extraction in Scientific Critique

Read the original on arXiv Computation and Language →

Tree-of-Concerns is a multi‑agent framework that uses specialized skeptic personas to conduct parallel debate trees, each focusing on a specific category of potential limitations in scientific papers. The system employs structured, evidence‑grounded argumentation and a panel review mechanism to correct drift and miscalibration, ultimately extracting unstated limitations. Experiments on the ToC‑Bench benchmark show that the approach improves precision by 79% and coverage by 11% over leading baselines, providing reviewers with specific, evidence‑based concerns for systematic evaluation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Aug 20

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

The paper investigates how large language models can extract contextualized data from scientific literature. It presents four workflows: expert‑written prompts, self‑generated prompts, autonomous literature discovery, and dataset creation from guidelines. While models perform well with prompts, they struggle with context, hallucinate references, and still need human oversight for final validation.

By Valentin Romanov, Monique Bax, Steven Niederer
Hugging Face Trending Papers
Aug 19

Self-prompting and cross-model consensus enable reproducible data extraction from scientific literature with large language models

The paper evaluates browser-based large language models (LLMs) for extracting detailed, contextualized data from scientific papers. It presents four workflows: (1) expert-curated prompts yield good extraction but struggle with nuance; (2) LLMs can generate effective prompts from simple instructions; (3) autonomous literature discovery is challenging, with missing or hallucinated references; (4) LLMs can build new datasets from guidelines that align closely with human experts, yet still need human oversight. The study outlines a practical, auditable workflow where experts set standards, models cross-check extractions, and researchers resolve disputes, enabling scalable scientific data curation.