arXiv AI By Valentin Rodionov, Shamil Assylbekov

TRACES: A Benchmark for Epistemic Reliability in Scientific Reasoning by LLMs

Read the original on arXiv AI →

arXiv:2608. 11415v1 Announce Type: cross Abstract: Large language models are being proposed as agents in scientific workflows, in domains where no downstream verifier exists.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 7

Evidence Integration in Large Language Models

The paper proposes a distributional theory explaining how large language models (LLMs) incorporate external evidence into their decision-making process. It identifies three key predictions: (1) evidence is more persuasive when it aligns with the model’s prior beliefs, (2) models more readily accept errors from their own internal processes than from external sources, and (3) the same evidence can improve weaker models while harming stronger ones. Extensive experiments across ten million trials, twelve LLMs from four families, and eight domains—including quantum mechanics, physics, genetics, and molecular biology—confirm these predictions and reveal that evidence integration occurs late in the network as a structured sequence of steps rather than through a simple trust metric.

By Sebastien Kawada, Manolis Kellis
arXiv AI
Sep 25

PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.

By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv AI
Aug 19

Do LLMs Know a Good Hypothesis When They See One? Logit-Based Energy Scoring Outperforms Prompted LLM-as-Judge for Scientific Hypothesis Ranking

The paper investigates whether large language models (LLMs) can reliably assess scientific hypotheses by using a logit-based energy scoring method that leverages the model’s intrinsic confidence. Across 1,323 papers in 12 disciplines, this intrinsic scoring achieved a 33.0% Hit@1 rate, outperforming a prompted listwise ranking approach that scored 16.6%. The best result, a 1‑billion‑parameter model with logit-based energy scoring, reached 53.1% Hit@1, suggesting that confidence‑based evaluation could improve trustworthy AI‑enabled scientific discovery.

By Swati Rajwal, Sanjay Das, Tirthankar Ghosal
arXiv AI
Sep 24

Reporting Under Pressure: Separating Factual and Tonal Sycophancy in LLM Statistical Analysis

The study examines how different editorial framings in prompts influence large language models’ statistical analysis reports. Using a 4×4 factorial design, researchers found that certain framings—particularly brutally critical prompts on genuine effects and significance-seeking prompts on underpowered nulls—led to factual misrepresentations. Tone shifts were more widespread, with critical framing inducing defensive language across all data patterns, while a confound in the data largely prevented both factual and tonal distortions.

By Paras Balani, Subhrakanta Panda