From Articles to Publishers: Aggregating Language Model Predictions for News Source Reliability Inference
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2607. 28282v1 Announce Type: cross Abstract: Evaluating the quality and relevance of textual outputs from Large Language Models (LLMs) remains challenging and resource-intensive.
The paper demonstrates that large language models (LLMs) used for forecasting real‑world events can be manipulated by simply publishing new articles, even without direct access to the model or its retriever. By injecting a small number of targeted news pieces into a common crawl corpus, an adversary can flip over half of the forecast probabilities and significantly degrade forecast accuracy. The study also shows that common defense strategies can be cheaply bypassed, highlighting the vulnerability of probabilistic LLM judgments to information‑supply‑chain attacks.
The paper introduces GPTBIAS, a framework that uses powerful large language models like GPT‑4 to evaluate bias in other LLMs. It employs specially crafted prompts called Bias Attack Instructions to probe for bias and outputs a bias score along with detailed information such as bias types, affected demographics, keywords, reasons, and improvement suggestions. Extensive experiments demonstrate the framework’s effectiveness and usability.
The paper introduces Strategic 16K, a 16,000‑document corpus of diplomatic cables from WikiLeaks’ Public Library of US Diplomacy, designed to eliminate label leakage. It benchmarks six models—both classical machine learning and transformer-based—on this leakage‑controlled dataset, finding BERT and ELECTRA top performers while TF‑IDF with Logistic Regression offers strong accuracy at lower cost. This work provides the first fully reproducible sensitivity‑classification benchmark built under explicit leakage‑control conditions.
The paper presents the Temporal Coherence Score (TCS), a continuous, interpretable metric for detecting temporal inconsistencies in political news. TCS is computed through a four‑stage pipeline that extracts temporal facts, builds a temporal knowledge graph, verifies consistency using internal rules and external references, and aggregates scores with explanations. On a benchmark of 100 political articles with injected errors, TCS achieves 0.909 precision, providing detailed explanations for each flagged inconsistency.
The paper introduces DECO, a diagnostic framework that factorises content into independent moderation criteria, allowing controlled evaluation of large language models (LLMs) at the criterion level. Using pairwise evaluation across four datasets and four LLMs, the authors find that high aggregate benchmark scores can mask significant failures when decisions hinge on specific content aspects required by individual criteria. The study underscores that aggregated labels do not guarantee reliable criterion-conditioned performance, highlighting the need for evaluation methods that explicitly assess this behavior.