Hugging Face Trending Papers

SocietyBench: Forecasting Counterfactual Social-World Evolution

Read the original on Hugging Face Trending Papers →

Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 22

The Corroboration Illusion: When More News Makes LLM Forecasts Less True

The paper demonstrates that large language models (LLMs) used for forecasting real‑world events can be manipulated by simply publishing new articles, even without direct access to the model or its retriever. By injecting a small number of targeted news pieces into a common crawl corpus, an adversary can flip over half of the forecast probabilities and significantly degrade forecast accuracy. The study also shows that common defense strategies can be cheaply bypassed, highlighting the vulnerability of probabilistic LLM judgments to information‑supply‑chain attacks.

By Yuan Lu, Yukuan Zhang
arXiv Machine Learning
1d ago

Do Your Own Research: Learning to Forecast by Learning to Search

The paper introduces an agentic forecasting environment built on 2,100+ resolved Polymarket questions, where a language model (Qwen3.5-35B-A3B) learns to gather evidence during rollout via web search, page reading, and financial time series, all filtered to avoid post‑cutoff leaks. Training with single‑epoch GRPO and a Brier‑score reward improves calibration by 30‑40% and reduces search attempts, while the trained policy outperforms four frontier models in evidence‑based forecasting, achieving lower soft‑Brier scores at roughly 5% of the inference cost. The authors release the environment, dataset, and per‑rollout records as a reusable harness for temporal forecasting agents.

By Yusuf Afifi, Artur Kiulian, Anton Polishko, Mykola Khandoga, Hamudi Naanaa, Alina Krasnobrizha