arXiv AI

A robust association between LLM use and scientific productivity: Assessing stopping-time selection

arXiv:2607. 28968v1 Announce Type: cross Abstract: Renault, Bergeaud, and Bosquet (hereafter RBB) argue that dating LLM adoption as the first month in which an author's abstract is flagged induces a stopping-time selection that can produce a positive event-study path even when there is no causal effect.

arXiv AI
Sep 24

Beyond the Illusion of Power: Calibrating Quasi-Experiments in Observational IS

Information systems researchers increasingly rely on quasi‑experimental methods such as difference‑in‑differences and instrumental variables to infer causal effects from observational panel data. A large Monte Carlo study of 9,837 parameter settings (≈9.8 million simulated datasets) shows that the gap between planned and achieved power is largely driven by serial correlation, panel attrition, staggered adoption bias, and parallel‑trend pre‑testing—factors that no closed‑form power calculator can fully capture. For IV designs, increasing sample size does not improve power or reduce exclusion bias unless instrument strength is enhanced, underscoring that identification hinges on the instrument rather than on larger N.

By Spandan Ghose Chowdhury
arXiv Machine Learning
Jun 9

Benchmark Datasets for Lead-Lag Forecasting on Social Platforms

arXiv:2511. 03877v2 Announce Type: replace Abstract: Social and collaborative platforms emit multivariate time-series traces in which early interactions -- such as views, likes, or downloads -- are followed, sometimes months or years later, by higher impact like citations, sales, or reviews.

By Kimia Kazemian (Department of Computer Science, Cornell University), Zhenzhen Liu (Department of Computer Science, Cornell University), Yangfanyu Yang (Department of Information Science, Cornell University), Katie Luo (Department of Computer Science, Stanford University), Shuhan Gu (Department of Computer Science, Cornell University), Audrey Du (Department of Computer Science, Cornell University), Xinyu Yang (Department of Information Science, Cornell University), Jack Jansons (Department of Computer Science, Cornell University), Kilian Q. Weinberger (Department of Computer Science, Cornell University), John Thickstun (Department of Computer Science, Cornell University), Yian Yin (Department of Information Science, Cornell University), Sarah Dean (Department of Computer Science, Cornell University)
arXiv Machine Learning
Sep 25

Time-Series Foundation Models That Understand Data Revisions

The paper introduces VINTAGE-TS, a revision‑aware time‑series foundation model that separates observation time from information‑availability time. It predicts both the next period’s first‑published value and the value available after a fixed delay, maintaining a joint distribution to capture their dependence and uncertainty. The authors provide a detailed evaluation protocol, software tools for validity‑interval reconstruction and delayed‑label filtering, and a synthetic demonstration with a 25‑configuration sensitivity suite to illustrate performance variability and the impact of hindsight contamination.

By Taimoor Ahmad