arXiv AI By Molood Arman

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Coupling Diagnostic for Machine Collectives

Read the original on arXiv AI →

arXiv:2608. 03722v1 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

AgentDiff: Meaning-Bearing Rewrites Trigger Deeper Divergence than Presentation Changes in LLM Agents

The paper introduces AgentDiff, a metric that quantifies how much LLM agents’ answers differ when inputs are altered by meaning‑bearing rewrites (paraphrases, synonym substitutions) versus presentation changes (reordering, formatting, distractors). Across 68 model–benchmark–scaffold combinations involving ten LLMs and over 1,500 questions, meaning‑bearing rewrites consistently produce a roughly 20‑percentage‑point higher inconsistency rate than presentation changes, a gap that persists across severity proxies and remains significant even outside the Qwen family. Trace analysis reveals that meaning‑bearing rewrites preserve the first action but reduce thought similarity from the second step onward, extending the divergence cascade—a phenomenon termed “stealth divergence.”

By Liyun Zhang, Jiayi Guo
arXiv AI
Sep 3

Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence

The paper introduces the concept of an epistemic Sybil problem in multi‑agent AI systems, where multiple agents may produce seemingly independent reports that actually stem from the same underlying evidence. It formalizes this issue using information‑theoretic measures and demonstrates through large‑scale experiments that naive aggregation of replicated reports can severely degrade inference accuracy unless the system accounts for shared evidence ancestry and correlated extraction errors. The study shows that aggregators that track evidential dependence rather than merely report multiplicity or similarity achieve better calibration and inference performance.

By Marc Bara