A benchmark score means nothing without knowing what a trivial method achieves and what the best possible method could achieve. We construct both bounds for a task with a rare kind of ground truth: predicting which sentences a crowd of readers -- highlighting for their own purposes, unpaid, uninstructed, and blind to each other -- marked in 120 web documents.
arXiv:2608. 15980v1 Announce Type: cross Abstract: Preference benchmarks are built by hiring annotators, and the identity of those annotators is treated as an implementation detail.
By Anik Jha
Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion.
arXiv:2608. 09093v1 Announce Type: cross Abstract: How a document's arrangement is written down, its notation, is a training variable that no dataset card records.
By E. M. Freeburg
arXiv:2608. 03722v2 Announce Type: replace Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise.
By Molood Arman
arXiv:2608. 03722v1 Announce Type: new Abstract: Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise.
By Molood Arman