arXiv:2609.13168v1 Announce Type: cross
Abstract: Speech AI, any AI system that recognizes, transforms, or generates speech, is built and evaluated across two communities with only a small overlap: t...
By Maria Teleki, Kimi Wenzel, Anna Seo Gyeong Choi, Tobias Weinberg, Shree Harsha Bokkahalli Satish, Stephanny Sanchez, Belu Ticona, Ariadna Sanchez, Yash Sonkar, Aarti Mathur, Christoph Minixhofer, Abraham Glasser, Raja Kushalnagar, James Caverlee, Minha Lee, Shaomei Wu, Alyssa Hillary Zisk, \'Eva Sz\'ekely, Dylan Gaines, Angelika Seeschaaf Veres, Seray Ibrahim, Nicholas Cummins, Allison Koenecke
arXiv:2608. 13577v1 Announce Type: new Abstract: This position paper argues that the dominant paradigm of AI evaluation (which focuses on superhuman autonomous performance and so implicitly targets the goal of replacing humans) is guiding AI development in the wrong direction.
By Jan Kulveit, Gavin Leech, Tom\'a\v{s} Gaven\v{c}iak, Raymond Douglas
Discover how Intercom built a scalable AI platform with 3 key lessons—from evaluations to architecture—to lead the future of customer support.
arXiv:2502. 20502v2 Announce Type: replace Abstract: Recent advances in Artificial Intelligence (AI) have yielded powerful computational models that, by learning from vast amounts of human-generated data, are increasingly posited as approximate models of human cognition.
By Lance Ying, Katherine M. Collins, Lionel Wong, Ilia Sucholutsky, Ryan Liu, Adrian Weller, Tianmin Shu, Thomas L. Griffiths, Joshua B. Tenenbaum
arXiv:2608. 07473v1 Announce Type: new Abstract: Evaluative AI (EAI) has been recently proposed as a way to support human decision-making, not by producing a single recommendation, but by presenting competing hypotheses together with evidence for and against each.
By Xiang Yin, Tim Miller, Nico Potyka, Antonio Rago, Francesca Toni
arXiv:2609.14236v1 Announce Type: cross
Abstract: With the rapid proliferation of large language model (LLM)-based systems, AI companions have emerged as conversational agents designed to cultivate e...
By Soobin Cho, Deveshi Modi, Divya Mavinkurve, Jieqiong Ding, Mark Zachry