arXiv AI By Babak Abbaschian

Cross-Disciplinary Taxonomy and Modeling of Misunderstanding Generation, Amplification, and Detection, from Pragmatics to AI Agents

Read the original on arXiv AI →

arXiv:2608. 13604v1 Announce Type: new Abstract: Detection of misunderstanding is an urgent problem to solve because communication has moved away from real-time, in-person interaction and is increasingly handled by AI-mediated channels.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Feedback That Backfires: Why Small Language Model Agents Repeat the Call They Just Watched Fail

The study investigates why small language model agents tend to repeat a tool call that just failed. By recording the failed call and its error message in the transcript, the authors measure a negative corrective gain—agents are more likely to repeat the failed action, with a drop of about 1.03 nats per token. The problem is traced to the harness design rather than the model’s understanding of errors, and the authors show that replacing the verbatim call with a runtime-generated description of the failure can reduce this backfiring effect by 76%.

By Esmail Gumaan
arXiv AI
2d ago

RealCompanion: Benchmarking Human Understanding from Reasoning over Longitudinal Real-World Conversations

RealCompanion is a benchmark that evaluates an AI companion’s ability to understand a human over long, real-world conversations. It consists of ten real relationships with 27,218 messages spanning up to 120 days, along with derived files such as a profile, persona, chat ground truth, and a question set that cites the relevant messages. The study finds that past context is rarely needed, memory detection is challenging, and agent systems vary widely in cost while achieving similar persona reconstruction.

By Arman Behnam, Sunglyoung Kim, Liangwei Yang
arXiv AI
Aug 24

Recognizing Artificial Minds: A Philosophical Defense of AI Cognition

The paper defends the 'Whole Hog Thesis', arguing that sophisticated large language models such as ChatGPT are full linguistic and cognitive agents, possessing understanding, beliefs, desires, knowledge, and intentions. It rejects low‑level computational starting points and instead builds its case from high‑level behavioral observations, using Holistic Network Assumptions to link actions to mental states. The authors systematically rebut common objections—such as hallucinations and planning errors—by showing these resemble human fallibility and by challenging the necessity of traditional conditions like embodiment or semantic grounding.

By Herman Cappelen, Josh Dever