TextQuests: How Good are LLMs at Text-Based Video Games?
Related stories
Introducing Agents.js: Give tools to your LLMs using JavaScript
LLMs help robots understand vague instructions and focus on key details
To help robots do chores in places like homes and factories, a new approach from MIT uses one language model to clarify users’ instructions, then another to ignore irrelevant info.
Game Arena: Strategic LLM Evaluation in Competitive Environments
Game Arena is an open, continuously expanding platform that evaluates large language models through competitive games, allowing head‑to‑head matchups in structured environments. Unlike static benchmarks, it prevents performance saturation by increasing gameplay difficulty as models improve. The report outlines the infrastructure and presents three pilot games—Chess, Poker, and Werewolf—covering perfect information, imperfect information, and multiplayer settings, and details evaluation metrics and competition results.
Judge Arena: Benchmarking LLMs as Evaluators
TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment
LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-wide adoption of LLMs necessitates continuous evaluation to ensure their safety and fairness.
Open-source LLMs as LangChain Agents
Open-Source Text Generation & LLM Ecosystem at Hugging Face
EMemBench: Interactive Benchmarking of Episodic Memory for VLM Agents
arXiv:2601.16690v2 Announce Type: replace Abstract: We introduce EMemBench, a programmatic benchmark generator for evaluating long-term episodic memory of agents through interactive games. Rather tha...
Zing-0.5: Toward Playable Worlds with Real-Time Joint Action and Text Control
Zing-0.5 is a 5B autoregressive world model that enables users to explore and influence generated worlds through joint keyboard and online text control. It integrates unified action and text conditioning, event-scale supervision for incremental generation, and low-cost real-time interaction, achieving high scores on WBench Navigation. The authors release model weights, inference code, and a serving implementation to support further research on playable generated worlds.
WaLLM -- Understanding Use and Engagement with a General-Purpose LLM on WhatsApp
The paper introduces WaLLM, a general‑purpose large language model chatbot deployed on WhatsApp to investigate how users interact with open‑ended AI in everyday settings. The study finds that health and well‑being queries dominate user interactions, and that proactive communication and communal lists—adapted to WhatsApp’s affordances—enhance engagement and content discovery. These insights inform the design of general‑purpose LLM services on messaging platforms.
Structured Feedback Improves Repair in an LLM Agent Loop
arXiv:2607. 14167v1 Announce Type: cross Abstract: LLM agents often retry after external validation rejects a candidate, but the interface between validation and the next model call remains underspecified.