TextQuests: How Good are LLMs at Text-Based Video Games?
Related stories
Introducing Agents.js: Give tools to your LLMs using JavaScript
LLMs help robots understand vague instructions and focus on key details
To help robots do chores in places like homes and factories, a new approach from MIT uses one language model to clarify users’ instructions, then another to ignore irrelevant info.
Judge Arena: Benchmarking LLMs as Evaluators
TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment
LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services. The domain-wide adoption of LLMs necessitates continuous evaluation to ensure their safety and fairness.
Open-source LLMs as LangChain Agents
Open-Source Text Generation & LLM Ecosystem at Hugging Face
How Long Prompts Block Other Requests - Optimizing LLM Performance
Structured Feedback Improves Repair in an LLM Agent Loop
arXiv:2607. 14167v1 Announce Type: cross Abstract: LLM agents often retry after external validation rejects a candidate, but the interface between validation and the next model call remains underspecified.
TriEval: A Resource-Efficient Pipeline for LLM Bias, Toxicity, and Truthfulness Assessment
arXiv:2606. 03036v1 Announce Type: new Abstract: LLMs have evolved from basic chatbots to the backbone of the AI ecosystem, now widely used in healthcare, schools, and government services.
Security and Privacy Prompts in the Wild: What Users Ask LLMs and How LLMs Respond
arXiv:2606. 18062v1 Announce Type: cross Abstract: Large language models (LLMs) are widely used to fulfill users' information needs; users ask LLMs about the weather, pose educational questions, and consult them for legal assistance.
Controlling Reasoning Effort in LLMs
How LLMs Learn Low-, Medium-, and High-Effort Reasoning Modes
