TraceBench: Controlled Evaluation of LLM Agents for Time-Series Root-Cause Attribution
Read the original on arXiv Machine Learning →TraceBench is a simulation-based framework that generates controlled root‑cause attribution tasks for time‑series data. In each task, an LLM agent must determine whether a system parameter was altered during a simulation of a physical dynamical system and identify the altered parameter. The authors evaluated four LLM agents on tasks derived from three interpretable mechanical systems, finding that agents perform better with domain context, rely mainly on numerical console output, and struggle more when required to produce Python scripts for labeling than when submitting direct predictions.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.