ChartBmkAgent: Harness-Governed Multi-Agent Construction of Chart QA Benchmarks from Sparse Error-Taxonomy Specifications
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2608.03464v2 Announce Type: replace Abstract: Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large...
arXiv:2609.24172v1 Announce Type: new Abstract: Legends are fundamental to chart understanding, as reliable interpretation requires correctly binding legend entries to corresponding visual marks. Whi...
arXiv:2606. 27383v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used as research assistants, yet it remains unclear whether they can calibrate research takeaways to the strength and scope of the supporting evidence.
arXiv:2605. 21347v3 Announce Type: replace Abstract: Diagnosing failures in LLM agents remains largely manual.
VisAudit is a new benchmark that tests multimodal agents on visual diagnosis, repair, and verification tasks. It presents agents with rendered charts and auxiliary evidence—such as source data, intended summaries, and code—to iteratively detect defects, modify the visualization, and confirm successful repairs. The benchmark includes 1,900 flawed charts across 21 types and 10 flaw categories, plus 300 correct charts, and shows that current models recover only about 47.4% of flawed charts autonomously.
arXiv:2607. 16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance.