arXiv AI

ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation

arXiv:2607. 12857v1 Announce Type: cross Abstract: A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty.

arXiv AI
Sep 15

ChartAnno: Benchmarking Multimodal Large Language Models for Chart Annotation Generation

arXiv:2608.03464v2 Announce Type: replace Abstract: Annotations are essential to communicative visualization, helping explain data, emphasize key findings, and guide attention. While multimodal large...

By Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen
arXiv AI
3d ago

ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via Code

ChartRevise is a new dataset and evaluation protocol designed for exact chart editing via code. It contains 92,438 records covering 344 edit types across 20 chart types and three plotting libraries, built using the grammar of graphics and source‑program checks to ensure applicability. The protocol measures atomic requirement completion, detects gratuitous changes and missed coupled updates, and combines these with execution and rendering success to determine exact‑edit success.

By Jiaxiang Tang, Yi Zhou, Chad DeLuca, Rogerio Feris, Ahmed Khalil Omran, Zhi-Li Zhang, Pengyuan Li, Ali Anwar
arXiv AI
Aug 5

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

arXiv:2608. 03464v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored.

By Zhenghan Chen, Zekai Shao, Lidan Tan, Xin Lin, Xingchen Zeng, Yi Shan, Ziyue Lin, Xiaoliang Fu, Xinyuan Liu, Yuetong Guo, Fen Wang, Bongshin Lee, Siming Chen
Hugging Face Trending Papers
Aug 4

ChartAnno: Evaluating MLLMs for Chart Annotation Generation

Multimodal large language models (MLLMs) have made significant progress in chart understanding, generation, and editing, but their ability to annotate existing charts remains underexplored. Annotating charts is a common yet challenging communicative task, requiring models to infer intended messages, interpret chart semantics, and place appropriate textual or graphical elements.

arXiv AI
Sep 16

LSREP: A Longitudinal State-Replay Protocol for Evaluating Conversational Memory, with ICE v2 as an Audited Local-First Architecture

The paper introduces LSREP, a Longitudinal State‑Replay Evaluation Protocol designed to assess how conversational memory evolves over time, incorporating ordered replay, lifecycle schedules, repeated probes, evolving reference answers, and mechanism‑fidelity checks. It applies LSREP to ICE v2, a local‑first memory middleware, and reports that on three ordinary‑density datasets ICE v2 achieves near‑zero mean quality difference from vector‑RAG while using fewer fragments but slightly more prompt tokens, yet fails catastrophically on a dense dataset. In a public diagnostic, ICE v2 underperforms pure vector‑RAG on LongMemEval, revealing significant multi‑session and temporal failures and a quality‑cost trade‑off rather than superior efficiency.

By Deepesh Sonar
Hugging Face Trending Papers
Aug 17

IndicQE-APE: A Benchmark for Quality Estimation and Automatic Post-Editing for Indic Languages

Indic quality estimation (QE) and automatic post-editing (APE) data is spread across separate releases, so no single resource supports training and evaluation across tasks and language pairs on one footing. We consolidate the WMT 2020--2024 shared-task lineage with an extended English--Malayalam resource into \indicqe: $126{,}754$ instances over nine directional pairs, with up to four label types aligned on the same segment, a direct assessment, a human post-edit, word-level OK/BAD tags and an error explanation, and a test set stratified over four difficulty axes.

arXiv Computation and Language
Sep 1

PaperBanana-Interact: Scientific Diagram Refinement with Multi-Turn Human Feedback

PaperBanana-Interact is a multi-agent system designed to refine scientific diagrams through multi-turn human feedback. The authors introduce MTPaperBananaBench, a benchmark with 292 images and 3,518 user requirements, and a user simulator that generates natural language feedback at each turn. Experiments show that PaperBanana-Interact consistently improves diagram quality, outperforming baseline systems by 11.9–18.6 points and reducing forgetting by 3.7–6.2 points.

By Xueqing Wu, Ashwin Balasubramanian, Bingxuan Li, Dawei Zhu, Kai-Wei Chang, Yale Song, Yiwen Song, Rui Meng, Tomas Pfister, Nanyun Peng