arXiv AI By Shih-Yu Lai, Wen-Fan Wang, Sai Ling, Shaune Jan, Bing-Yu Chen, Xiang Anthony Chen

MOONWALK: Mediating Operations with Intent-Evidence-Action Alignment Across Junior-Supervisor Review Workflows in Animation/VFX Pre-Production

Read the original on arXiv AI →

MOONWALK is a pre‑production review system for animation and VFX that aligns creative intent, evidence, and action across junior‑supervisor workflows. It records shared intent, anchors judgments to evidence, and translates decisions into clear revision tasks tied to reference notes, with AI handling administrative coordination. A studio study shows that MOONWALK improves intent alignment, decision traceability, and checklist executability compared to a chat‑only interface, while keeping aesthetic authority with practitioners.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 3

DeskCraft: Benchmarking Desktop Agents on Professional Workflows and Human-in-the-Loop Collaboration

arXiv:2606. 03103v1 Announce Type: new Abstract: Real-world professional desktop workflows in specialized creative and engineering software unfold over long horizons and often require human-in-the-loop coordination, where agents proactively seek necessary information and users provide additional instructions, clarifications, feedback, or corrections as the task progresses.

By Wenkai Wang, Tao Xiong, Jingchen Ni, Yunpeng Bao, Xiyun Li, Tianqi Liu, Hongcan Guo, Zilong Huang, Shengyu Zhang
arXiv AI
Aug 25

Designing Benchmarks for Knowledge Work

The paper proposes a new way to describe benchmarks for AI systems that perform knowledge work, outlining four explicit fields: represented activity, tested setting, required work product, and evaluated result. It builds an inventory of 18 work activities from O*NET to enable activity-level reporting across occupations, and evaluates these activities for semantic coherence, algorithm sensitivity, ontology legibility, and human interpretability. The authors apply their framework to three existing benchmarks—GDPval, OfficeQA Pro, and APEX-SWE—to show how different aspects of work can be captured within the same representation.

By Yining Hua, Hongbin Na, Cyrus Ayubcha, Levi Lian