arXiv AI

Bounded-Fidelity Sim-as-Demo-Stage: Mocap Handoff for Governance Benchmarks

The paper introduces a bounded‑fidelity sim‑as‑demo‑stage design pattern that suppresses contact physics during object handoffs in simulators, using MuJoCo’s mocap‑body primitive and a lightweight Python adapter. This approach ensures audit‑chain stability, producing identical event‑log hashes across 1,000 replays per posture, whereas a contact‑force baseline yields significant divergence. The authors demonstrate that the pattern maintains reproducibility across various timesteps and sequential handoffs with minimal overhead, and they identify contexts where it should not be applied.

arXiv AI
Aug 18

DeepInsight II: One Trace from Benchmark to Robot

arXiv:2608. 16556v1 Announce Type: new Abstract: Across a Physical AI stack, evaluation maturity is inversely aligned with deployment risk: foundation models enjoy mature, standardized harnesses, while the embodied layers on which deployment actually turns remain fragmented across benchmark-specific simulators, embodiments, and interfaces.

By Siyi Li, Yuchen Kang, Wuliang Wang, Zhengjie Zhang, Jiangpin Liu, Jianhao Yao, Jie Chen
Hugging Face Trending Papers
Sep 8

Hi-FLoop: Hierarchical State-Feedback Loops for Multi-Timescale World Modeling

Hi-FLoop introduces a hierarchical state‑feedback framework for multi‑agent traffic simulation that reconciles decision time scales over an 8‑second rollout. The model uses eight scene‑level Worlds to maintain joint hypotheses, with an 8‑second Goal, 2‑second Preview, and 1‑second Control hierarchy, and commits only executed prefixes every 0.5 seconds to preserve factual consistency. A joint preview interaction graph and a prefix‑frozen A‑to‑B cascade enable sparse interaction refinement and accurate state recovery, achieving an overall score of 0.689987 on the H‑D public‑validation split and strong oracle‑minADE performance. whyItMatters":"The paper presents a novel multi‑timescale approach that improves consistency and realism in long‑horizon traffic simulations, as evidenced by its competitive evaluation metrics."

arXiv AI
Sep 25

Stale Does Not Mean Unsafe: Guard Precision for Tool-Using LLM Agents under Infrastructure State Races

The paper investigates how tool‑using language‑model agents can safely commit changes to infrastructure when external state may change between read and commit. By distinguishing invalidating races from predicate‑preserving and irrelevant ones, the authors evaluate three commit‑time guard granularities—global epoch, read‑set version, and semantic commit predicate—using a deterministic simulator and three quantized model families. The study finds that only the complete predicate guard consistently eliminates unsafe commits, while freshness‑based guards block a large proportion of benign races and model‑side signals fail to replace precise semantic enforcement.

By Zihao Zheng, Jiayu Long, Baichuan Li, Junyi Yao
arXiv AI
4d ago

Mingbird: A Local-First Agent Harness Enabling Small Open Models to Complete Real Tasks

Mingbird is a local‑first agent harness designed for small open‑weight language models (2–9 B) that run on ordinary laptops. It introduces ten mechanisms—such as a byte‑level net‑zero prefill budget, a finish gate that re‑reads the task before accepting completion, and signature‑level loop detection—to address common failure modes that arise from the harness rather than the model itself. In controlled experiments on the LRAB benchmark and the $ au^2$‑bench, Mingbird achieves higher overall scores (0.886 and 0.856 respectively) compared to other harnesses, and its ablation studies show that each mechanism contributes measurable performance gains.

By Hao Wang, Ting Huang