arXiv Machine Learning

Long-Term Behavioral Evaluation for Trusted Collaborator Selection via Bidirectional Mamba

The paper introduces a bidirectional Mamba-enabled model (BM) for long‑term behavioral evaluation of devices in collaborative tasks. By constructing short‑time‑slot graphs of device interactions and aggregating behavioral features, BM integrates forward and backward temporal dependencies across all intervals. Experiments show that BM outperforms baseline methods, improving the accuracy of selecting trustworthy collaborators to maximize task completion value.

arXiv Machine Learning
Aug 27

TrustFormer: Cross-Temporal and Cross- Dimensional Transformer for Task-Specific Multi-Dimensional Trust Evaluation

TrustFormer is a task‑specific framework that evaluates trust across multiple dimensions in dynamic collaborative systems. It synchronizes heterogeneous trust data using task identifiers and timestamps, then applies cross‑temporal and cross‑dimensional attention to model both temporal dynamics and inter‑dimensional correlations. By combining these multi‑dimensional trust profiles, the system selects optimal collaborators and achieves a 40.8% improvement in trust evaluation accuracy over existing methods.

By Botao Zhu, Xianbin Wang
arXiv Machine Learning
Aug 27

Multi-View Trust Evaluation for Collaborator Selection via Evidential Deep Learning

The paper introduces Multi-View Evidential (MVE) learning for evaluating trustworthiness of collaborators in distributed systems. It models each task owner’s interaction as an independent view, uses the Mamba model to capture temporal trust dynamics, and applies evidential deep learning to quantify uncertainty. A dynamic fusion strategy then combines view-specific evidence to produce a final trust assessment, outperforming baselines in accuracy and task success rate.

By Botao Zhu, Xianbin Wang
arXiv AI
Sep 25

Qwen-Planner-Agent: A Closed-Loop AI-for-AI Framework for Real-World Mobile Planner Agents

Qwen‑Planner‑Agent is a closed‑loop AI‑for‑AI framework that enables large language models to act as both developers and participants in building advanced AI systems. The framework integrates data production, model training, and deployment through a shared action‑feedback‑verification contract, employing AI‑for‑Data, AI‑for‑Training, and AI‑driven model‑harness co‑evolution. It achieves top performance on MobilePA‑Bench by improving tool use, memory, skills, and sub‑agent coordination, while also showing gains on non‑mobile benchmarks.

By Tingyu Qu, Weigao Sun, Yuecheng Liu, Yucheng Zhao, Yi Zhu, Yifeng Ding, Qiyi Wang, Sihan Cao, Pengkun Jiao, Hanlei Xie, Xiongwei Wu, Qichao Wang, Haodong Zhang, Jiajun Liu, Yuhao Wang, Yuqing Xie, Junpeng Zhao, Long Chen, Ming Ma, Sihan Yang, Ziwang Zhao, Yanhao Jia, Liangquan Gong, Feida Zhu, Yiran Zhong, Steven Hoi
arXiv AI
Jun 6

Humans' ALMANAC: A Human Collaboration Dataset of Action-Level Mental Model Annotations for Agent Collaboration

arXiv:2606. 06388v1 Announce Type: new Abstract: Recent advances in LLM agents have enabled complex cognitive capabilities, such as multi-step reasoning, planning, and tool use, that increasingly position these agents as human collaborators.

By Jiaju Chen, Yuxuan Lu, Jiayi Su, Chaoran Chen, Songlin Xiao, Zheng Zhang, Yun Wang, Yunyao Li, Jian Zhao, Tongshuang Wu, Toby Jia-Jun Li, Dakuo Wang, Bingsheng Yao
arXiv AI
Sep 21

DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-Refinement

arXiv:2609.21423v1 Announce Type: new Abstract: Online agent deployments produce abundant execution traces, while task-specific verification and expert annotation are costly to scale. We study how to...

By Siyuan Liu (Fudan University, Meituan Longcat Team), Fan Yu (Fudan University, Meituan Longcat Team), Dongyu Ru (Meituan Longcat Team), Yizhu Liu (Meituan Longcat Team), Yifan Yang (Meituan Longcat Team), Xuezhi Cao (Meituan Longcat Team), Xunliang Cai (Meituan Longcat Team), Yixin Cao (Fudan University)