CollabSkill: Evaluating Human-Agent Collaboration On Real-World Tasks
arXiv:2606. 09833v1 Announce Type: cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work.
Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.
arXiv:2606. 09833v1 Announce Type: cross Abstract: AI agents are reshaping the workspace, leading to drastic change of how humans work.
arXiv:2606. 10862v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models achieve strong performance on standard manipulation benchmarks, but most evaluations assume that task-relevant objects are fully visible.
arXiv:2606. 10229v1 Announce Type: cross Abstract: We study whether demonstration-curation metrics that detect defective training episodes also improve the downstream behavior-cloning policy that trains on the curated data.
arXiv:2606. 10554v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed in real-world applications that require access to up-to-date knowledge.
arXiv:2606. 10614v1 Announce Type: cross Abstract: Robotic foundation models pre-trained on human demonstration videos have shown promise, but a significant embodiment gap remains when the resulting policies are deployed on real robots.
arXiv:2606. 11042v1 Announce Type: new Abstract: Recent years have witnessed the rapid evolution of AI agents toward handling increasingly complex, real-world tasks.
arXiv:2606. 10768v1 Announce Type: new Abstract: The success of Large Language Models in mathematical reasoning relies heavily on the generation of diverse and valid solution paths during the rollout phase.
arXiv:2606. 10611v1 Announce Type: new Abstract: Traditional heuristic solvers for the 2D irregular nesting problem share a fundamental limitation: they are blind to polygon geometry, relying on guided brute-force to navigate the continuous placement space with minimal geometrical guidance.
arXiv:2605. 20347v2 Announce Type: replace Abstract: Labeling a training set is often expensive and susceptible to errors, making the design of robust loss functions for label noise an important problem.
arXiv:2606. 09917v1 Announce Type: new Abstract: Multivariate time series forecasting requires capturing the continuously evolving correlation structure among interacting variables.
arXiv:2606. 10774v1 Announce Type: new Abstract: Decentralized Federated Learning (DFL) over lossy wireless networks faces two key challenges: selection bias, where updates from poor-quality links are systematically underrepresented due to partial model reception, and update staleness, where asynchronous nodes contribute outdated information.
arXiv:2602. 09639v2 Announce Type: replace Abstract: Denoising diffusion models (DDMs) are state-of-the-art methods for learning densities from data across numerous domains, yet many aspects of the training and sampling pipeline remain poorly understood.
arXiv:2606. 10087v1 Announce Type: cross Abstract: Pre-training on raw code teaches syntax but provides sparse signal for diverse real-world task formats.
arXiv:2606. 10466v1 Announce Type: cross Abstract: In time-series generation, existing approaches typically handcraft ortrain a separate model for each dataset, which hinders their scalability and fails to leverage shared temporal structures across domains.
arXiv:2606. 10799v1 Announce Type: new Abstract: Large Language Models (LLMs) struggle to rigorously verify complex mathematical proofs.
arXiv:2606. 10099v1 Announce Type: cross Abstract: The rapid development of large language models (LLMs) has raised concerns about misuse such as plagiarism, misinformation, and automated influence operations, motivating the need for robust detectors.
arXiv:2606. 09884v1 Announce Type: cross Abstract: We study two reproducible failure modes of deep multi-agent reinforcement learning in continuous-time pricing markets: (i) tacit cartel formation between competing DDPG agents, and (ii) actor--critic instability at high event rates.
arXiv:2606. 10254v1 Announce Type: new Abstract: While Large Language Models (LLMs) have achieved near-perfect performance in \emph{solving} high-school mathematics, their ability to \emph{evaluate} the diverse reasoning processes of real human students remains under-examined.
arXiv:2606. 11182v1 Announce Type: cross Abstract: In this paper, we propose EEVEE, the first multi-dataset test-time prompt learning framework for LLM agents, enabling test-time prompt learning under real-world task streams.
arXiv:2606. 10824v1 Announce Type: new Abstract: The Euler Characteristic Curve (ECC) records the Euler characteristic of a linearly embedded cell complex as a function of filtration height in a given direction, and the Euler Characteristic Transform (ECT) is the injective shape descriptor obtained by collecting ECCs over many directions.