Temporal Logic Guided Universal Task Representations for Reinforcement Learning
arXiv:2608. 15509v1 Announce Type: cross Abstract: Task guided agents demonstrate strong performance in a wide range of complex tasks.
Policy optimisation, reward modelling and RLHF — how models are trained by feedback rather than by labels.
arXiv:2608. 15509v1 Announce Type: cross Abstract: Task guided agents demonstrate strong performance in a wide range of complex tasks.
arXiv:2608. 16492v1 Announce Type: cross Abstract: This paper studies the regret analysis for parallel Gaussian process (GP) bandit optimization.
arXiv:2608. 14945v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) gives language agents dense token-level supervision from a privileged self-teacher on the policy's own trajectories.
arXiv:2608. 14795v1 Announce Type: new Abstract: An AI that can only give advice seems safe: the human is always free to ignore it.
arXiv:2608. 14828v1 Announce Type: new Abstract: Aligning a language agent to several objectives at once is a persistent failure mode of preference-based training: when objectives are combined additively, optimization collapses onto whichever is cheapest to improve and sacrifices the rest, so a support agent learns to sound warm while giving no real help.
arXiv:2608. 15817v1 Announce Type: new Abstract: The growing ecosystem of large language models (LLMs) offers huge potential to optimize performance-cost trade-offs.
arXiv:2608. 15592v1 Announce Type: new Abstract: Efficient LLM serving is often bottlenecked by the need to pad sequences to a fixed maximum length, and this wastes compute and degrades throughput.
arXiv:2608. 15265v1 Announce Type: new Abstract: Constructing an interactive 3D open world from a user query is important.
arXiv:2608. 15372v1 Announce Type: new Abstract: We study generating game-theoretically optimized Courses of Action (COAs) for a Blue UAS swarm against an adaptive Red adversary in a communication-degraded environment, motivated by (but not derived from) a public U.
arXiv:2608. 15291v1 Announce Type: new Abstract: Demand forecasting increasingly requires combining two complementary sources of information: historical sales reveal recurring numerical dynamics, while future promotions, holidays, price changes, and platform interventions provide forward-looking knowledge.
arXiv:2608. 16156v1 Announce Type: new Abstract: Long-horizon large language model (LLM) agents are typically optimized with sparse terminal outcomes, making fine-grained credit assignment across multi-step interactions difficult.
arXiv:2608. 14559v1 Announce Type: new Abstract: Effective communication in multi-agent reinforcement learning requires agents to decide not only \textit{what} to communicate, but when?
arXiv:2608. 15930v1 Announce Type: new Abstract: Foundation GUI agents can automate complex digital tasks, but deployment is hindered by scarce and biased training data, ambiguous prompts, and unreliable execution.
arXiv:2608. 14791v1 Announce Type: new Abstract: Abductive reasoning, often characterized as inference to the best explanation, is central to explanation under uncertainty, from everyday sense-making and investigation to scientific discovery.
arXiv:2608. 16482v1 Announce Type: new Abstract: The dosing of intravenous fluids and vasopressors in sepsis is a sequential decision made under uncertainty and guided largely by clinical judgment, which makes it a natural target for reinforcement learning from historical care.
arXiv:2608. 16626v1 Announce Type: new Abstract: Radio frequency identification (RFID) technology has been widely implemented for real-time data collection in manufacturing shop floors, which, in turn, can be used to support dynamic shop floor production planning and scheduling.
arXiv:2608. 16435v1 Announce Type: new Abstract: In collection operations, accumulating payload progressively slows the vehicle, imposing a cumulative penalty on routing efficiency.
arXiv:2311. 02629v5 Announce Type: replace Abstract: We introduce the Pointer Q-Network (PQN), a hybrid neural architecture that integrates model-free Q-value policy approximation with Pointer Networks (Ptr-Nets) to enhance the optimality of attention-based sequence generation, focusing on long-term outcomes.
arXiv:2608. 15088v1 Announce Type: cross Abstract: Human-in-the-loop (HIL) online reinforcement learning for real robots must absorb human interventions quickly while continuing to improve beyond the human prior.
arXiv:2608. 14642v1 Announce Type: new Abstract: Reinforcement Learning (RL) agents trained on a single reward signal exploit the gap between the designed reward and the intended behavior.