LLM-Derived Priors for Thompson Sampling in Cold-Start Comment Recommendation
arXiv:2608. 03382v1 Announce Type: cross Abstract: Multi-armed bandit algorithms, especially Thompson sampling, are widely used in online recommendation.
Policy optimisation, reward modelling and RLHF — how models are trained by feedback rather than by labels.
arXiv:2608. 03382v1 Announce Type: cross Abstract: Multi-armed bandit algorithms, especially Thompson sampling, are widely used in online recommendation.
arXiv:2608. 03437v1 Announce Type: cross Abstract: While human evaluation is the gold standard in many NLP tasks, it suffers from prohibitive costs and poor scalability.
arXiv:2608. 03483v1 Announce Type: cross Abstract: Existing chunk-based Vision-Language-Action (VLA) models execute a fixed number of actions (i.
arXiv:2602. 17976v2 Announce Type: replace-cross Abstract: In active sequential testing, also termed pure exploration, a learner is tasked with the goal to adaptively acquire information so as to identify an unknown ground-truth hypothesis with as few queries as possible.
arXiv:2608. 03674v1 Announce Type: new Abstract: Causal diagnostic models must explain how conclusions follow from evidence because diagnoses guide repairs and treatments.
arXiv:2608. 03682v1 Announce Type: new Abstract: Physical AI policies require inference throughout their lifecycle, including model evaluation, cloud reinforcement learning rollout, edge GPU serving, and onboard deployment.
arXiv:2608. 03527v1 Announce Type: cross Abstract: Retrieval systems help deep research agents generate high-quality answers by providing relevant documents.
arXiv:2607. 27703v2 Announce Type: replace Abstract: Vision-language models (VLMs) are increasingly used in embodied agents to interpret visual inputs, reason about spatial relationships, and make task-level decisions based on that reasoning.
arXiv:2608. 03506v1 Announce Type: new Abstract: Self-consistency assumes the most frequent answer among sampled reasoning traces is the most reliable, but this can fail in causal reasoning: samples often repeat the same confounding error, and votes fragment across multiple valid answers, letting an invalid answer win despite a valid minority trace.
arXiv:2608. 03606v1 Announce Type: new Abstract: Clinical development is sequential decision-making under uncertainty, where a sponsor must plan a portfolio of experiments from heterogeneous evidence.
arXiv:2608. 03119v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability.
arXiv:2608. 03929v1 Announce Type: new Abstract: Aligning diffusion models with human preferences usually relies on a sparse terminal reward evaluated on the final generated samples, presenting a severe temporal credit-assignment challenge across the multi-step denoising process.
arXiv:2511. 09173v3 Announce Type: replace-cross Abstract: External trajectories can improve offline decision-sequence learning, but dynamics shift may make some source subsequences inconsistent with the target environment.
arXiv:2608. 03468v1 Announce Type: new Abstract: Historical tool-use trajectories provide valuable experience for large language model (LLM) agents to plan and coordinate tool usage.
arXiv:2608. 03036v1 Announce Type: cross Abstract: Large Language Models (LLMs) are integrated into software systems and AI services, making efficient LLM serving a concern for software engineering.
arXiv:2608. 02993v1 Announce Type: new Abstract: (Flat) Reinforcement Learning (RL) agents face significant challenges in environments with sparse rewards that require long-horizon reasoning.
arXiv:2608. 03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features.
arXiv:2608. 03875v1 Announce Type: cross Abstract: Designing effective reward functions remains a major bottleneck in Reinforcement Learning (RL).
arXiv:2508. 20697v4 Announce Type: replace Abstract: As large language models (LLMs) continue to grow in capability, so do the risks of harmful misuse through fine-tuning.
arXiv:2608. 03092v1 Announce Type: cross Abstract: We aim to improve model performance in multi-reward reinforcement learning training process.