Large language models

Model releases, architecture work and prompting research on large language models — from frontier-lab announcements to the arXiv papers behind them.

25,737 stories · RSS feed

arXiv AI
1d ago

Natural Language Questions as an Interface for Knowledge Graphs: QRAKEN Graph Distillation and Semantic Self-Healing

The paper introduces QRAKEN, a training‑free, ontology‑agnostic neurosymbolic pipeline that grounds natural‑language queries to RDF knowledge graphs using empirical graph evidence instead of schema expectations. QRAKEN distills a compact TTQL representation offline, then uses it online to guide large language models, providing deterministic checks and an iterative refinement loop. On the CK25 benchmark, QRAKEN achieves a strict F1 of 0.643–0.652 with GPT‑4.1 mini and GPT‑5.4, outperforming state‑of‑the‑art systems and demonstrating the effectiveness of empirical pattern distillation over schema‑only approaches.

By Remo Grillo, Lukas Klic, Giovanni Colavizza
arXiv AI
1d ago

PlaySuite: A Large-Scale Benchmark for Interactive Visual Intelligence

PlaySuite is a large-scale benchmark that evaluates interactive visual intelligence by using over 5,000 open-source video games from platforms like PyWeek and itch.io. The benchmark covers diverse game engines (Pygame, HTML5, Godot, Unity) and introduces a unified closed-loop interaction framework and a Video-LLM-as-a-judge protocol to standardize progress measurement. Evaluation of fourteen recent models shows a perception-action gap, with strong reasoning but poor sustained progress, spatial grounding, action execution, and self-correction.

By Dheeraj Varghese, Anna Vettoruzzo, Walter Simoncini, Michelle Lorena Acevedo Callejas, Mohammad Mahdi Derakhshani, Kristof Meding, Joaquin Vanschoren, Cees G. M. Snoek
arXiv AI
1d ago

Contrastive Learning for Aspect Representation towards Explainable Recommendation

The paper introduces CLARER, a recommendation model that fuses aspect features extracted from textual reviews with rating data to enhance recommendation accuracy and explainability. It learns user and item representations by combining rating-based features via an MLP and aspect-based features via a transformer encoder with contrastive learning. A transformer decoder then generates explanations using the combined representations, and experiments on three benchmark datasets show superior performance over baseline methods in both recommendation accuracy and explanation generation.

By Emrul Hasan, Chen Ding
arXiv AI
1d ago

Tool-calling retrieval versus vector RAG for a small Greek--English knowledge base: accuracy and robustness to how users type Greek

The study compares two methods for grounding assistants in a small Greek–English agricultural knowledge base: tool‑calling retrieval via a live data interface and vector retrieval‑augmented generation (RAG). Using the KyGround benchmark of 198 questions, vector RAG achieved 95.3% accuracy on canonical Greek questions, outperforming the tool agent’s 71.6% and revealing that the tool agent’s failures stem from literal searches that miss non‑verbatim matches. The results show that search tolerance to user typing variations—such as accents, capitalization, and Greeklish transliterations—is crucial for reliable community knowledge interfaces.

By Nikolaos D. Tantaroudas, Ilias Karachalios, Andrew J. McCracken
arXiv AI
1d ago

Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution under Analysis Budgets

The paper introduces a cost‑aware Hierarchical Multi‑Agent System (HMAS) for ransomware detection and family attribution that prioritizes static analysis and escalates to dynamic and memory analysis only when necessary, thereby reducing analysis time and resource usage. In experiments on 12,439 samples from 16 ransomware families, the deterministic HMAS achieved high F1 scores (0.93) while resolving nearly 58% of cases with static evidence alone and cutting average internal analysis cost by 44.6% compared to exhaustive methods. The system also records a complete provenance trace for each decision and includes optional local LLM review for limited verdict adjustment.

By Mubashar Iqbal, Asifullah Khan, Hifsa Asif, Saddam Hussain Khan, Umme Zahoora, Asifullah Khan
arXiv Machine Learning
1d ago

When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO

The paper introduces Guidance-Augmented GRPO (GA‑GRPO), a theoretical framework that unifies several external‑guidance methods for reinforcement learning with verifiable rewards (RLVR) used to improve large language model reasoning. GA‑GRPO treats external guidance as a stochastic operator that rewrites the question distribution, yielding a biased on‑policy policy‑gradient estimator whose bias is bounded by the total‑variation guidance divergence. The authors prove convergence rates, derive an optimal guidance‑weight formula, and validate their predictions experimentally on a math‑reasoning benchmark, showing that the optimal‑weight GA‑GRPO outperforms existing methods while saving GPU time.

By Sofia Torres, Gabriel Almeida, Carter Adams, Camila Rocha
arXiv Machine Learning
1d ago

Improving Synthetic Data Generation for Argument Mining via Adversarial Reinforcement Learning

The paper introduces an adversarial reinforcement learning framework to generate synthetic data for Argument Mining (AM). By jointly training a generator and a discriminator, the system produces structured AM instances that are both accurate and diverse. Experiments show consistent performance gains on three benchmark datasets in both full-data and low-resource scenarios.

By Zhijun Zhang, Qianlong Wang, Keyang Ding, Genan Dai, Bowen Zhang, Bin Liang, Ruifeng Xu, Yongsheng Liang
arXiv Machine Learning
1d ago

A Systematic Study of Small Language Models on Abstract Reasoning Tasks

The paper investigates how small language models acquire abstract reasoning skills on the ARC‑TGI benchmark, which groups grid‑transformation tasks into controllable families and allows resampling, spatial shifts, and cross‑benchmark transfer. Over 1,000 supervised fine‑tuning runs across decoder‑only, encoder‑decoder, and mixture‑of‑experts families, the study finds that high in‑distribution accuracy is possible but depends heavily on optimization and is uneven across task families. Performance drops sharply outside the training distribution, and gains from larger training sets or additional in‑context examples vary by model family; attention diagnostics reveal distinct patterns but do not explain causal mechanisms.

By Nur A Zarin Nishat, Jens Lehmann, Andrei Aioanei, Sahar Vahdati
arXiv Machine Learning
1d ago

AdaLoop: Adaptive-Depth Latent Reasoning for Audio Language Models

AdaLoop is a lightweight recurrent module that adaptively determines how many latent refinement steps are needed for audio–question pairs, allowing deeper reasoning only when necessary. It shares a transformer block that iterates over the audio representation guided by the question, with a learned halting mechanism that exits the loop once the representation is ready. Adding fewer than 3 % of the base model’s parameters, AdaLoop improves average accuracy by 2.9 to 3.8 points across three distinct models, especially on perception-heavy subtasks.

By Lee Seung-woo, Bowen Qi
arXiv Machine Learning
1d ago

Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

The paper investigates whether frozen video‑language models inherently encode a signal indicating whether sufficient evidence has been observed to answer a question. By training linear probes on seven byte‑identical models, the authors demonstrate that these models contain a readable evidence‑readiness signal with AUROC ranging from 0.733 to 0.905, even when the probe is trained without any footage from the benchmark family. The signal is question‑conditioned, remains robust when the model answers incorrectly, and outperforms traditional uncertainty estimators; it can be leveraged as a Readiness Gating policy that improves answer accuracy by up to 9.75 percentage points without extra computational cost.

By Dan Ben-Ami, Kobi Cohen, Chaim Baskin
arXiv Computation and Language
1d ago

Large Language Model Orchestration under Heterogeneous Preferences via Explicit Persona Inference

The paper introduces HARP, a framework for orchestrating heterogeneous agents with hidden preferences by maintaining numeric posterior beliefs updated via Bayes’ rule, rather than embedding beliefs in prompts. HARP achieves ∼O(√K) Bayesian regret and, with the HARP+ variant, adds a bonus for informative actions to keep inference active even when optimal actions are uninformative. Experiments across three problem settings show HARP+ outperforms other non‑oracle methods in scenarios where explicit joint inference is infeasible.

By Shuqing Shi, Ziyan Wang, Milind Tambe, Yali Du
arXiv Computation and Language
1d ago

AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking

AVMeme Exam is a human‑curated benchmark featuring over a thousand iconic Internet audio‑visual clips—including speech, songs, music, and sound effects—each paired with a unique Q&A that probes understanding from surface content to context, emotion, usage, and world knowledge. The benchmark also provides metadata such as original year, transcript, summary, and sensitivity. Evaluations of state‑of‑the‑art multimodal large language models (MLLMs) and human participants reveal that current models perform poorly on textless music and sound effects and struggle to think in cultural and contextual terms compared to surface content.

By Xilin Jiang, Qiaolin Wang, Junkai Wu, Xiaomin He, Zhongweiyang Xu, Yinghao Ma, Minshuo Piao, Kaiyi Yang, Xiuwen Zheng, Riki Shimizu, Yicong Chen, Arsalan Firoozi, Gavin Mischler, Sukru Samet Dindar, Richard Antonello, Linyang He, Tsun-An Hsieh, Xulin Fan, Yulun Wu, Yuesheng Ma, Chaitanya Amballa, Weixiong Chen, Jiarui Hai, Ruisi Li, Vishal Choudhari, Cong Han, Yinghao Aaron Li, Adeen Flinker, Mounya Elhilali, Emmanouil Benetos, Mark Hasegawa-Johnson, Romit Roy Choudhury, Nima Mesgarani
arXiv Computer Vision
1d ago

Event Detection in Table Tennis Videos using 2D Keypoints

The paper introduces EventNet, a two‑stage pipeline that uses 2D keypoints of players, table corners, and the ball to detect key events in table tennis videos. First, a keypoint transformer condenses the pose and ball information into a robust representation; second, a transformer encoder predicts how close each frame is to the next and previous ball‑racket contact using a novel temporal cosine‑like target signal. Experiments on Latte‑MV and TTHQ datasets show high accuracy, with an F1 score of 91.16% and a mean frame deviation of 0.42 on Latte‑MV, and 73.08% / 1.16 on TTHQ.

By Rainer Lienhart, Daniel Kienzle, Shin'ichi Satoh, Anastasiia Bilinska
arXiv AI
1d ago

Explore, Then Commit: Measurement-Efficient Scientific Law Discovery with Language Models

The paper presents an explore‑then‑commit protocol that uses a large language model to generate hypotheses, a programmatic planner to collect measurements, and a fresh prompt to synthesize a scientific law from fixed observations. In 576 NewtonBench trials across 12 physics modules, the protocol—especially when interpreter‑enabled planners are used—reduces the number of measurements needed and improves root‑mean‑squared logarithmic error for both GPT‑4.1‑mini and GPT‑4.1. The study demonstrates measurement savings in every module, though it notes that the causal components and generalization beyond noiseless direct‑equation tasks remain unresolved.

By Kautik Mandve, Dileepa Fernando
arXiv AI
1d ago

Evidence Before Sampling: Interpretable Implicit Negative Candidate Discovery for Recommendation

The paper introduces a method for discovering implicit negative candidates in recommender systems by extracting symbolic rules from observed customer behavior. These rules are scored on support, informativeness, and product relevance, then interpreted by a large language model to align with business objectives. Experiments in an industrial B2B setting and on five public datasets show that the approach improves precision and downstream PR-AUC compared to baseline negative sampling methods.

By Shreya Rajpal, Sonia Sharma, Swapnil Parekh, Lisa Li, Jeyendran Balakrishnan, Nagaraj Janardhana, Andrew Mattarella-Micke
arXiv AI
1d ago

DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks

DAEDALUS is a method that builds reusable memory for large‑language‑model agents by having an explorer agent generate self‑created tasks and a solver agent attempt them. When the solver fails, a heuristic is extracted and only accepted after repeated successful use, then added to a memory bank for future test‑time use. Experiments on AppWorld, τ²‑bench, and AutomationBench show that DAEDALUS raises mean success rates by up to 15.9 points and pass⁵ by up to 2.2× compared to a no‑memory baseline, while also providing a cost‑effective alternative to training‑task or oracle‑verifier approaches.

By Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud