arXiv:2603.17815v2 Announce Type: replace
Abstract: Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) pr...
By Corentin Royer (International Business Machines), Anna Hedstr\"om (ETH AI Center), Debarun Bhattacharjya (Lirio), Gaetano Rossiello (International Business Machines), Andrea Giovannini (International Business Machines), Mennatallah El-Assady (Department of Computer Science, ETH Zurich)
arXiv:2609.37603v1 Announce Type: cross
Abstract: A common safeguard for a data pipeline is redundant computation: derive each published number by two routes built on different technology and refuse...
By Fabio Rovai
arXiv:2609.37223v1 Announce Type: new
Abstract: Credit-risk prediction is important in banking, but a prediction alone does not explain why an applicant is risky or how it should be combined with oth...
By Aakash Kumar Tiwari
arXiv:2512.00329v2 Announce Type: replace-cross
Abstract: Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose an approach that recasts the task...
By Ashish Thanga, Vibhu Dixit, Abhilash Shankarampeta, Vivek Gupta
AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.
By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
arXiv:2609.37989v1 Announce Type: new
Abstract: Tabular foundation models achieve strong zero-shot accuracy on structured data by pretraining on synthetic tables, but they ignore the column names, ta...
By Deqing Fu, Huangyuan Su, Rajat Sen, Taman Narayan, Sujay Sanghavi, Abhimanyu Das, Weihao Kong
IronLLM-0.6B is a 654‑million‑parameter language model engineered for efficient on‑device inference, featuring a hybrid attention architecture, X‑MTP multi‑token prediction, and a lightweight verification head that yields a 1.48× decoding speedup. Trained on roughly 6.2 trillion tokens with a quality‑oriented pipeline and further refined via Multi‑Domain On‑Policy Distillation, the model adopts an Instruct‑Only design to meet low‑latency requirements. A lighter variant, IronLLM‑0.6B‑Light, replaces RMSNorm with Dynamic Tanh and streamlines costly components to enhance inference and quantization efficiency, offering a strong performance‑efficiency trade‑off for resource‑constrained deployment.
By Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao
arXiv:2605.14854v3 Announce Type: replace-cross
Abstract: Human Mesh Recovery (HMR) is fundamentally ambiguous: under occlusion or weak depth cues, multiple 3D bodies can explain the same image evide...
By Patrick Kwon, Chen Chen
InsightMap is a framework that uses top‑down maps as explicit spatial memory and action‑conditioned prediction targets for language‑guided navigation. It links historical views to labeled map locations and employs a shared multimodal backbone to jointly learn navigation action prediction and post‑action map generation, providing auxiliary training supervision. The approach supports a unified RGB‑D pipeline for navigation, visual question answering, situated reasoning, and 3D grounding, achieving state‑of‑the‑art results on R2R‑CE, RxR‑CE, ScanQA, SQA3D, ScanRefer, and outperforming baselines on the Unitree Go2 platform.
By Hongpei Zheng, Hujun Yin
arXiv:2604.22893v2 Announce Type: replace-cross
Abstract: Traditional ``row-count $\times$ quality coefficient'' approaches fail to capture the nonlinear utility of data for Large Language Model (LLM...
By Minghui Xu, Qi Luo, Kun Li, Zhengyang Shan
The article reports on a benchmark that reduced 1,000 Apache Iceberg files to just six, then measured how this consolidation affected query performance across three SQL workloads. It details the methodology and results of the experiment, highlighting changes in execution speed and resource usage. The findings illustrate the trade‑offs between file count and query efficiency in large data systems.
By Thomas Reid
The paper introduces Q-Target Pretrained Transformers (QTPT), a method that replaces supervised behavior cloning with a Bellman-style Q‑target objective for in‑context reinforcement learning. QTPT retains the context‑conditioned Transformer architecture but learns to estimate action values using rewards and transitions from the context, rather than merely imitating offline actions. The authors provide theoretical analysis in stochastic linear bandits and finite‑horizon MDPs, demonstrating improved robustness to weak or suboptimal data, and empirically show gains over supervised pretraining on controlled RL benchmarks and extensions to D4RL Kitchen and AntMaze.
By Yichen Lin, Xuyuan Xiong, Xue Wang, Xiangfu Meng, Mike Mingcheng Wei, Tao Yao
The Epstein Files Engine is an AI agent developed by the New York Times to help journalists investigate a massive mixed‑media collection released by the U.S. Department of Justice on January 30, 2026, which contains about three million pages of PDFs related to Jeffrey Epstein. The Engine translates reporter questions into Google BigQuery SQL queries across three corpora—Epstein‑related releases, the Times’s archive, and external Epstein‑related news headlines—using an LLM to plan queries and return citation‑rich answers that reporters can verify. Over 100 journalists used the Engine, contributing to at least 20 published stories, and the system includes a Diff method for text‑and‑visual duplicate matching to surface genuinely new information.
By Duy K. Nguyen, Teresa Mondr\'ia Terol, Dylan Freedman, Zach Seward
The paper presents a pipeline for generating multi‑turn synthetic conversations and a self‑improvement loop that uses variance‑based contrastive optimization and a coding agent to refine planning and tool‑use in conversational recommendation agents. This approach improves agent quality by 8% over a manually optimized prompt and has been deployed at Spotify, where it accelerated development cycles. In production, the system achieved a 14% increase in user listening, a 5% rise in weekly active users, and a 5% reduction in skip rate compared to a prior session‑only experience.
By Enrico Palumbo, Alexandre Tamborrino, Victor Ode, Ben Lacker, Adri\`a Casas Escoda, Jeremy Hopple, Marcus Better, James Leoni, Hugo Galv\~ao, Hugues Bouchard, Mounia Lalmas, Jos\'e Luis Redondo Garc\'ia, Abenezer Abebe, Ann Clifton, Anton Blomberg, Henrik Lindstr\"om, Dani Doro, Christine Doig Cardet
REALMS is a conversational system that provides real‑time, exact audience sizing for digital marketers. It uses embedding‑based vector search to retrieve relevant categorical attributes, an LLM‑powered NL2SQL pipeline for accurate query generation over complex nested schemas, and schema standardization for industry‑agnostic deployment. Evaluations on real enterprise data show high recall, accurate SQL execution, and low latency, enabling interactive audience insights that previously required hours.
By Haixu Ma, Aditya Bansal, Shubham Lohiya, Sumit Ranjan
The study examines how practitioners in AI-driven systems define, assess, and manage data quality, revealing six key themes. It highlights shifts in traceability, the use of models as quality assessors, and the emergence of new data objects such as agent context and synthetic data. The research proposes a lifecycle assurance framework to provide evidence that data supports specific AI claims throughout model behavior, judgments, and agent actions.
By Hariharan Gopinath, Jan Bosch, Helena Holmstr\"om Olsson
The paper investigates how two signals—input‑conditional uncertainty and prediction‑label loss—detect different types of data corruption in federated learning. Experiments on ResNet‑20 with CIFAR‑10 and SVHN show that prediction‑label loss excels at spotting persistent random label flips, while expected‑entropy uncertainty better identifies additive image noise. The authors argue that effective federated data‑quality assessment must match the chosen signal to the specific corruption type rather than rely solely on uncertainty measures.
By Bradley Scott, Zeqi Luo, Edmond S. L. Ho
The paper evaluates a production text‑to‑SQL pipeline that uses an LLM as a judge, finding that the deployed gpt‑4o‑mini judge agrees with human annotators only weakly (Cohen’s kappa 0.04 on a disagreement‑enriched set and 0.42 on a random spot‑check). The authors identify a specific failure mode, GRADE‑HALLUCINATION, responsible for most over‑flags, and demonstrate that a self‑hosted Qwen3.6‑27B model achieves substantially higher agreement (kappa 0.72) at a much lower cost. They also show that ensembling judges does not improve performance, and that their audit method flags a significant portion of out‑of‑domain SQLs as potential issues.
By Haowei Liu, Hsin-Tai Wu, Yi Fang
The paper introduces FAST-CAP, a causality‑aware framework for simultaneous speech‑to‑speech translation that combines a factorized S2ST architecture, an adaptive policy, and a new latency metric. It employs a novel data pipeline to generate high‑fidelity, causally aligned segments, improving voice transfer and reducing the need for large training datasets. Experiments on Spanish, German, and French demonstrate that FAST‑CAP outperforms fixed‑policy baselines, achieving up to +1.2 BLEU, 26% lower latency, and a 38.8% relative latency reduction while maintaining speaker fidelity.
By Amir Hussein, Enas Albasiri, Travis M. Bartley, Nourchene Ferchichi, Ke Hu, Harishchandra Dubey, Myungjong Kim, Zhehuai Chen, Oluwatobi Olabiyi, Sanjeev Khudanpur
The paper introduces SPARK, a method for privacy‑preserving continual learning that decouples knowledge retention from privacy correction. SPARK freezes the post‑task distribution and then selectively corrects it to reduce the likelihood of sensitive content while maintaining strong performance on current and past tasks. Experiments show that this approach effectively suppresses PII and preserves continual‑learning utility across various settings.
By Shengtao Wen, Yunying Yang, Xiang Chen, Lingbing Guo, Yu Tian, Sheng-Jun Huang