arXiv:2610.01039v1 Announce Type: new
Abstract: While recent video generative models can synthesize high-fidelity videos, they struggle to portray plausible physical interactions and the resulting st...
By Jiho Jang, Jinyoung Kim, Nojun Kwak, Kyungjune Kim
Enterprise Representation Simplification (ERS) proposes reducing unnecessary representational complexity in enterprise information while preserving essential data within a defined scope. The paper introduces Enterprise Representation Complexity (ERC), a representation‑neutral model that measures complexity across four dimensions—Objects, Interactions, Behaviors, and Supporting Sources—at both representation and task levels. ERC enables comparison of architectural simplification versus retrieval optimization, supports an economic model of maintenance costs, and demonstrates that lower task‑level ERC can improve AI reasoning accuracy, as shown in Text‑to‑SQL research.
By Terry Dorsey, Kevin Huggins
BudgetSchemaBench is a diagnostic tool for evaluating how different schema‑context budgets affect text‑to‑SQL systems. It automatically derives relevance labels from gold SQL, tests four budgets across 80 databases, and compares three schema representations while keeping table rankings fixed. The study shows that increasing the budget improves execution accuracy, especially for lexical retrieval, and that dense retrieval already captures most needed tables at low budgets.
By Chen Shen
arXiv:2610.01064v1 Announce Type: cross
Abstract: Retrieving the right tables is a prerequisite for Text-to-SQL over realistic databases. Dense table retrievers rank schema elements independently, bu...
By Sandipan De, Abhijit Chakraborty, Sambaran Bandyopadhyay, Vivek Gupta
arXiv:2610.02122v1 Announce Type: cross
Abstract: Real-world enterprise data science and analytics workflows require reasoning across dozens of tables, performing statistical analyses, and acting on...
By Gabriel Tomitsuka, Arman Raayatsanati, Emma Xing, Duke Gand, Joseph J Ma
arXiv:2610.01352v1 Announce Type: new
Abstract: Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven da...
By Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu
arXiv:2610.00873v1 Announce Type: cross
Abstract: In many industrial applications, 1) tabular data is scarce and imbalanced and thus requires synthetic expansion; 2) input distributions drift between...
By Hongyu Cao, Xinyuan Wang, Arun Vignesh Malarkkan, Kunpeng Liu, Haifeng Chen, Yanjie Fu
arXiv:2510.24046v2 Announce Type: replace-cross
Abstract: Existing tabular data generation methods primarily focus on matching statistical distributions between real and synthetic data, often overloo...
By Tu Anh Hoang Nguyen, Dang Nguyen, Tri-Nhan Vo, Thuc Duy Le, Trung Le, Sunil Gupta
arXiv:2609.39257v1 Announce Type: cross
Abstract: Accurate forecasting of electricity production is essential for maintaining the operational efficiency and strategic planning of energy utilities. In...
By Nicolas Vautier, Paul Caron, Nardi Xhepi, F\'elicie Bizeul, Manel Boumghar, Christophe Degouy, Paul Boniol
arXiv:2609.39640v1 Announce Type: cross
Abstract: Cross-lingual transfer describes how knowledge in a source language benefits a target language. Measuring it quantitatively requires broad multilingu...
By Dalton Raphael Harmsen, Swier Garst, Thomas van Osch, Zar\`e Palanciyan, Joaquin Vanschoren
The paper introduces a co‑evolving framework where a target agent improves by learning from its own failures, and a separate failure agent is trained to generate hard negative trajectories. These hard negatives, derived from plausible but incorrect attempts, help the target agent better distinguish successful behavior from subtle errors. Experiments on online shopping, scientific reasoning, and interactive SQL querying show a 5.7% average reward improvement over baseline methods.
By Yeonsung Jung, Trilok Padhi, Sina Shaham, Dipika Khullar, Joonhyun Jeong, Ninareh Mehrabi, Eunho Yang
arXiv:2605.07111v3 Announce Type: replace-cross
Abstract: Recent literature on fine-tuning Large Language Models highlights a fundamental debate. While Full Fine-Tuning (FFT) provides greater represe...
By Haozhan Tang, Xiuqi Zhu, Xinyin Zhang, Boxun Li, Virginia Smith, Kevin Kuo
arXiv:2601.12310v2 Announce Type: replace
Abstract: Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic dr...
By Jennifer Dodgson, Alfath Daryl Alhajir, Michael Joedhitya, Akira Rafhael Janson Pattirane, Surender Suresh Kumar, Joseph Lim, C. H. Peh, Adith Ramdas, Steven Zhang Zhexu
arXiv:2609.39371v1 Announce Type: new
Abstract: In hospital workflows, electronic health records (EHRs) are often noisy, and may not contain the evidence needed to confirm events or measurements refe...
By Yitong Qiao, Yancheng Jin, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Zhixuan Chu
InsightMap is a framework that uses top‑down maps as explicit spatial memory and action‑conditioned prediction targets for language‑guided navigation. It links historical views to labeled map locations and employs a shared multimodal backbone to jointly learn navigation action prediction and post‑action map generation, providing auxiliary training supervision. The approach supports a unified RGB‑D pipeline for navigation, visual question answering, situated reasoning, and 3D grounding, achieving state‑of‑the‑art results on R2R‑CE, RxR‑CE, ScanQA, SQA3D, ScanRefer, and outperforming baselines on the Unitree Go2 platform.
By Hongpei Zheng, Hujun Yin
arXiv:2609.37705v1 Announce Type: cross
Abstract: Goodness-of-fit testing is a basic tool for assessing whether a fitted procedure has captured the systematic information contained in the covariates....
By Xingwei Liu, Yuhong Yang, Wangli Xu
arXiv:2604.22893v2 Announce Type: replace-cross
Abstract: Traditional ``row-count $\times$ quality coefficient'' approaches fail to capture the nonlinear utility of data for Large Language Model (LLM...
By Minghui Xu, Qi Luo, Kun Li, Zhengyang Shan
IronLLM-0.6B is a 654‑million‑parameter language model engineered for efficient on‑device inference, featuring a hybrid attention architecture, X‑MTP multi‑token prediction, and a lightweight verification head that yields a 1.48× decoding speedup. Trained on roughly 6.2 trillion tokens with a quality‑oriented pipeline and further refined via Multi‑Domain On‑Policy Distillation, the model adopts an Instruct‑Only design to meet low‑latency requirements. A lighter variant, IronLLM‑0.6B‑Light, replaces RMSNorm with Dynamic Tanh and streamlines costly components to enhance inference and quantization efficiency, offering a strong performance‑efficiency trade‑off for resource‑constrained deployment.
By Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao
arXiv:2609.35917v1 Announce Type: new
Abstract: Graph neural networks are increasingly applied to road-level crash prediction, but the stability of their reported gains has received little scrutiny....
By Maurya Patel
arXiv:2512.00329v2 Announce Type: replace-cross
Abstract: Temporal reasoning over evolving semi-structured tables poses a challenge to current QA systems. We propose an approach that recasts the task...
By Ashish Thanga, Vibhu Dixit, Abhilash Shankarampeta, Vivek Gupta