Data engineering

Pipelines, warehouses, feature stores and the query engines that feed everything above.

535 stories · RSS feed

arXiv AI
1d ago

Enterprise Representation Simplification (ERS): Reducing Representational Complexity for Enterprise AI

Enterprise Representation Simplification (ERS) proposes reducing unnecessary representational complexity in enterprise information while preserving essential data within a defined scope. The paper introduces Enterprise Representation Complexity (ERC), a representation‑neutral model that measures complexity across four dimensions—Objects, Interactions, Behaviors, and Supporting Sources—at both representation and task levels. ERC enables comparison of architectural simplification versus retrieval optimization, supports an economic model of maintenance costs, and demonstrates that lower task‑level ERC can improve AI reasoning accuracy, as shown in Text‑to‑SQL research.

By Terry Dorsey, Kevin Huggins
arXiv AI
1d ago

BudgetSchemaBench: A Budget-Swept Diagnostic for Schema Context in Text-to-SQL

BudgetSchemaBench is a diagnostic tool for evaluating how different schema‑context budgets affect text‑to‑SQL systems. It automatically derives relevance labels from gold SQL, tests four budgets across 80 databases, and compares three schema representations while keeping table rankings fixed. The study shows that increasing the budget improves execution accuracy, especially for lexical retrieval, and that dense retrieval already captures most needed tables at low budgets.

By Chen Shen
arXiv Computer Vision
1d ago

MMVistaReason: Toward Open-Data and Post-Training Recipes for Multimodal Reasoning

arXiv:2610.01352v1 Announce Type: new Abstract: Open multimodal reasoning models have benefited from large-scale reasoning supervision, yet reliable post-training remains challenging due to uneven da...

By Juekai Lin, Honglin Lin, Yuqian Yuan, Xiaolong Wu, Jie Cao, Liang Liang, Yunqi Cao, Yun Zhu, Wenqiao Zhang, Lijun Wu
arXiv AI
2d ago

Co-Evolving Agents: Learning from Failures as Hard Negatives

The paper introduces a co‑evolving framework where a target agent improves by learning from its own failures, and a separate failure agent is trained to generate hard negative trajectories. These hard negatives, derived from plausible but incorrect attempts, help the target agent better distinguish successful behavior from subtle errors. Experiments on online shopping, scientific reasoning, and interactive SQL querying show a 5.7% average reward improvement over baseline methods.

By Yeonsung Jung, Trilok Padhi, Sina Shaham, Dipika Khullar, Joonhyun Jeong, Ninareh Mehrabi, Eunho Yang
arXiv AI
2d ago

Survival is the Only Reward: Sustainable Self-Training Through Environment-Mediated Selection

arXiv:2601.12310v2 Announce Type: replace Abstract: Self-training systems often degenerate due to the lack of an external criterion for judging data quality, leading to reward hacking and semantic dr...

By Jennifer Dodgson, Alfath Daryl Alhajir, Michael Joedhitya, Akira Rafhael Janson Pattirane, Surender Suresh Kumar, Joseph Lim, C. H. Peh, Adith Ramdas, Steven Zhang Zhexu
arXiv Computer Vision
3d ago

InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning

InsightMap is a framework that uses top‑down maps as explicit spatial memory and action‑conditioned prediction targets for language‑guided navigation. It links historical views to labeled map locations and employs a shared multimodal backbone to jointly learn navigation action prediction and post‑action map generation, providing auxiliary training supervision. The approach supports a unified RGB‑D pipeline for navigation, visual question answering, situated reasoning, and 3D grounding, achieving state‑of‑the‑art results on R2R‑CE, RxR‑CE, ScanQA, SQA3D, ScanRefer, and outperforming baselines on the Unitree Go2 platform.

By Hongpei Zheng, Hujun Yin
arXiv AI
3d ago

IronLLM: Forging Compact Edge-Native Language Models for Real-Time Embodied Intelligence

IronLLM-0.6B is a 654‑million‑parameter language model engineered for efficient on‑device inference, featuring a hybrid attention architecture, X‑MTP multi‑token prediction, and a lightweight verification head that yields a 1.48× decoding speedup. Trained on roughly 6.2 trillion tokens with a quality‑oriented pipeline and further refined via Multi‑Domain On‑Policy Distillation, the model adopts an Instruct‑Only design to meet low‑latency requirements. A lighter variant, IronLLM‑0.6B‑Light, replaces RMSNorm with Dynamic Tanh and streamlines costly components to enhance inference and quantization efficiency, offering a strong performance‑efficiency trade‑off for resource‑constrained deployment.

By Changdi Yang, Fengquan Jiao, Haochih Lin, Haoran Yang, Jing Xiao, Liangyu Huo, Suxin Lu, Tiance Chen, Wei Liu, Yinggan Xu, Yunxiang Lu, Zai Zheng, Zhirui Xie, Zhongyang Che, Ziyan Tang, Zuoxiang Zhao, Jian Yao