arXiv:2608. 14742v1 Announce Type: cross Abstract: Pandas has emerged as the de facto library for data processing and machine learning, widely used for tasks, such as data loading, transformation, and analysis.
By Syrym Abdikhan, Mazhar Hameed
NeMo Data Designer (NDD) is an open‑source framework for generating multimodal synthetic data. It uses a declarative configuration format that lets users define dataset columns—text, code, structured outputs, images, embeddings, and statistical samplers—to steer diversity. The system supports a preview‑and‑revision workflow, dependency resolution, and retry logic, and can be extended via plugins. Case studies demonstrate its use for structured, agentic, multimodal, and domain‑specialized tasks, including datasets for Nemotron model development and enterprise deployments.
By Johnny Greco, Nabin Mulepati, Andre Manoel, Eric Tramel, Kirit Thadaka, Mike Knepper, Dhruv Nathawani, Dane Corneil, Yev Meyer, Alex Watson, Maarten Van Segbroeck
arXiv:2608.30514v1 Announce Type: cross
Abstract: We present TSExplorer, a cross-platform tool for interactive annotation and exploration of time-series data. The tool enables users to inspect high-d...
By Einari Vaaras, Manu Airaksinen, Okko R\"as\"anen
arXiv:2602. 01588v3 Announce Type: replace-cross Abstract: Multimodal time series forecasting is crucial in real-world applications, where decisions depend on both numerical data and contextual signals.
By Huu Hiep Nguyen, Minh Hoang Nguyen, Dung Nguyen, Hung Le
arXiv:2603. 07053v3 Announce Type: replace Abstract: Scientists face significant visualization challenges as time-varying datasets grow in speed and volume, often requiring specialized infrastructure and expertise to handle massive datasets.
By Ishrat Jahan Eliza, Xuan Huang, Aashish Panta, Alper Sahistan, Zhimin Li, Amy A. Gooch, Valerio Pascucci
arXiv:2606. 14941v1 Announce Type: new Abstract: Time series forecasting models often benefit from historical patterns.
By Shiqiao Zhou, Zipeng Wu, Holger Sch\"oner, Edouard Fouch\'e, IAG Wilson, Shuo Wang
The article titled "A Practical Introduction to PySpark Window Functions" explains why the standard groupBy function isn’t enough for certain data processing tasks. It introduces PySpark window functions as a more powerful alternative, providing readers with a practical guide to implementing these functions in their data workflows.
By Thomas Reid
DEEPCHART is a new benchmark that evaluates large language models (LLMs) on faithful data‑science chart generation. It contains 1,482 expert‑annotated instances from scientific papers, financial filings, and ecosystem reports, and assesses chart creation through an Extract–Reason–Visualize pipeline. Experiments show that while LLMs can produce visually plausible charts, they frequently hallucinate data at the extraction and reasoning stages, especially in long, noisy, and multimodal contexts.
By Jiahui tang, Kuicai Dong, Dexun Li, Hongchao Gu, Haocheng Yu, Wei Han, Chen Zhang, Yong Liu, Hao Wang, Enhong Chen
arXiv:2607. 21117v1 Announce Type: cross Abstract: Preprocessing blood glucose time-series data is a critical yet often overlooked step in developing data-driven methods for diabetes management, particularly for type 1 diabetes.
By Davide Marelli, Giorgia Rigamonti, Mirko Paolo Barbato, Paolo Napoletano
Recent years have witnessed the emergence of multivariate modeling using time series foundation models (TSFMs), which achieve advanced zero-shot generalization. Modern multivariate TSFMs are predominantly pretrained on multivariate synthetic data, which is easier to scale but may fail to capture the complex temporal dynamics and cross-variable relationships present in real-world time series.
Take the next step to building real workflows with Spark on your laptop The post PySpark for Beginners: Beyond the Basics appeared first on Towards Data Science .
By Thomas Reid