arXiv AI

Task-Conditioned Synthetic Data Generation for Improving Machine Learning Performance in Agricultural Prediction Tasks

arXiv:2607. 09751v1 Announce Type: new Abstract: Machine Learning (ML) algorithms have been widely used to estimate agricultural variables across diverse contexts.

arXiv Machine Learning
2d ago

Synthesis Without Training: An Inference-Only Pipeline for Tabular, Temporal, and Relational Synthetic Data

GENSCRIPT is an inference‑only pipeline that generates synthetic data without training a generative model. It creates a deterministic statistical profile of the source data, uses a language model to infer field semantics and cross‑column constraints, and then compiles these into an executable sampler that works for single‑table, temporal, and relational data. The method builds generators in minutes, samples large datasets quickly, and achieves fidelity comparable to leading methods while preserving key data relationships such as 1‑to‑1 mappings and primary‑foreign key constraints.

By Zilong Zhao, Abdul Raheem, Jiayu Li, Sohei Arisaka, Darius Lim Hong Yi, Milad Abdollahzadeh, Uzair Javaid, Biplab Sikdar
arXiv Machine Learning
2d ago

An Input-Frugal Deep Learning Framework for Weather-Driven National Crop-Yield Forecasting: A Case Study of Brazilian Soybean

The paper introduces a lightweight deep learning framework that forecasts Brazilian soybean yields using only routine weather data and two simple static inputs (crop year and agro-environmental label). Across 20 seasons, transformer-based models achieved the highest accuracy, outperforming traditional ridge regression and a moving‑average baseline by nearly 48%. Ablation studies show that the static inputs and spatial expansion improve performance without adding complexity, and SHAP analysis highlights the importance of crop year and weather variables in driving yield variations.

By Fernando Dupin da Cunha Mello (Stricto Sensu Department, SENAI CIMATEC University, Salvador, Bahia, Brazil), Prashant Kumar (Global Centre for Clean Air Research), Erick G. Sperandio Nascimento (Stricto Sensu Department, SENAI CIMATEC University, Salvador, Bahia, Brazil)
arXiv Machine Learning
Aug 19

Evaluating and improving crop-yield forecasting methods during extreme drought

The study evaluates crop‑yield forecasting methods for the 2012 Midwestern US drought, comparing non‑deep learning machine learning models with a deep learning model (VITA) using 16 meteorological predictors. It highlights challenges such as distributional dissimilarity between training and test data, spatial and temporal sparsity, and demonstrates that sample weighting and feature selection improve non‑deep learning models but not VITA. The work contrasts deep versus non‑deep learning approaches and shows how modifications can mitigate issues arising from extreme drought conditions.

By Shrey Gupta, Yi Ming, George Mohler
arXiv Machine Learning
Aug 24

On the Transferability of Agricultural Weed Detection Under Cross-Field Distribution Shift

arXiv:2608.21254v1 Announce Type: cross Abstract: Accurate agricultural weed detection in real-world field conditions is essential for precision agriculture, enabling targeted intervention and reduci...

By Nikhilesh Prabhakar, Pranuthi Tenali, Wilfredo Abudeye Fernandez, Shekhar Borah, Athresh Karanam, Erik Blasch, Prabha Sundaravadivel, Sriraam Natarajan