arXiv AI

A Multimodal Dataset for Large Language Model Applications in the Energy Domain

arXiv:2607. 11459v1 Announce Type: cross Abstract: This paper presents the mAIEnergy dataset, an open-access, multimodal corpus developed to support Large Language Model (LLM) applications in the energy sector.

arXiv AI
Jul 14

WattCouncil: Context-Aware Household Energy Scenario Generation With Governed LLMs

arXiv:2607. 10720v1 Announce Type: new Abstract: The accelerating shift toward low-carbon power systems, together with the widespread adoption of behind-the-meter technologies such as rooftop solar and electric vehicles, is placing new operational and analytical demands on electricity grids.

By Mohannad Takrouri, Nicolas M. Cuadrado A., Martin Tak\'a\v{c}
arXiv Machine Learning
Aug 19

Open datasets and machine learning for two-phase heat transfer: a review following a spatial-temporal taxonomy

The review discusses how two‑phase heat transfer—critical for boiling, condensation, and thermal management—poses challenges for data reuse due to its complex interfacial physics. It surveys open datasets, machine‑learning techniques, and reusable software, organizing them with a spatial‑plus‑temporal dimensionality taxonomy (S+TD) that links data types to AI tasks such as regression, sequence learning, and image/video analysis. The paper proposes a roadmap for physics‑aware open data, including metadata standards, maturity labels, benchmark splits, and community databanks, emphasizing that progress in two‑phase AI relies as much on robust data infrastructure as on model design.

By Christy Dunlap, Ridwan Olabiyi, Firas Al-Hindawi, Hari Pandey, Stephen Pierson, Daniel Curl, Braden Stevens, Mohammad Ishraq Hossain, Annapurna Parjuli, Chinmaya Joshi, Ashif Iquebal, Han Hu
arXiv Machine Learning
Jul 31

Bridging AI and Energy Forecasting: An Autonomous Workflow with Customized Toolkit

arXiv:2307. 07191v3 Announce Type: replace Abstract: Energy forecasting is crucial for the power grid, but fundamentally different from general time series analysis: it highly relies on covariates like meteorological factors, and its goals must align with actual power grid operations, such as risk assessment and system reliability.

By Zhixian Wang, Leandro Von Krannichfeldt, Qingsong Wen, Chaoli Zhang, Liang Sun, Shirui Pan, Yi Wang
arXiv Machine Learning
Aug 19

TiMi: Empower Time Series Transformers with Multimodal Mixture of Experts

The paper introduces TiMi, a framework that enhances time series transformers with a Multimodal Mixture-of-Experts (MMoE) module to incorporate multimodal data, especially textual information, into forecasting. TiMi leverages large language models to generate future inferences that guide predictions, eliminating the need for explicit representation alignment. Experiments show TiMi achieves state‑of‑the‑art performance on sixteen real‑world multimodal forecasting benchmarks, outperforming advanced baselines while maintaining adaptability and interpretability.

By Jiafeng Lin, Yuxuan Wang, Huakun Luo, Jianmin Wang, Zhongyi Pei
arXiv AI
Jun 10

GCA Framework: A GCC Countries-Grounded Dataset and Agentic Pipeline for Climate Decision Support

arXiv:2604. 12306v3 Announce Type: replace-cross Abstract: Climate decision-making in the GCC states increasingly demands systems that can translate heterogeneous scientific and policy evidence into actionable guidance, yet general-purpose large language models (LLMs) remain weak both in region-specific climate knowledge and grounded interaction with geospatial and forecasting tools.

By Muhammad Umer Sheikh, Khawar Shehzad, Salman Khan, Fahad Shahbaz Khan, Muhammad Haris Khan
arXiv AI
Jun 11

Sustainability assessment using multimodal AI agents

arXiv:2507. 17012v2 Announce Type: replace Abstract: Reducing the rapidly growing environmental impact of the computing industry requires assessing the emissions of electronics at scale.

By Zhihan Zhang, Alexander Metzger, Yuxuan Mei, Felix H\"ahnlein, Zachary Englhardt, Tingyu Cheng, Gregory D. Abowd, Shwetak Patel, Adriana Schulz, Vikram Iyer
arXiv Machine Learning
Sep 18

The Environmental Impacts of Language Model Training Keep Rising Now is the Time to Catch Impacts on the Rebound

The paper analyzes the environmental footprint of machine learning model training, focusing on large language models and their hardware. It finds that energy use and environmental impacts have risen exponentially over the past decade, even when employing carbon‑efficient electricity and more efficient hardware. The study argues that optimization strategies alone cannot curb these impacts due to a rebound effect, and stresses the need to evaluate hardware life‑cycle impacts and integrate environmental metrics into NLP research practices.

By Cl\'ement Morand (STL), Anne-Laure Ligozat (ENSIIE, LISN, STL), Aur\'elie N\'ev\'eol (STL, LISN)
arXiv AI
Sep 16

little m: An AI Agent for Industrial Process Optimization

The paper introduces little m, an AI agent that helps formulate industrial process control models by combining a domain-specific knowledge repository with LLM-driven interaction. It tackles the challenge of converting messy real-world specifications, including natural language and spatial diagrams, into rigorous mathematical optimization models. The authors also present IPC-Bench, a multimodal dataset of 50 canonical scenarios, and show through automated and human evaluations that little m outperforms state‑of‑the‑art LLMs in generating semantically correct models.

By Yongchao Ye, Xinyu He, Dutliff Boshoff, Way Kuo, Lishuai Li