arXiv:2607. 09084v1 Announce Type: new Abstract: The rapid expansion of large-scale AI models has led to significant performance breakthroughs across diverse domains, yet it has also raised critical concerns regarding computational costs, energy consumption, and environmental sustainability.
By Linhui Xiao, Guiping Cao, Mingyue Guo, Xianchao Guan, Fan Yang, Ming Tao, Xin Li, Yuxin Peng, Yaowei Wang
arXiv:2603. 14147v2 Announce Type: replace Abstract: The generative artificial intelligence (AI) ecosystem is undergoing rapid transformations that threaten its sustainability.
By Margarita Belova, Yuval Kansal, Yihao Liang, Jiaxin Xiao, Niraj K. Jha
The survey "From Models to Systems: A Comprehensive Survey of Efficient Multimodal Learning" reviews over 300 works on efficient multimodal learning (EML), proposing a structured taxonomy that spans model, algorithm, and system layers. It synthesizes how cross‑layer co‑design addresses the Efficiency‑Utility‑Privacy trade‑off and illustrates this through a case study of multimodal large language models. The paper also offers optimization blueprints for various domains, discusses a shift toward self‑regulating intelligence, and outlines open challenges for future EML research.
By Pan Wang, Siwei Song, Hui Ji, Siqi Cao, Heng Yu, Zhijian Liu, Huanrui Yang, Yingyan Celine Lin, Beidi Chen, Mohit Bansal, Xiaoming Liu, Pengfei Zhou, Ming-Hsuan Yang, Tianlong Chen, Jingtong Hu
The paper introduces agentic-eCAL, an extension of the Energy Cost of AI Lifecycle metric to evaluate multi‑agent AI workflows across the edge‑cloud continuum. By combining a two‑rate energy model with OSI‑layer transport analysis, the authors quantify that inter‑agent text transfer accounts for only 0.25% of total workflow energy, highlighting that the main energy cost lies in additional inference and context processing triggered by communication. The study uses extensive GPU benchmarks on NVIDIA A100/H100 with 16 open‑weight models and 8 orchestration topologies to validate the metric and explore placement implications.
By Carolina Fortuna, Vid Han\v{z}el, Tim Strnad, Bla\v{z} Bertalani\v{c}
arXiv:2606. 12688v1 Announce Type: cross Abstract: We are entering a new era of composite model architectures that integrate diverse components such as vision encoders, language backbones, diffusion and flow heads, audio codecs, action generators, and world-model predictors.
By Atindra Jha, Naomi Sagan, Keisuke Kamahori, Irmak Sivgin, Rohan Sanda, Steven Gao, Mark Horowitz, Luke Zettlemoyer, Olivia Hsu, Jure Leskovec, Baris Kasikci, Stephanie Wang
The article discusses how vision generative AI models, while rapidly advancing, have largely been developed with a focus on output quality, leading to hardware that adapts reactively to increasing model demands. It evaluates the parameter cost and energy efficiency of these models across various accelerator platforms and aligns four generative model families with seven real-world application domains. The authors propose a software‑hardware co‑design strategy that considers deployment constraints from the outset, ensuring that the appropriate model runs on suitable hardware for specific applications, thereby making generative AI deployment more sustainable and widely accessible.
By Eleni Tselepi, Cristian Sestito, Shady Agwa, Themis Prodromakis
arXiv:2609.24526v1 Announce Type: new
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints and...
By Foundation Model, Li Auto Inc
arXiv:2608. 06723v1 Announce Type: cross Abstract: The rapid scaling of Large Language Models (LLMs) has significantly increased computational cost, energy consumption, and inference latency, making accurate estimation essential for sustainable artificial intelligence deployment and hardware-aware design.
By Saeid Shokoufa, Mohammad Erfan Sadeghi, Mehdi Kamal, Massoud Pedram
arXiv:2609.24526v2 Announce Type: replace
Abstract: Physical AI requires models to ground visual and linguistic understanding in real-world environments while accounting for environmental constraints...
By Foundation Model, Li Auto Inc
arXiv:2607. 20468v1 Announce Type: new Abstract: AI agents are increasingly used to automate research and development tasks, yet existing benchmarks typically evaluate them on prescribed workflows or narrow action spaces.
By Jehyeok Yeon, Ben Rank, Maksym Andriushchenko
arXiv:2601. 21570v2 Announce Type: replace Abstract: The field of Embodied AI is witnessing a rapid evolution toward general-purpose robotic systems, fueled by high-fidelity simulation and large-scale data collection.
By Zixing Lei, Genjia Liu, Yuanshuo Zhang, Qipeng Liu, Yuzhu Cai, Sixiang Chen, Jixian Wu, Yunhong Wang, Weixin Li, Chuan Wen, Bo Zhao, Shanghang Zhang, Wenzhao Lian, Siheng Chen
MineExplorer is a benchmark designed to assess the open‑world exploration abilities of multimodal large language models (MLLMs) in Minecraft. It filters out tasks that rely heavily on Minecraft‑specific knowledge, organizes tasks into ReAct‑style capabilities, and composes atomic tasks into implicit multi‑hop challenges. A multi‑agent synthesis workflow creates reliable task graphs, sandbox scenes, and rule‑based milestone evaluators, and human evaluation confirms its superiority over a single‑agent baseline. Experiments show that while advanced MLLMs can handle many single‑hop tasks, they struggle with longer trajectories that require coordinating hidden prerequisites, and larger models or different thinking modes do not consistently improve performance.
By Tianjie Ju, Yueqing Sun, Zheng Wu, Wei Zhang, Yaqi Huo, Xi Su, Qi Gu, Xunliang Cai, Gongshen Liu, Zhuosheng Zhang