Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

18,270 stories · RSS feed

arXiv Machine Learning
Jul 7

SMART: A Machine Learning and Monte Carlo Framework for Rapid Analysis of Stochastic Transistor Aging and Process Variation in Digital Circuits

arXiv:2607. 05187v1 Announce Type: new Abstract: As CMOS technology scales into the deep nanometer regime, digital circuit reliability is increasingly threatened by the combined stochastic effects of Bias Temperature Instability (BTI) and Process Variation (PV).

By Arash Esshaghi, Siavash Es'haghi, Gholamreza Shahabadi, Alireza Moradi
arXiv AI
Jul 7

ACPO: Adaptive Credit Policy Optimization via Fine-Grained Surrogate Entropy

arXiv:2607. 03126v1 Announce Type: cross Abstract: Reinforcement Learning (RL) has substantially improved the reasoning ability of large language models (LLMs), but sparse outcome rewards still make token-level credit assignment difficult.

By Zijun Xie, Yuyang You, Yongzhi Li, Enlei Gong, Zeyu Chen, Quan Chen, Yanhua Cheng, Peng Jiang, Yadong Mu
arXiv AI
Jul 7

Lights, Camera, Carbon: Architectural Scaling Laws for Video Generation Energy Consumption

arXiv:2607. 04553v1 Announce Type: cross Abstract: We present a bidirectional framework for estimating the energy consumption of text-to-video (T2V) and text-to-video-audio (T2VA) models from architectural first principles and observable generation parameters such as resolution and duration, requiring no access to weights, model size, or implementation details.

By Nidhal Jegham, Boris Gamazaychikov, Sasha Luccioni
arXiv AI
Jul 7

A Step Towards Robust Unsupervised Domain Adaptation via Fine-Tuning and Reinforcement Learning

arXiv:2607. 03600v1 Announce Type: cross Abstract: Adversarial robustness in Unsupervised Domain Adaptation (UDA) remains a significant challenge due to noisy pseudo labels and inherent distributional shifts between the clean source and adversarially perturbed target domains.

By Sushant Dagaji Desale, Rahul Mishra, Ashutosh Kumar Sinha
arXiv AI
Jul 7

Efficient Discovery of Conditional Dependencies with Desbordante

arXiv:2607. 04030v1 Announce Type: cross Abstract: Conditional functional dependencies (CFDs) are functional dependencies with a restricted scope: they specify the context in which a dependency holds and are useful for data-quality tasks, specifying complex integrity constraints, and extracting valuable insights from data.

By Ivan Kozhukov, Dmitry Fedoseev, Maksim Emelyanov, Artem Smola, Pyotr Senichenkov, Pavel Anosov, George Chernishev
arXiv AI
Jul 7

TacReasoner: A Dynamic Tactile-Language Framework for Interactive Reasoning in Real-World Scenarios

arXiv:2607. 05131v1 Announce Type: new Abstract: Among the five primary human senses, tactile is arguably the most fundamental to survival, as it enables the perception of physical contact and interaction in real-world environments.

By Kailin Lyu, Di Wu, Long Xiao, Jianning Zeng, Jianwei He, Chang Lin, Lianyu Hu, Lin Shu, Jie Hao, Ce Hao