arXiv Machine Learning

Tensor Methods: A Unified and Interpretable Approach for Material Design

arXiv:2602. 10392v2 Announce Type: replace Abstract: When designing new materials, it is often necessary to tailor the material design to have some desired properties.

arXiv AI
Sep 1

Tensor Methods for Language Models: From Token Representation to Training, Adaptation, Inference, Compression, and Interpretability

This survey reviews tensor methods applied to large language models, framing them through a seven‑stage lifecycle (tokenization, embeddings, pre‑training, adaptation, compression, inference, interpretability) and a component view (embeddings, attention, feed‑forward networks). It offers unified notation, theoretical foundations, and comparative analyses of tensorization strategies for Transformer components, while highlighting evaluation protocol differences and model scale effects. The paper also introduces a new metric, ρ_gap, to quantify the gap between theoretical memory savings and actual system‑level speedup, and connects tensor techniques to related efficiency and probabilistic methods.

By Matvei Tarasov, Salman Ahmadi-Asl, Andre L. F. de Almeida, Andrzej Cichocki
arXiv AI
Aug 26

Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

The article proposes treating large language model (LLM) data mixing as a classical mixture experiment, where data domains are components, token shares are proportions, and proxy-training runs serve as design points. Using sparse second‑order Scheffé response‑surface models, the authors construct model‑robust Σ‑optimal designs that efficiently identify optimal data mixtures and reveal strong interaction effects, especially between weak domains and web‑derived text. Empirical results on RegMix show that these designs recover mixture rankings while reducing proxy runs by about 25%, demonstrating that data mixing can be optimized through experimental design rather than solely prediction.

By Yicheng Mao, Hongru Du
arXiv Machine Learning
Sep 2

LLM-driven design of physics-constrained constitutive models: two agents are better than one

The paper presents a novel multi‑agent approach to generating physics‑constrained constitutive models using large language models (LLMs). A Creator agent proposes a model tailored to the data, while an Inspector agent audits each proposal against nine physical constraints and requests refinement if violations occur. Experiments with constitutive artificial neural networks on brain tissue, rubber, and porcine skin show that adding the Inspector increases the proportion of models passing all checks from 90 % to 95 % for Claude Opus 4.7 and from 47 % to 60 % for Kimi K2.5, while the resulting models match or exceed expert‑designed models in accuracy and generalization.

By Marius Tacke, Matthias Busch, Kian Abdolazizi, Jonas Eichinger, Kevin Linka, Roland Aydin, Christian Cyron
arXiv Machine Learning
Sep 10

Leveraging Discrete Function Decomposability for Scientific Design

The paper introduces Decomposition-Aware Distributional Optimization (DADO), a new algorithm that exploits decomposability in property predictors to improve in‑silico design of discrete objects such as proteins, circuits, and materials. DADO uses a soft‑factorized search distribution and graph message‑passing to coordinate optimization across linked factors defined by a junction tree over design variables. The method aims to make distributional optimization over combinatorial design spaces more efficient by leveraging the structure of the predictive model.

By James C. Bowden, Sergey Levine, Jennifer Listgarten
arXiv Machine Learning
Aug 20

A single design choice determines whether machine learning models of materials make physically impossible predictions

The article demonstrates that a single design decision—whether a machine‑learning model’s features include parity labels—determines if the model can ever predict physically impossible values for material properties. Using group‑theoretical analysis, the authors introduce the parity gap criterion to identify which properties and crystal symmetries are affected. Experiments on 2,000 centrosymmetric crystals show that parity‑labelled models achieve exact zero predictions for forbidden piezoelectric responses, whereas models lacking parity labels produce large errors, yet both maintain comparable overall accuracy.

By Can Polat, Mustafa Kurban, Erchin Serpedin, Hasan Kurban