MOCLIP (Metasurface Optics Contrastive Learning Pretrained) is a nanophotonic foundation model that unifies metasurface geometry and spectra in a shared latent space using contrastive learning on an ImageNet‑1K–sized experimental dataset. It enables high‑throughput zero‑shot inverse design, predicting 0.2 million samples per second and allowing the design of a full 4‑inch wafer of high‑density metasurfaces in minutes. The model also supports generative latent‑space optimization with 97 % accuracy and demonstrates an optical information storage concept achieving 0.1 Gbit/mm², six times higher than commercial optical media.
By S. Rodionov, A. Burguete-Lopez, M. Makarenko, Q. Wang, F. Getman, A. Fratalocchi
arXiv:2608. 11860v1 Announce Type: cross Abstract: Data-driven inverse design enables efficient generation of nanophotonic structures with prescribed optical responses, but spectrum-to-geometry mapping remains challenging due to non-uniqueness and fine geometric features.
By Waleed Waseer, Muhammad Shahid Jabbar, Muhammad Sohail Ibrahim, Shujaat Khan
The paper proposes using ultrathin metalenses to physically encode metric depth cues into two polarized optical wavefronts, enabling accurate monocular depth estimation. By aligning a pretrained depth foundation model with these optical signals through direct fine‑tuning, and by creating a simulation pipeline to generate realistic metalens responses from RGB‑D data, the authors bridge the gap between nanophotonics and learned depth priors. Experiments show that this method surpasses conventional monocular metric depth estimation and depth‑from‑defocus baselines.
By Bingxuan Li, Jiahao Wu, Yuan Xu, Zezheng Zhu, Yunxiang Zhang, Kenneth Chen, Yanqi Liang, Nanfang Yu, Qi Sun
The paper presents a method for training single‑step neural surrogates that can handle wave‑scattering inverse problems with tens of thousands of controllable variables. By dynamically generating training examples through gradient ascent and using a replay dataset with normalization, the authors achieve a surrogate that accurately models two‑dimensional wave scattering for up to 41,772 variables and can generalize to over 3 million variables without retraining. The surrogate demonstrates comparable or better performance than traditional FDTD simulations for large‑scale forward simulations and inverse design of photonic devices, achieving speedups up to 26.5×.
The paper introduces a unified large language model workflow for modeling and inverse-design of metasurfaces across multiple families. By converting geometries, design parameters, and optical responses into a shared instruction‑following text format, the authors fine‑tune Gemma‑2‑9B on eight distinct metasurface families. Compared to single‑family models, the joint model predicts all families’ optical responses simultaneously and reduces mean‑squared error by an average of 56.5%.
"whyItMatters":"The approach demonstrates that a shared sequence‑based LLM interface can streamline cross‑family metasurface design, eliminating the need for separate surrogate architectures for each geometry class."
By Huanshu Zhang, Lei Kang, Yuyan Chen, Luxiang Wang, Zhaolong Cao, Douglas H. Werner
The paper presents a method for training single‑step neural surrogates that can handle wave‑scattering problems with tens of thousands of controllable variables. By dynamically generating training examples that highlight surrogate errors and using a replay dataset with normalization, the authors achieve a surrogate that accurately simulates two‑dimensional wave scattering for up to 41,772 variables and generalizes to over 3 million variables without retraining. The surrogate is applied to forward simulations and inverse design of freeform beam splitters and gradient‑index lenses, achieving speedups up to 26.5× compared to traditional FDTD methods.
By Charles Dove, Laura Waller