arXiv Computer Vision

Collaborative On-Sensor Array Cameras

The paper presents a collaborative on‑sensor array camera that uses a distributed meta‑optics learning method to jointly optimize a 100‑million‑nanopost metasurface array for broadband visible imaging. By training the array end‑to‑end with a learned meta‑atom proxy and a parallax‑aware, noise‑aware reconstruction algorithm, the design overcomes the wavelength‑dependent limitations of traditional metalenses. Experimental results show that the camera delivers consistent image quality across varying scene illumination spectra without relying on generative reconstruction.

arXiv AI
Aug 25

MOCLIP: A Foundation Model for Large-Scale Nanophotonic Inverse Design

MOCLIP (Metasurface Optics Contrastive Learning Pretrained) is a nanophotonic foundation model that unifies metasurface geometry and spectra in a shared latent space using contrastive learning on an ImageNet‑1K–sized experimental dataset. It enables high‑throughput zero‑shot inverse design, predicting 0.2 million samples per second and allowing the design of a full 4‑inch wafer of high‑density metasurfaces in minutes. The model also supports generative latent‑space optimization with 97 % accuracy and demonstrates an optical information storage concept achieving 0.1 Gbit/mm², six times higher than commercial optical media.

By S. Rodionov, A. Burguete-Lopez, M. Makarenko, Q. Wang, F. Getman, A. Fratalocchi
arXiv AI
Aug 13

Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra

arXiv:2608. 11860v1 Announce Type: cross Abstract: Data-driven inverse design enables efficient generation of nanophotonic structures with prescribed optical responses, but spectrum-to-geometry mapping remains challenging due to non-uniqueness and fine geometric features.

By Waleed Waseer, Muhammad Shahid Jabbar, Muhammad Sohail Ibrahim, Shujaat Khan
arXiv Computer Vision
2d ago

Physically Grounded Monocular Depth via Nanophotonic Wavefront Encoding

The paper proposes using ultrathin metalenses to physically encode metric depth cues into two polarized optical wavefronts, enabling accurate monocular depth estimation. By aligning a pretrained depth foundation model with these optical signals through direct fine‑tuning, and by creating a simulation pipeline to generate realistic metalens responses from RGB‑D data, the authors bridge the gap between nanophotonics and learned depth priors. Experiments show that this method surpasses conventional monocular metric depth estimation and depth‑from‑defocus baselines.

By Bingxuan Li, Jiahao Wu, Yuan Xu, Zezheng Zhu, Yunxiang Zhang, Kenneth Chen, Yanqi Liang, Nanfang Yu, Qi Sun
Hugging Face Trending Papers
Aug 18

Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems

The paper presents a method for training single‑step neural surrogates that can handle wave‑scattering inverse problems with tens of thousands of controllable variables. By dynamically generating training examples through gradient ascent and using a replay dataset with normalization, the authors achieve a surrogate that accurately models two‑dimensional wave scattering for up to 41,772 variables and can generalize to over 3 million variables without retraining. The surrogate demonstrates comparable or better performance than traditional FDTD simulations for large‑scale forward simulations and inverse design of photonic devices, achieving speedups up to 26.5×.

arXiv Machine Learning
Aug 28

Towards a universal meta-optics solver via large language models

The paper introduces a unified large language model workflow for modeling and inverse-design of metasurfaces across multiple families. By converting geometries, design parameters, and optical responses into a shared instruction‑following text format, the authors fine‑tune Gemma‑2‑9B on eight distinct metasurface families. Compared to single‑family models, the joint model predicts all families’ optical responses simultaneously and reduces mean‑squared error by an average of 56.5%. "whyItMatters":"The approach demonstrates that a shared sequence‑based LLM interface can streamline cross‑family metasurface design, eliminating the need for separate surrogate architectures for each geometry class."

By Huanshu Zhang, Lei Kang, Yuyan Chen, Luxiang Wang, Zhaolong Cao, Douglas H. Werner
arXiv AI
Aug 19

Inductively Scalable, Single-Step Neural Surrogates for Wave-Scattering Inverse Problems

The paper presents a method for training single‑step neural surrogates that can handle wave‑scattering problems with tens of thousands of controllable variables. By dynamically generating training examples that highlight surrogate errors and using a replay dataset with normalization, the authors achieve a surrogate that accurately simulates two‑dimensional wave scattering for up to 41,772 variables and generalizes to over 3 million variables without retraining. The surrogate is applied to forward simulations and inverse design of freeform beam splitters and gradient‑index lenses, achieving speedups up to 26.5× compared to traditional FDTD methods.

By Charles Dove, Laura Waller
arXiv Machine Learning
Aug 20

A Comprehensive Review of Large Language Models for Nanophotonics: From Surrogate Modeling to Autonomous Design

The review examines how Large Language Models (LLMs) enhance nanophotonics design by providing semantic interfaces, code generation, and tool orchestration. It traces the evolution from classical neural networks to transformer-based models and categorizes LLM applications into surrogate models that map structure to spectrum and agentic systems that generate code and orchestrate simulations for closed-loop optimization. The article also highlights potential cross-disciplinary uses of LLMs in materials science and wireless communications, and envisions future multimodal foundation models that actively collaborate in autonomous scientific discovery.

By Huanshu Zhang, Kegeng Tang, Lei Kang, Sawyer D. Campbell, Zihao Wang, Douglas H. Werner
arXiv AI
Aug 10

Optimizing Spectral Prediction in MXene-Based Metasurfaces Through Multi-Channel Spectral Refinement and Savitzky-Golay Smoothing

arXiv:2602. 08406v2 Announce Type: replace-cross Abstract: The prediction of electromagnetic spectra for MXene-based solar absorbers, where MXenes are a family of two-dimensional transition metal carbides and nitrides, is a computationally intensive task traditionally addressed using full-wave solvers.

By Shujaat Khan, Waleed Iqbal Waseer, Muhammad Shahid Jabbar
arXiv Machine Learning
Sep 14

A unified self-supervised framework for single-frame Fresnel CDI and overlapped ptychography

The paper introduces a self‑supervised neural network that unifies single‑frame Fresnel coherent diffraction imaging (CDI) and overlapped ptychography. By using a fixed, pre‑estimated probe and optimizing with a Poisson negative log‑likelihood objective, the method reconstructs object patches from either a single diffraction frame or multiple overlapping measurements, achieving high SSIM scores and a ten‑fold improvement in photon‑dose efficiency. Demonstrations on synthetic patterns and real datasets from APS and LCLS show robust, high‑throughput reconstructions, with a 36× speedup over iterative solvers for a 10,304‑frame workload.

By Oliver Hoidn, Steven Henke, Albert Vong, Aashwin Mishra, Apurva Mehta, Matthew Seaberg
arXiv Machine Learning
Jun 3

Towards Blind Lens Aberration Correction via Large LensLib Pre-training and Discrete Degradation Priors

arXiv:2511. 17126v4 Announce Type: replace-cross Abstract: Emerging deep-learning-based lens library pre-training (LensLib-PT) pipeline offers a new avenue for blind lens aberration correction by training a universal neural network, demonstrating strong capability in handling diverse unknown optical degradations.

By Xiaolong Qian, Qi Jiang, Yao Gao, Lei Sun, Kailun Yang, Xian Wang, Zhonghua Yi, Wenyong Li, Ming-Hsuan Yang, Luc Van Gool, Kaiwei Wang
arXiv Computer Vision
Sep 3

SlowFast-SCI: Slow-Fast Deep Unfolding Learning for Spectral Compressive Imaging

SlowFast‑SCI introduces a dual‑speed deep‑unfolding framework for spectral compressive imaging that combines a slow, pre‑trained backbone with a fast, test‑time adaptation stage. The slow phase distills a priors‑based model into a compact fast‑unfolding network, while the fast phase embeds lightweight modules that self‑supervise at test time without retraining the backbone. This design yields significant reductions in parameters and FLOPs, improves out‑of‑distribution PSNR by up to 5.79 dB, and accelerates adaptation four‑fold, all while remaining modular enough to integrate with any existing deep‑unfolding system.

By Haijin Zeng, Xuan Lu, Jiezhang Cao, Kai Zhang, Yurong Zhang, Qiangqiang Shen, Guoqing Chao, Li Jiang, Yongyong Chen, Jingyong Su, Jie Liu