arXiv:2608.00152v3 Announce Type: replace-cross
Abstract: AI evaluation can support the wrong inference when an in-domain benchmark success does not survive distribution shift, or when the benchmark...
By Mehrdad Shoeibi, Niloofar Yousefi
arXiv:2608. 00152v1 Announce Type: new Abstract: Predicting the magnitude of a CRISPRi perturbation's transcriptomic effect on held-out target genes is an important open problem in single-cell biology.
By Mehrdad Shoeibi, Niloofar Yousefi
SCALE is a conditional transport model that treats cells as unordered sets to predict treated cell populations without requiring cell-level matching. It uses a shared set-aware encoder and a conditional DiT backbone to learn latent transport, enabling endpoint supervision that is directly delta-aligned. Across diverse perturbation types—including genetic, chemical, developmental, and immune—SCALE accurately recovers gene‑expression changes, response directions, and population structure, outperforming competing methods on CRISPR data and successfully prioritizing cytokines that elicit distinct immune responses.
By Shuizhou Chen, Lang Yu, Xueqin Lin, Xinjie Mao, Songming Zhang, Xinyu Gu, Hao Wu, Sheng Xu, Kedu Jin, Lei Bai, Quan Qian, Qin Chen, Qiang Gao, Siqi Sun, Zhangyang Gao
The paper introduces CELLAUDIT, a method for auditing whether inputs claimed to influence predictive models actually do so. By testing if an input can enter the computation, whether predictions depend on it, and if that dependence improves observed responses, the authors evaluate agent-generated predictors on a morphology‑transcriptomics benchmark (BBBC047). Their findings show that many models claim compound contributions that are not supported by the data, and that falsification‑guided revisions can recover genuine input effects while improving performance.
By Mengran Li, Bo Li, Chengyang Zhang, Yang Yan, Jinfeng Xu, Zhenchao Tang
The paper introduces CP‑BG‑Bench, a paired‑view evaluation framework for Cell Painting vision encoders that fixes a central cell across four matched views (raw crop, segmented, and density‑augmented variants). Using this framework on three datasets and three encoders, the authors show that standard single‑metric rankings (e.g., replicate mAP) vary systematically across protocols, revealing disagreements along axes of cell versus background, morphology versus context, and within‑study versus across‑batch performance. The study demonstrates that segmented views can outperform crops in certain tasks and that background‑driven gains are largely determined by experimental design rather than encoder choice.
By Tim Treis, Nikita Moshkov, Johan Fredin Haslum, Shantanu Singh, Fabian J. Theis
The paper introduces AssayBench-Loop, a large benchmark of 1,389 CRISPR screens across five phenotype categories, and builds on it to develop AssayLoop, a sequential experimental design framework that combines a transformer-based acquisition policy (AssayFormer) trained on historical data with LLM-derived biological priors. AssayLoop achieves a 5.67‑fold enrichment over random selection, recovering 27.7% of hits after testing only about 5% of the library, and outperforms existing adaptive-design methods and standalone LLMs. The authors also present AssayLLM, extending the approach directly to an LLM via task‑specific post‑training, and show that performance improves with more historical training data and transfers to unseen phenotype categories.
By Carl Edwards, Edward De Brouwer, Xiner Li, Namkyeong Lee, Ehsan Hajiramezanali, Anne Biton, Sara Mostafavi, Gabriele Scalia
arXiv:2607. 17671v1 Announce Type: new Abstract: Large-scale single-cell perturbation atlases make it possible to ask an inverse question: given an observed transcriptional response, which annotated targets and compounds in a fixed library are most consistent with that response?
By Kseniia Vaniushkina, Jeongmin Lim, Jinyong Park
Compositional analysis of frozen vision encoders should determine both what changed and where it changed. Standard factor probes score these axes separately, however, and can reward multiple operations that reuse the same predicted slot.
The study investigates whether independently trained language‑model societies share a common packet language and how inherited interface states affect learning. A comprehensive audit of 30 pairwise interactions among six restricted societies shows that only one pair is fully interoperable, another is partially compatible, and the remaining 26 pairs fail across all alignment levels. Further experiments reveal that a globally trained communication interface can act as a severe negative‑transfer prior, but inherited interfaces never outperform fresh‑interface controls by the preregistered margin.
By Narcis Marincat
arXiv:2608. 10145v1 Announce Type: new Abstract: LeWorldModel trains a latent world model with a prediction loss and a single anti-collapse regulariser, and reports approximately 87% of goals reached on TwoRoom, its simplest diagnostic environment.
By Joyjeet Singh
arXiv:2607. 18553v1 Announce Type: cross Abstract: Can a language model read the quality of ongoing computation, and can an external intervention turn that readout into better outcomes?
By Jan Kirin
arXiv:2607. 27651v1 Announce Type: new Abstract: Adaptive laboratories choose measurements during experiments, yet most methods begin after adaptation is permitted.
By Jia Bi, Samuel Pinilla, Chenyang Zhu