arXiv Computer Vision

HiFiC-G: Adapting HiFiC for Hi-C Contact Matrices

The paper investigates adapting the HiFiC generative image compression model, originally designed for natural photographs, to compress Hi‑C chromatin contact maps while preserving biologically relevant features. By replacing HiFiC’s distortion term with a spatially‑weighted MSE that emphasizes loops, TAD boundaries, stripes, and compartments, and adding an insulation‑score loss, the authors fine‑tune a pretrained HiFiC checkpoint in a three‑phase strategy to create HiFiC‑G. Evaluation on two cell lines shows HiFiC‑G better retains local structures such as stripes and TAD boundaries, though long‑range A/B compartment preservation remains limited due to architectural constraints.

arXiv Machine Learning
Sep 7

When Genomic Masking Priors Fail to Transfer: Strong Variant Prediction, Weak Functional Generation

The paper introduces GenDA, a bidirectional discrete diffusion model designed for genomic sequence reconstruction, hypothesizing that entropy-guided span placement would improve variant-effect prediction and functional sequence generation. While the 202‑million‑parameter GenDA model achieves a higher ClinVar SNV AUROC (0.774) than a comparable autoregressive model, the improvement is not attributable to entropy guidance, and the model fails to outperform a shuffled‑gap baseline in zero‑shot functional inpainting across various genomic regions. The authors identify limitations such as tokenization granularity, span length caps, and the mismatch between local sequence complexity and functional importance, concluding that variant prediction, corruption priors, and functional generation are distinct tasks requiring separate validation.

By Susu Hu, Preetam Gattogi, Jens Lehmann, Sahar Vahdati, Stefanie Speidel, Julien Vibert
arXiv Machine Learning
Sep 14

Same Encoder, Different Winner: A Paired-View Framework for Cell Painting Encoder Evaluation

The paper introduces CP‑BG‑Bench, a paired‑view evaluation framework for Cell Painting vision encoders that fixes a central cell across four matched views (raw crop, segmented, and density‑augmented variants). Using this framework on three datasets and three encoders, the authors show that standard single‑metric rankings (e.g., replicate mAP) vary systematically across protocols, revealing disagreements along axes of cell versus background, morphology versus context, and within‑study versus across‑batch performance. The study demonstrates that segmented views can outperform crops in certain tasks and that background‑driven gains are largely determined by experimental design rather than encoder choice.

By Tim Treis, Nikita Moshkov, Johan Fredin Haslum, Shantanu Singh, Fabian J. Theis