What Images Cannot Say: Language-Guided Olfactory Representation Learning
arXiv:2607. 06402v1 Announce Type: cross Abstract: Images tell us what a scene looks like, but rarely what it would feel like to be there.
arXiv:2511. 20544v2 Announce Type: replace-cross Abstract: While olfaction is central to how animals perceive the world, this rich chemical sensory modality remains largely inaccessible to machines.
arXiv:2607. 06402v1 Announce Type: cross Abstract: Images tell us what a scene looks like, but rarely what it would feel like to be there.
arXiv:2605. 07821v2 Announce Type: replace-cross Abstract: Out-of-distribution (OOD) detection is crucial for ensuring the reliability of deep learning models.
arXiv:2606. 28399v1 Announce Type: cross Abstract: The structure of human visual representations underpins our capacity for adaptive behaviour.
arXiv:2512. 15748v2 Announce Type: replace Abstract: Visual Species Recognition (VSR) is a fundamental task in scientific disciplines that require species-level identification, including ecology, palynology, evolutionary biology, systematics, and phylogenetics.
arXiv:2607. 14721v1 Announce Type: cross Abstract: Cross-modal learning, i.
arXiv:2510. 13774v2 Announce Type: replace Abstract: Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data.
arXiv:2607. 16789v1 Announce Type: new Abstract: Real-world perception and decision making are inherently multimodal, integrating complementary signals across modalities.
arXiv:2606. 06458v1 Announce Type: new Abstract: Multiple Instance Learning (MIL) addresses problems where supervision is available at the level of bags of instances and has been successfully applied in fields ranging from computational pathology to satellite imagery.
arXiv:2606. 08204v1 Announce Type: new Abstract: Neural fields parameterize data as functions from coordinates to values, providing a unified framework for representation learning across modalities.
arXiv:2602. 24181v2 Announce Type: replace-cross Abstract: Pre-trained vision encoders like DINOv2 have demonstrated exceptional performance on unimodal tasks.
arXiv:2605. 05627v2 Announce Type: replace-cross Abstract: Sustainable forest management relies on precise species composition mapping, yet traditional ground surveys are labour-intensive and geographically constrained.
arXiv:2606. 27886v1 Announce Type: new Abstract: Recent advances in Human Activity Recognition (HAR) from wearable sensors have shown that multi-modal deep learning models consistently outperform their uni-modal counterparts.