Posted by Dave Steiner, Clinical Research Scientist, Google Health, and Rory Pilgrim, Product Manager, Google Research There’s a worldwide shortage of access to medical imaging expert interpretation across specialties including radiology , dermatology and pathology . Machine learning (ML) technology can help ease this burden by powering tools that enable doctors to interpret these images more accurately and efficiently.
By Google AI
The paper examines how multi‑level fairness techniques—combining several bias‑mitigation steps—can reduce biases across patient demographics in health informatics. It reviews current literature, identifies gaps in implementation and reporting of health equity outcomes, and evaluates the role of reporting standards such as MINIMAR and TRIPOD in enhancing transparency. The authors conclude with recommendations to improve reporting transparency, broaden adoption of multi‑level fairness methods, and explicitly prioritize health equity in future research.
By Nick Souligne, Vignesh Subbian
Posted by Pooja Rao, Research Scientist, Google Research Health datasets play a crucial role in research and medical education, but it can be challenging to create a dataset that represents the real world. For example, dermatology conditions are diverse in their appearance and severity and manifest differently across skin tones.
By Google AI
The study examines why dermatology AI models, largely trained on light‑skinned, cancer‑focused images, perform poorly when applied to diverse patient populations. By comparing a cancer‑trained baseline, two dermatology foundation models, and a general‑purpose vision model on tone‑stratified and disease‑shifted datasets, the authors find that disease‑distribution shift, rather than skin‑tone underrepresentation, is the primary cause of generalization failure. Representation analysis shows that cancer‑specialized features lack transferable structure, while dermatology‑pretrained features maintain stronger clustering, and lightweight adaptation with about ten labeled examples per category can recover most performance.
By Nirajan Kunwor, Sanjaya Poudel, Quoc-Huy Trinh, Jahidul Arafat, Sunil Kumar Gaire
arXiv:2608.30568v1 Announce Type: cross
Abstract: Background: Population level assessments of predictive artificial intelligence (AI) can conceal performance disparities across subgroups. Fairness ev...
By Jo\~ao Matos, Ben Van Calster, Richard D. Riley, Paula Dhiman, Gary S. Collins
Posted by Rishabh Tiwari, Pre-doctoral Researcher, and Pradeep Shenoy, Research Scientist, Google Research Machine learning models in the real world are often trained on limited data that may contain unintended statistical biases . For example, in the CELEBA celebrity image dataset, a disproportionate number of female celebrities have blond hair, leading to classifiers incorrectly predicting “blond” as the hair color for most female faces — here, gender is a spurious feature for predicting hair color.
By Google AI
arXiv:2604. 23786v2 Announce Type: replace Abstract: In recent years, the integration of multimodal machine learning in wellbeing assessment has offered transformative potential for monitoring mental health.
By Sophie Chiang, Tom Brennan, Fethiye Irmak Dogan, Jiaee Cheong, Hatice Gunes
The paper investigates intersectional biases in multimodal clinical predictions using Electronic Healthcare Records (EHR). It introduces datasets MIMIC-Eye1 and MIMIC-IV ED, applies unified text representations from pre‑trained clinical language models, and benchmarks bias mitigation at the intersectional subgroup level. Results show that subgroup‑specific mitigation is robust across datasets, subgroups, and embeddings, effectively addressing intersectional biases in multimodal settings.
By Ayaazuddin Mohammad, Kishore Sampath, Resmi Ramachandranpillai
arXiv:2607. 08953v1 Announce Type: new Abstract: Algorithmic fairness methods are increasingly used to identify and mitigate bias in machine learning models, yet most approaches are evaluated in isolation and along single demographic axes.
By Nick Souligne, Isabella Mixton-Garcia, Vignesh Subbian
arXiv:2605. 02942v2 Announce Type: replace Abstract: Fairness studies of medical imaging AI often explain subgroup performance gaps through under-representation in the training data.
By Aya Elgebaly, Joris Fournel, Benjamin Laine J{\o}nch Jurgensen, Kamil Mikolaj, Anders Christensen, Martin Tolsgaard, Claes Ladefoged, Aasa Feragen
The paper introduces FRAME, a two‑step framework for auditing fairness claims in medical imaging. First, it derives a fair‑model reference distribution that captures the portion of subgroup performance differences attributable to sampling variation. Second, it tests the remaining difference using operators in representation space to assess whether demographic information or disease entanglement drives the bias. Across a large dataset of 702,206 images and 36 encoders, the reference explains a substantial median share of race and age differences, while interventions such as injecting demographic decodability or entangling disease direction have limited impact on the residual bias.
By Mahshad Lotfinia, Daniel Truhn, Andreas Maier, Soroosh Tayebi Arasteh
arXiv:2606. 04971v1 Announce Type: new Abstract: Machine learning engineering (MLE) agents promise to automate end-to-end ML pipeline development from raw data and natural language instructions, potentially making ML accessible to non-technical domain experts.
By Anna Richter, Julia Stoyanovich, Sebastian Schelter