The paper introduces Thomson, a frontier AI model developed through continual learning on open-weight models, aiming to democratize access to high-performance AI. It argues that institutions with limited resources can achieve frontier-level performance by applying a modern mid- & post-training stack, preserving model plasticity and stability while minimizing high-impact interventions. Thomson demonstrates competitive performance across agentic tasks, safety, legal, tax, multilingualism, and large-scale deep research, exhibiting a distinctive π-shaped improvement pattern and effectively mitigating the forgetting problem seen in narrow domain adaptation.
By Shengzhuang Chen, Jerrod Parker, Yejin Bang, Andrew M. Bean, Nabeel Seedat, Stefan Winzeck, Daniil Glazko, Jannik Zgraggen, Fangyi Yu, Scott Arnott, Dietrich Trautmann, Luca Ciuffreda, Guglielmo Bonifazi, Davide Romano, Bradley Bell, Kirsty Fielding, Daniele Giofr\`e, Tom Zielund, Ipshita Chatterjee, Sneha Murthy Ghantasala, Manpreet Nanreh, John Scoville, Maciej Sakowicz, Wassim Seifeddine, Lukas Thede, Jonathan Richard Schwarz
arXiv:2608. 09548v1 Announce Type: cross Abstract: Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators.
By Yilin Jiang, Xiaorong Zhu, Fei Tan, Zicheng Zhang, Kaiyi Huang, Yang Yu, Zexuan Fei, Yiming Luo, Keqian Li, Hao Hao, Aimin Zhou, Guangtao Zhai
The paper examines how large language models (LLMs) can be biased by irrelevant social contexts when evaluating teachers, using a large U.S. classroom transcript dataset. It shows that spurious contexts can shift model ratings by up to 1.48 points on a 7‑point scale and that standard mitigation methods like SFT and DPO are insufficient. The authors introduce Debiasing‑DPO, a method that combines contrastive reasoning‑augmented DPO with SFT, which reduces bias by 84% and improves predictive accuracy by 52% on Llama and Qwen Instruct models.
By Hyunji Nam, Dorottya Demszky
The paper demonstrates that large language model (LLM) evaluators, whether reward‑model based or prompted LLM‑as‑a‑Judge, exhibit significant language bias in multilingual settings. Experiments with semantically identical instruction‑response pairs across 23 languages reveal that lower‑resource languages receive higher scores, a bias that persists across eight open‑weight evaluators and is not detectable by standard pairwise accuracy metrics. The authors link the bias to model uncertainty and language identity, showing it cannot be explained by content difficulty alone.
By Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen
The paper reports the first systematic audit of open‑weight large language models (LLMs) in hiring contexts, examining how job‑posting language influences recruiter and job‑seeker simulations across six models. It finds that agentic language lowers recruiter scores for female candidates while communal language mitigates this effect, and that coded‑exclusion language sharply reduces recruiter scores for non‑White candidates and discourages non‑White personas from applying. The study also identifies the explicit demographic label as the main causal factor and proposes a concrete pre‑deployment audit protocol aligned with EU and U.S. regulatory requirements.
By Kosuke Kitahara, Nobuhiro Yamaguchi
arXiv:2609.22169v1 Announce Type: new
Abstract: Employers are increasingly using large language models (LLMs) to automate their hiring process. This paper investigates the risk of monocultural biases...
By Matthew Bone, Fabian Stephany, Maria del Rio-Chanona
The study examines whether large language models (LLMs) discriminate against candidates based on institutional prestige and geographic location. Three factorial experiments (4,320 API calls across four LLMs and five domains) reveal a significant institution‑tier gradient (+0.297 points on a 10‑point scale), a stronger prestige effect than country‑of‑origin effect, and a dominant journal prestige effect (Nature vs. a peripheral journal). The authors also identify a ‘rescue effect’ where publishing in Nature mitigates low institutional prestige, and quantify bias using the Neutrosophic Bias Index, highlighting evaluation inconsistency for low‑prestige profiles.
By Maikel Leyva-Vazquez, Florentin Smarandache
Warning: This paper studies stereotypes and biases, and contains potentially disturbing examples, used for illustration purposes only. Our findings should not be interpreted as an argument against alignment.
The study audits demographic bias across four deep knowledge tracing architectures—DKT, DKVMN, SAKT, and AKT—using two large public datasets (Eedi and OULAD). It finds that bias is context‑dependent: socioeconomic bias is significant on Eedi, while gender bias appears on OULAD for most models. The most accurate model, AKT, also exhibits the greatest bias, and standard mitigation techniques such as reweighting and adversarial debiasing fail to reduce bias without sacrificing accuracy.
By Dang Quang Minh, Nguyen Dung Son, Nguyen Huu Loi, Truong Viet Vu, Nguyen Thai Anh
arXiv:2607. 28934v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender.
By Martin Lukk (University of Toronto)
arXiv:2606. 26099v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly deployed in artificial intelligence (AI) governance analysis across national and international organisations.
By Jason Hung
arXiv:2604. 14514v2 Announce Type: replace Abstract: Healthcare disparities persist across socioeconomic boundaries, often attributed to unequal access to screening, diagnostics, and therapeutics.
By Michal Rosen-Zvi, Yoav Kan-Tor, Michael Danziger, Agata Ferretti, Javier Aula-Blasco, Julia Falcao, Ron Shamir, Mira Marcus-Kalish, Mordechai Muszkat