arXiv:2606. 09672v1 Announce Type: new Abstract: Ask a pretrained biomedical language model whether "cortisol 28 ug/dL" and "stock-market volatility" are related, and it returns a cosine similarity of 0.
By Suraj Biswas, Saurabh Gupta, Pritam Mukherjee
The paper evaluates nine on‑device named‑entity recognition models ranging from classical taggers to large language models, measuring not only accuracy but also latency and output validity. Using a silver‑gold benchmark derived from an LLM judge panel and a human‑validated corpus, the study shows that encoder‑based models achieve comparable accuracy to a 4 B instruct LLM while being much smaller, faster, and producing no malformed output. Confidence calibration of GLiNER is analyzed, revealing over‑confidence but improved reliability after temperature scaling and thresholding.
By Vinay Kumar Chaganti
The study investigates whether particular attention heads and individual neurons within those heads in language models are responsible for detecting network infrastructure information—specifically hostnames paired with IP addresses. Using causal ablation and selective testing across five models from three architecture families, the authors find that a small subset of heads reliably identifies such information with near-perfect accuracy. However, the extent to which this responsibility is concentrated in a single neuron varies by model; in some cases a single neuron suffices, while in others the signal is distributed across the head. The findings generalize to an independent reverse‑DNS dataset, though single‑neuron detectors are less robust.
By Abdul Kadir (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence), Md Mohasin Hossain (German Research Center for Artificial Intelligence, Saarland University, Saarbrucken, Germany), Daniel Sonntag (University of Oldenburg, Oldenburg, Germany, German Research Center for Artificial Intelligence)
arXiv:2608. 12365v1 Announce Type: cross Abstract: For fifty years, data systems have answered two questions.
By Ganesh S
The paper reports the ABAI submission to COLIEE 2026 Task 1, a case law retrieval challenge that suppresses cited passages, and details a four‑stage retrieval pipeline: multi‑view BM25 with reciprocal rank fusion, neural reranking, graph‑based features via a graph attention network, and a LightGBM meta‑learner over 34 features. The best run achieved an F1 score of 0.177 on the official test set, compared to a cross‑validated 0.311, and the authors attribute the gap to a recall ceiling, temporal distribution shift, and threshold miscalibration. A controlled post‑hoc study examined the impact of threshold transfer, decision quality across time, and query similarity, and identified specific remedies—such as BM25 length‑normalisation tuning, event‑triple views, and dense fusion—that improved recall, while other interventions had no effect.
By Minhan Cho, Soyoung Park, Daejin Choi, Jinyoung Han
arXiv:2608. 02877v1 Announce Type: new Abstract: Causal discovery recovers directed structure from observational data and is increasingly used in clinical settings to support mechanism reasoning and fairness audits of predictive models.
By Nitish Nagesh, Elahe Khatibi, Thomas Dean Hughes, Mahdi Bagheri, Pratik Gajane, Amir M. Rahmani
arXiv:2606.22419v3 Announce Type: replace
Abstract: A recent Nature Medicine study reports that general-purpose frontier LLMs outperform specialized retrieval-augmented clinical tools on medical benc...
By Madhulatha Mandarapu, Sandeep Kunkunuru
arXiv:2609.07595v2 Announce Type: replace-cross
Abstract: The same underlying computational problem is solved across unrelated fields under different names: recursive Bayesian state estimation appear...
By Eryk Kulikowski
The paper introduces SBERT2S1, a method that transforms Sentence-Transformers encoders into typed decision models for biomedical text, and presents BIODECIDE, a suite for evaluating such models, along with MEDLINE‑S1, a large training set of 243k decisions derived from NLM indexing. Experiments show that retrieval‑trained encoders improve zero‑shot matching of content‑bearing options, and that a prior‑fused residual (PFR) head benefits most from retrieval pre‑training, while a cross‑head (C) head generally outperforms PFR across all objectives. The authors also release code, data, and a model, and discuss calibration and reward‑normalisation effects in their RLCD recipe.
By Pritam Deka
arXiv:2606. 19625v2 Announce Type: replace-cross Abstract: We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B.
By Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, Mark Riedl
The paper reports that in English all‑words word sense disambiguation (WSD), the scarcity of high‑quality labels—not the models—has become the limiting factor. The authors introduce lexEN, a human‑adjudicated correction layer over the Maru2022 ALL_NEW benchmark, and SenseBench, a living leaderboard for LLM WSD evaluation. They show that frontier large language models reach about 95 % accuracy on lexEN‑v1, that relabeling corpora with these models improves downstream systems, and that fine‑grained WordNet senses are often ill‑posed, with coarsening improving both annotator agreement and model performance.
"whyItMatters":"The study highlights that improving label quality and managing annotation costs are now the critical challenges for advancing WSD performance, as model accuracy is already near its theoretical ceiling."
By Vassili Philippov, Amro Salman, Dmitrii Andreev, Penny Hands, Emil Kaiumov, Pavel Katunin, Anton Nikolaev
The paper introduces Murmur2Vec, a lightweight, alignment‑free embedding that uses k‑mer counts hashed with MurmurHash to create a compact representation for biological sequences. It provides a full theoretical analysis, including bias/variance formulas, a Johnson–Lindenstrauss‑style concentration bound, and an excess‑risk bound that clarifies the trade‑off between hash‑table size and classifier performance. Empirically, Murmur2Vec matches or surpasses a fine‑tuned 650M‑parameter ESM‑2 protein language model across several classification tasks, including SARS‑CoV‑2 spike lineage and HIV‑1 Env subtype identification.
By Sarwan Ali, Taslim Murad, Imdadullah Khan, Safi Faizullah