← Back to all news
arXiv Machine Learning October 5, 2026 By Hamidreza Saghir

Useful Features, Backward Scores: OOD in Language-Model Trajectories

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

  • safety

One email a morning, machine-written

One email a day, machine-written, one click to leave. We never share your address.

Related stories

arXiv AI
Jun 4

Luminol-AIDetect: Fast Zero-shot Machine-Generated Text Detection based on Perplexity under Text Shuffling

arXiv:2604. 25860v2 Announce Type: replace-cross Abstract: Machine-generated text (MGT) detection requires identifying structurally invariant signals across generation models, rather than relying on model-specific fingerprints.

By Lucio La Cava, Andrea Tagarelli
llmsbenchmarkssafety
More like this →
arXiv AI
Aug 25

Text-ADBench: Text Anomaly Detection Benchmark Based on LLM Embeddings

arXiv:2507.12295v2 Announce Type: replace-cross Abstract: Text anomaly detection is a critical task in natural language processing (NLP), with applications spanning fraud detection, misinformation id...

By Feng Xiao, Jicong Fan
llmsragnlpbenchmarks
More like this →
arXiv AI
Sep 15

Rubrics as an Attack Surface: Stealthy Preference Drift in LLM Judges

arXiv:2602.13576v2 Announce Type: replace-cross Abstract: Evaluation and alignment pipelines for large language models increasingly rely on LLM-based judges, whose behavior is guided by natural-langu...

By Ruomeng Ding, Yifei Pang, He Sun, Yizhong Wang, Zhiwei Steven Wu, Zhun Deng
llmsbenchmarkssafety
More like this →
arXiv AI
Aug 18

HarmProfile: Characterizing Harmful Distributions in Frontier LLMs

arXiv:2608. 14577v1 Announce Type: cross Abstract: Frontier large language models (LLMs) safety evaluation has largely treated harmful generation as an attack outcome rather than as an object of analysis.

By Zhouyuan Ma, Yutao Wu, Hanxun Huang, Xiang Zheng, Xiao Liu, Yixin Cao, Zuxuan Wu, Xingjun Ma, Yu-Gang Jiang
llmsbenchmarkssafety
More like this →
arXiv Computation and Language
Sep 15

Data Attribution of Emergent Misalignment with Persona Features

arXiv:2608.11025v2 Announce Type: replace Abstract: Emergent misalignment (EM) is the phenomenon where fine-tuning a language model on a narrow task leads to harmful behavior in unrelated domains. A...

By Clemens Vetter, David Kacz\'er, Lucie Flek, Florian Mai
llmsroboticsfine-tuningsafety
More like this →
arXiv Machine Learning
Jun 30

ANVIL: Anomaly-based Vulnerability Identification without Labelled Training Data

arXiv:2408. 16028v4 Announce Type: replace-cross Abstract: Supervised-learning-based vulnerability detectors often fall short due to limited labelled training data.

By Weizhou Wang, Eric Liu, Xiangyu Guo, Xiao Hu, Ilya Grishchenko, David Lie
llmsbenchmarks
More like this →
About Pricing API Newsletter Sources Privacy Terms Refunds Accessibility Provider info Contact RSS

The Flow links to publishers and never republishes their articles. Summaries are machine-generated.

v1.1.0 · 5f852ea