Hugging Face Trending Papers

Behavioral Reprogramming of Open-Weights Models: Cognitive Plasticity and Alignment Bounds

Read the original on Hugging Face Trending Papers →

Large language models (LLMs) are predominantly aligned to function as passive, sycophantic assistants. We challenge this default paradigm by empirically evaluating the cognitive plasticity of open-weight architectures when subjected to rigorous behavioral reprogramming.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
4d ago

PADM\'E: Preference Alignment Data Synthesis for Meta-Evaluation of LM Agent Evaluators

PADM'E is a method for synthesizing preference‑aligned data to meta‑evaluate language‑model (LM) evaluators of agentic behaviors. It reframes meta‑evaluation as a preference judgment problem, generating criterion‑based data with small LMs and no human involvement. In a prototype, PADM'E produced 1,000 samples across four domains and three criteria, and human validation showed agreement with human judgment rising from 73% to 85% compared to a naive baseline.

By Cheng Chang, Yining Mao, Peng Qi
arXiv Computation and Language
Sep 21

Cultural Alignment in Large Language Models Using Soft Prompt Tuning

The paper proposes a method for culturally aligning large language models (LLMs) using soft prompt tuning optimized via Differential Evolution (DE). Unlike traditional fine‑tuning or reinforcement learning, this approach keeps model weights frozen and requires no preference data, instead leveraging aggregated survey scores from Hofstede's Value Survey Module (VSM13). Experiments on four countries and four instruction‑tuned models show that DE‑optimized prompts reduce cultural discrepancy, improve agreement with the World Values Survey, and are preferred in blinded pairwise evaluations by LLM judges.

By Reem I. Masoud, Martin Ferianc, Philip Treleaven, Miguel Rodrigues
arXiv AI
Aug 11

Unified Hallucination Fuzzing for Multimodal Large Language Models

arXiv:2608. 07525v1 Announce Type: cross Abstract: Hallucination remains a persistent challenge for Multimodal Large Language Models (MLLMs), severely limiting their reliability in high-stakes applications.

By Pengfei Zhou, Jiajun Song, Zhiwei Tang, Yixing Ma, Xiaopeng Peng, Donghui Si, Yuhang Xu, Huiqi Song, Yiyuan Miao, Yichen Qian, Weihua Chen, Wangbo Zhao, Bohan Zhuang, Jiasheng Tang, Yang You