ToolAlignBench: Investigating Alignment Conflicts in Tool-Calling Enabled LLMs
arXiv:2607. 14285v1 Announce Type: cross Abstract: Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict?
arXiv:2606. 08381v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly released and deployed through opaque development and deployment pipelines, enabling model providers to inject intentional, provider-specific policies without officially announcing them.
arXiv:2607. 14285v1 Announce Type: cross Abstract: Safety alignment in LLMs aims to align models with human values, but which values take precedence when they conflict?
arXiv:2606. 00023v1 Announce Type: cross Abstract: The rapid development of Language Diffusion Models (LDMs) challenges the dominant position of auto-regressive competitors in language processing.
arXiv:2607. 18295v1 Announce Type: new Abstract: We study whether alignment schemes that reshape a base model's output distribution, combined with bounded safety filters, can drive the probability of harmful behavior to zero in modern large language models.
arXiv:2607. 22766v1 Announce Type: cross Abstract: The alignment of Large Language Models (LLMs) is increasingly bottlenecked by data quality.
arXiv:2607. 01208v1 Announce Type: cross Abstract: Language models deployed in high-stakes roles can potentially favor certain entities, brands, or viewpoints, steering user decisions at scale.
arXiv:2503.04332v2 Announce Type: replace-cross Abstract: The tremendous commercial potential of large language models (LLMs) has heightened concerns over their unauthorized use. To address this, we...
arXiv:2606. 08451v1 Announce Type: cross Abstract: Safety-aligned large language models often exhibit sycophancy, which is the tendency to affirm users' opinions regardless of factual accuracy.
Conformal Privacy Auditing (CPA) is a distribution‑free framework that calibrates re‑identification risk for each released document against large language model (LLM)‑empowered adversaries. It outputs a conformal ambiguity set of candidate identities that is guaranteed to contain the true identity with a user‑chosen confidence level under exchangeability, along with an interpretable leakage proxy derived from the set size. CPA supports both logit‑access and sampling‑only attackers, enabling audits of both open‑source and proprietary models, and demonstrates calibrated coverage across various benchmarks and attacker configurations.
The paper introduces Word-level Probability MIA (WPMIA), a black-box membership inference attack that estimates word-level generation probabilities via Monte Carlo sampling and local kernel smoothing, then aggregates them into a sequence-level likelihood estimator. By conditioning on different prefixes, WPMIA amplifies distributional differences between member and non-member texts, outperforming existing black-box baselines on open-source LLMs and achieving an average TPR@5%FPR of 42.0 on proprietary models such as GPT‑5‑Chat, Gemini‑2.5‑Flash, and Claude‑4.5‑Haiku.
arXiv:2607. 10252v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly consumed through opaque serving chains - API aggregators, resellers, and inference providers - in which the client has no technical means to confirm that the model answering is the model advertised, and recent audits show that a substantial fraction of commercial endpoints deviate from the vendor's reference weights.
arXiv:2601. 22313v2 Announce Type: replace Abstract: Large Language Models (LLMs) are rarely static and are frequently updated in practice.
The article surveys fake review detection research, focusing on how pre‑trained language models (PLMs) and large language models (LLMs) influence both the generation of deceptive reviews and their detection. It reviews 211 studies from 2018 to early 2026, categorizing methods by evidence source—such as review text, sentiment, rating behavior, temporal metadata, user‑product graphs, multimodal content, external knowledge, and LLM‑generated signals—and by fusion level. The survey traces the evolution from traditional machine learning to PLM‑based and LLM‑based approaches, evaluates performance on Amazon, Yelp, and OpSpam benchmarks, and highlights open challenges including adversarial generation, cross‑domain transfer, uncertainty‑aware fusion, robustness to missing sources, interpretability, and trustworthy evaluation of AI‑generated deceptive content.