arXiv:2510. 21891v2 Announce Type: replace-cross Abstract: To deploy large language models (LLMs) in high-stakes application domains that require substantively accurate responses to open-ended prompts, we need reliable, computationally inexpensive methods that assess the trustworthiness of long-form responses generated by LLMs.
By Dhrupad Bhardwaj, Julia Kempe, Tim G. J. Rudner
arXiv:2608. 06908v1 Announce Type: cross Abstract: We propose Zero-phase Component Analysis (ZCA) whitening as a geometric pre-processing step for the Word Embedding Association Test (WEAT).
By Seitaro Ono, Senna Ross, Jun Saiki
arXiv:2606. 03029v1 Announce Type: cross Abstract: A core goal of computational social science is to discover interpretable differences in how language varies across outcomes of interest, such as political affiliation or instructional quality.
By Paiheng Xu, Jing Liu, Wei Ai
arXiv:2606. 07226v1 Announce Type: cross Abstract: Human creativity has emerged as a critical competency in the era of large language models.
By Tongzhou Yu, Mingjia Li, Hong Qian, Wenkai Wang, Zongbao Zhang, Yaoyu Jiang, Xiangfeng Wang, Aimin Zhou, Jiajun Guo
arXiv:2606. 05972v1 Announce Type: new Abstract: Causal graphs provide a high-level language for making mechanisms transparent.
By Nirit Nussbaum-Hoffer, Nitay Calderon, Liat Ein-Dor, Roi Reichart
arXiv:2607. 05679v1 Announce Type: cross Abstract: Language models (LMs) exhibit problematic biases, such as stereotypes.
By Damian Hodel, Jevin West, Aylin Caliskan
arXiv:2606. 19625v2 Announce Type: replace-cross Abstract: We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B.
By Glenn Matlin, Chandreyi Chakraborty, Saehee Eom, Mika Okamoto, Rayan Castilla, Louis Jaburi, Alvin Deng, Taywon Min, Lucia Quirke, Stella Biderman, Mark Riedl
The paper introduces MUtE, a dual framework that simultaneously erases concept-specific information from representations and generates counterfactual mappings. By deriving new erasure functions based on optimal bounds, MUtE imposes a translational bias on counterfactual trajectories, aligning with geometric properties of concepts in language models. The authors demonstrate that this approach improves downstream algorithmic fairness and enables the generation of counterfactual texts.
By Antoine Saillenfest
arXiv:2607. 10212v1 Announce Type: new Abstract: Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance.
By Nipun Misra, Vikranth Udandarao, Aanchal Gupta, Yogender Kumar, Manuj Mukherjee, Raghava Mutharaju
arXiv:2404.06349v3 Announce Type: replace
Abstract: The ability to understand causality significantly impacts the competence of large language models (LLMs) in output explanation and counterfactual r...
By Yu Zhou, Xingyu Wu, Jibin Wu, Liang Feng, Kay Chen Tan
The paper introduces C$^{3}$T, a Counterfactual Causal Conversation Transformer that models sentiment shifts in social‑media conversation trees by treating discourse moves such as denial, evidence, and toxicity as interventions. It adds a causal sentiment reasoning layer, CaSiRe, to public rumor datasets, providing sentiment, shift, intervention, and causal‑source annotations. Experiments show that C$^{3}$T outperforms text‑only, graph‑based, and temporal baselines in predicting sentiment and attribution, revealing that denials and evidence reduce negativity while toxicity increases it.
Knowledge Graphs (KGs) are increasingly constructed through automated extraction pipelines; however, such systems often introduce spurious or incomplete triples, which degrade downstream performance. Existing evaluation practices rely heavily on task-specific metrics or small-scale manual verification, offering limited insight into the structural and semantic fidelity of extracted graphs.