arXiv Computation and Language

Shrome at Touch\'e: Soft-Vote Ensembling and Counter-Causal Augmentation for Causality Extraction

The paper presents a system for the Touché 2026 causality extraction challenge, focusing on counter‑causal claims—sentences that appear causal but actually deny causation. It tackles three subtasks: detecting causal sentences, extracting cause and effect spans, and labeling polarity (procausal, counter‑causal, or uncausal). The approach uses a fine‑tuned classifier with a cross‑task rule for detection, an ensemble of three RoBERTa‑large BILOU+CRF taggers for extraction, and counter‑causal data augmentation via a large language model for polarity classification, achieving state‑of‑the‑art scores on the CCNC test set.

arXiv Machine Learning
Aug 11

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

arXiv:2608. 09209v1 Announce Type: cross Abstract: Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs.

By Chidaksh Ravuru, Shashank Srivastava
arXiv AI
4d ago

Sentence-Level Context Sensitivity as a Training-Free Detector of Unsupported Content, Evaluated Against Trained Verifiers

The paper proposes a training‑free detector that uses sentence‑level context sensitivity to identify unsupported content in retrieval‑augmented generation (RAG) answers. By re‑scoring each sentence with full context, no context, and each chunk removed, the method flags the chunk whose removal most reduces a sentence’s likelihood as the likely source. Evaluated on RAGTruth, TofuEval, and RAGBench, the detector outperforms answer‑level faithfulness scores, achieving AUCs up to 0.745 and matching per‑chunk fact‑checkers while using only a fraction of the compute required by large‑language‑model judges.

By Mohamed Aly Bouke
arXiv Computation and Language
4d ago

Writerslogic at the CLEF 2026 SimpleText Track: Multi-Candidate LLM Simplification and Stacked Complexity Spotting

The Writerslogic team participated in the CLEF 2026 SimpleText shared task, tackling both text simplification (Task 1) and complexity spotting (Task 2). For simplification, they built a multi‑candidate pipeline with GPT‑4o‑mini, selecting the best candidate via a reference‑free heuristic, and their Claude Sonnet 4 submission achieved a SARI of 47.43 and BLEU of 14.21, ranking third overall on the Task 1 leaderboard. For complexity spotting, they fine‑tuned a DeBERTa‑v3‑large NLI model on 350 K labeled pairs, achieving a macro F1 of 0.8081 (0.8085 in an ensemble) on binary over‑generation identification and 0.804 accuracy on multi‑class error classification, placing them second among unique teams.

By David L. Condrey
Hugging Face Trending Papers
Aug 10

UNMASK: Discovering and Causally Verifying Spurious Shortcuts in Text Classifiers

Neural language models trained on large crowdsourced corpora frequently exploit spurious surface patterns tied to target labels without true linguistic or causal relevance, boosting benchmark performance while failing on adversarial or out-of-distribution inputs. Existing approaches either require manual specification of the feature vocabulary or automate discovery only partially, leaving the gap between dataset-level correlation and model-level exploitation unaddressed.

arXiv Computation and Language
Sep 25

Framing by Wording, Framing by Selection: A Large-Scale Two-Dimensional Audit of French News Headlines, 2022-2025

The study introduces a two‑dimensional framework to audit French news headlines, distinguishing salience framing—captured by four wording devices—from selection framing—captured by outlet‑level story form and high‑charge distributions. Using a 10,000‑headline supervision set annotated by LLMs and human arbitration, the authors classify 902,111 headlines from 25 outlets (2022‑2025) and find that salience and selection diverge yet correlate, that default thresholds inflate salience estimates, and that headlines mentioning Jews, the Far‑right, and Muslims exhibit the highest salience rates. The authors release their dataset, lexicons, and code, claiming it is the largest French headline framing audit to date.

By Amr Sobhy
arXiv Machine Learning
6d ago

When a Data Artifact Isn't a Shortcut: Causal Auditing of Synthetic RLVR Corpora

The paper audits whether synthetic distractors in RLVR corpora act as shortcuts for learning policies. A classifier using only surface statistics barely outperforms chance, and manual inspection reveals that code distractors are almost identical to correct answers. Experiments with a paraphrase‑matched control show no exploitation advantage for the unmodified data, indicating that the detectable artifact was not used by the policy.

By Esther Xin
arXiv AI
Sep 3

C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees

The paper introduces C$^{3}$T, a Counterfactual Causal Conversation Transformer that models sentiment shifts in social‑media conversation trees. It treats discourse moves such as denial, evidence, and toxicity as interventions, predicts node sentiment and shifts, and attributes sentiment changes to specific ancestor messages. The authors also present CaSiRe, a causal sentiment reasoning layer that enriches rumor conversation datasets with sentiment, shift, intervention, and causal‑source annotations, and demonstrate that C$^{3}$T outperforms baseline models in robustness and interpretability.

By S M Rafiuddin, Atriya Sen
arXiv AI
Sep 7

MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate

MABPD (Multi‑Agent Bias Probing & Detection) is a training‑free pipeline that uses three specialized large language model agents to analyze news articles from complementary perspectives and resolve disagreements via a Structured Argument Debate (SAD) protocol. SAD imposes an asymmetric burden of proof—biased claims lacking grounded textual evidence receive zero weight—along with role‑weighted voting and post‑consensus verification, replacing task‑specific supervised decision boundaries. Ablation studies show that the debate module alone accounts for up to a 10.6‑point F1 gain, and on the BABE benchmark MABPD attains 83.4% macro F1, within 0.7 percentage points of the supervised state‑of‑the‑art, while achieving 75.0% zero‑shot accuracy on the SemEval 2019 HyperPartisan corpus.

By Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India)
Hugging Face Trending Papers
Sep 2

C$^{3}$T: Counterfactual Causal Reasoning for Sentiment Shifts in Social-Media Conversation Trees

The paper introduces C$^{3}$T, a Counterfactual Causal Conversation Transformer that models sentiment shifts in social‑media conversation trees by treating discourse moves such as denial, evidence, and toxicity as interventions. It adds a causal sentiment reasoning layer, CaSiRe, to public rumor datasets, providing sentiment, shift, intervention, and causal‑source annotations. Experiments show that C$^{3}$T outperforms text‑only, graph‑based, and temporal baselines in predicting sentiment and attribution, revealing that denials and evidence reduce negativity while toxicity increases it.