arXiv AI By Xiang Wang, Wei Wei

What Does Chain-of-Thought Contribute at Probe Time? Evidence for Local Co-Occurrence Activation

Read the original on arXiv AI →

arXiv:2605. 26795v2 Announce Type: replace Abstract: Chain-of-thought (CoT) prompting enhances large language model performance, yet what drives these gains remains unclear.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Words Speak Louder Than Order: A Behavioral Evaluation of Gemma 4

The study investigates how Google’s Gemma 4‑e4b language model resolves conflicts between two documents. Using a counterbalanced design, researchers found that the semantic framing of a source (e.g., labeling it as an official guideline) dominates over the order in which documents appear. While the model shows a primacy bias toward the first document, this bias varies widely with wording and is amplified only when the documents are structurally identical.

By Amanda Fitch