CRAFT: Clustering Rubrics to Diagnose Weak LLM Capabilities and Generate Targeted Fine-Tuning Data
arXiv:2607. 16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance.
arXiv:2605. 14568v3 Announce Type: replace-cross Abstract: Context.
arXiv:2607. 16122v1 Announce Type: new Abstract: Evaluations should do more than measure a models current performance.
arXiv:2607. 19442v1 Announce Type: cross Abstract: Machine unlearning is commonly evaluated by matching a retrained oracle on trained probes.
Diff Mining is a framework that identifies what a finetuned language model has learned by comparing its logits to those of its base model. It extracts per-context logit differences on a reference corpus and aggregates them into an interpretable token set using either a Top‑K frequency method or Non‑negative Matrix Factorization. The approach outperforms existing model‑diffing methods in domain detection and bias identification, and it requires only access to output logits, making it scalable to large models.
arXiv:2608. 10216v1 Announce Type: cross Abstract: Agent frameworks ship quality gates that compare text blocks by embedding-cosine similarity and decide at a fixed cutoff.
arXiv:2608. 00144v2 Announce Type: replace Abstract: Membership inference (MIA) on language models is usually summarised by aggregate ROC-AUC, but such evaluations are confounded: model-free blind baselines can separate members from non-members using surface text alone.
arXiv:2607. 04108v1 Announce Type: new Abstract: Large language models are increasingly used as evolutionary engines for scientific discovery: generate candidates, select winners, feed them back as parents, and repeat.
arXiv:2609.14976v1 Announce Type: new Abstract: Long-horizon LLM agents accumulate memory across sessions, creating sparse but high-impact risks: stale facts, conflicting updates, cross-user leakage,...
arXiv:2607. 18360v1 Announce Type: cross Abstract: Large language models (LLMs) now routinely draft literature reviews and assist with academic writing, which means a higher risk of fabricated references: GPTZero found 53 papers with hallucinated citations among NeurIPS 2025's accepted set.
arXiv:2605. 03534v2 Announce Type: replace-cross Abstract: Retrieval-augmented generation (RAG) grounds answers in retrieved passages, yet relevance does not guarantee sufficiency: a topical passage may still fail to justify the answer.
arXiv:2608. 05162v1 Announce Type: cross Abstract: Pooling is a consequential but under-examined design choice in decoder-only concept representation work: practitioners must collapse token-level hidden states into a passage-level vector, yet no shared protocol exists for comparing this choice across concepts, models, and tasks.
The study evaluates how the choice of test boundary affects feature‑based hardware Trojan detection across Trust‑Hub families. Using a corpus of 49,124 gates from 16 netlists, the authors compare three test settings—pooled gates, a single netlist held out, and an entire host family held out—showing that performance drops markedly when a host family is excluded. The results demonstrate that sibling benchmark variants can inflate detection metrics, and the authors recommend reporting family‑aware holdouts alongside pooled scores.
The paper investigates code-level autonomous research loops (ARLs) where a language model edits training pipelines to improve an in-loop metric. It identifies a failure mode called algorithmic mode collapse, where edits become semantically uniform despite surface diversity, leading to a growing gap between in-loop gains and independent evaluation. The authors propose Diversity‑Aware Proposal Sampling (DAPS), a lightweight method that reduces semantic decay by 69.1% and boosts faithfulness by over 80% while maintaining optimization speed.