S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation introduces a hybrid diffusion approach that first applies autoregressive diffusion at high noise levels and then switches to parallel diffusion at low noise levels. This method coordinates interdependent events to produce valid state transitions while reducing sampling time compared to fully serial generation. Implemented with a pixel‑space diffusion transformer and a LoRA‑fine‑tuned pretrained video model, S2PD outperforms bidirectional baselines in rule adherence and achieves greater temporal stability and sampling efficiency across games, physical simulations, and real video.
The paper highlights that AI voice assistants using ASR and LLMs struggle with regional British accents because most ASR models are trained on American English. It introduces CavaBench, a benchmark of spoken financial queries, to evaluate ASR models and their impact on downstream tool‑calling accuracy across British accents. The study finds that while WER predicts tool‑calling accuracy, it may not fully capture task‑level performance, revealing accent‑related failures that vary by model and acoustic conditions.
The paper investigates how frontier AI systems perform on tasks related to cyber security, specifically exploit generation, vulnerability repair, and subsequent attacks, using five nonpublic software environments. It evaluates both open‑weight and proprietary models, employing deterministic graders rather than LLM judges to score performance. Results show significant variation across systems and vulnerability types, with repair scores higher than attack scores in two environments and lower in three, and highlight that passing an initial security test does not guarantee long‑term defense, as 92 of 524 defender test intervals still experienced successful exploits after the first one was stopped.
The paper introduces VERA, an automated framework that audits large language model (LLM) reasoning in software vulnerability analysis. Instead of relying on free‑form explanations, VERA requires models to produce a Structured Reasoning Record (SRR) that captures pointers, memory operations, and state transitions in machine‑readable fields. A multi‑stage judge then checks each SRR against eight reasoning failure modes, revealing that reasoning flaws are as common in correct verdicts as in incorrect ones and that VERA detects 87% of errors missed by free‑form LLM‑as‑judge evaluations.
GraphDecide is a model‑independent benchmark designed to evaluate System One models—such as Jev—that make decisions directly from supplied options on graph‑related tasks. The benchmark combines structural task profiles, matched graph‑text input contrasts, and heuristic‑proposal controls to diagnose graph decision performance. In testing fourteen model‑interface configurations, GraphDecide shows that accurate adjacency recognition does not guarantee broader structural correctness, joint graph‑text input does not consistently improve prediction, and feasible construction does not ensure high solution quality.
The paper introduces DUCB-OGD, an algorithm that couples a Discounted Upper‑Confidence‑Bound sampler with Online Gradient Descent to address dynamic minimax regret in robust large‑language‑model post‑training. It operates under instantaneous mini‑batch‑only bandit feedback, tracking worst‑source performance without re‑evaluating historical data. Experiments on fine‑tuning, preference optimization, and reinforcement learning demonstrate that DUCB‑OGD improves worst‑group robustness with negligible computational overhead.
DeferKV rethinks when to evict key‑value (KV) cache entries in long‑context large language models. By delaying eviction until the first decoding step and combining prompt‑side and decode‑side attention signals, it aligns KV importance estimation with actual generation needs. The method requires no extra training or modules and consistently improves performance on benchmarks while keeping latency low.
The paper introduces Guidance‑TTT, a method that separates strategic planning from execution in test‑time training for large language models. A small guidance model is trained at test time to propose high‑level changes, while a frozen, larger execution model implements these changes, reducing the cost of maintaining gradients and optimizer states. Guidance‑TTT achieves strong results across four domains—combinatorial optimization, heuristic programming, machine learning, and GPU kernel optimization—outperforming prior work and matching state‑of‑the‑art leaderboard scores.
Anlu introduces counterfactual supervision for in‑context time series anomaly detection, pairing each query with two contrasting reference records to enforce reference‑dependent learning. By adding a reference memory and gated adapters to a frozen time‑series foundation model, Anlu improves the mean VUS‑PR from 0.542 to 0.607 on 350 evaluation sequences. Replacing the reference with zeros drops performance to 0.499, highlighting the importance of reference conditioning.
The study evaluates an auditable patient‑timeline reconstruction system that tracks provenance, records revisions, and refuses to answer when evidence is missing. Using a synthetic corpus of 1,000 patients and 3,353 notes, two provenance‑aware Evidence Graph operators reduced graph size by 33–37% while preserving all answers across 6,813 query points; a fixed‑window baseline failed to answer over half of the points. The system’s evidence‑gating mechanisms (BioClinicalBERT and a zero‑shot LLM) responded appropriately to evidence‑unavailable controls, but performance varied on marker‑free controls, with BERT maintaining high accuracy but the LLM’s coverage dropping sharply.
"whyItMatters":"The results demonstrate that provenance‑aware evidence graphs can significantly reduce data complexity while maintaining answer integrity, highlighting a practical approach to building auditable clinical NLP systems."
MS-Exam-Gen is a reproducible framework that builds a source‑grounded multiple‑choice question benchmark for evaluating large language models on knowledge about multiple sclerosis MRI. The pipeline uses expert‑indexed sources, topic induction, evidence‑grounded MCQ generation, automated quality audits, and consistency checks to produce a 3,058‑item benchmark covering 16 topics and 53 subtopics. Evaluation of 12 LLM endpoints on this benchmark revealed a wide accuracy range (89.7% to 46.9%) and identified items frequently missed by models, while audits showed reduced answer cues and position‑sensitivity in scoring.
OpenAI has launched a new visual advertising format within ChatGPT, enhancing how ads are presented to users. The update also expands measurement tools, establishes attribution partnerships, and improves brand suitability options for advertisers.
The paper introduces TrustMI, a method to causally control how large language model assistants decide to trust their users. By creating 2,000 contrastive conversations that vary in ability, benevolence, and integrity, the authors learn steering matrices that adjust trust decisions along linear directions in model activations while keeping the model parameters frozen. Experiments across six instruction‑tuned models show that these steering changes reliably alter trust decisions and affect safety‑related behaviors such as compliance with harmful requests, prompt injections, and insider threats.
StagQ is a multi‑precision weight format for large language models that uses a 2‑bit group‑wise affine base followed by optional 1‑bit refinement planes. Each supported precision can be read as a prefix of the main stream, decoded via a shared affine map without per‑weight lookups, and a sparse side record stores the few weights that the grid handles poorly. Experiments show that StagQ outperforms baseline multi‑precision schemes on Llama‑3.1‑8B, Phi‑4, and OLMo‑2‑7B across various bit‑widths, and its GPU kernel is faster than baseline kernels for most shape‑precision combinations.
ThunderSyncRL is a training framework that eliminates idle time in agentic reinforcement learning by starting gradient computation immediately once all necessary inputs are available, thereby avoiding policy staleness. It applies to group relative policy optimization (GRPO) by computing trajectory score gradients as soon as rewards arrive, and to on‑policy distillation (OPD) by updating gradients for completed agentic turns while tool calls execute. Experiments on SWE‑bench Verified and Terminal Bench 4.0 show that ThunderSyncRL matches synchronous training performance up to 1.9× faster and outperforms asynchronous training by up to 2.47 percentage points at a fixed budget.
The paper investigates the challenges of optimizing compacted context models, particularly the KV cache, in continual learning scenarios. It identifies the optimization landscape as brittle and flat, and proposes a simplified Perceiver-based architecture that matches or surpasses full Perceiver transformers in continuous context compaction. Experiments on MCQ tasks in Finance, Legal, Gutenberg, and Code demonstrate the effectiveness of this approach.
The paper introduces a game-theoretic approach to text revision, treating token positions as players and vocabulary items as actions, with utilities based on a language model’s log conditional probability. It shows that Nash equilibria can yield exponentially higher likelihoods than autoregressive outputs as sequence length increases, and proposes Nash decoding, an algorithm that finds an ε-Nash equilibrium in O(1/ε) time. Experiments on CLAPNQ, PubMedQA, and CoQA demonstrate that equilibria derived from masked language models achieve higher F1 and ROUGE scores than autoregressive models, up to 18× larger, without fine-tuning, though with extra test-time computation.
The paper compares text‑based and feature‑based models for recognizing compound emotions in real‑world videos. It proposes textualizing non‑verbal cues from audio and visual modalities into text to leverage large language models, while feature‑based models directly combine extracted multimodal features. Experiments on the C‑EXPR‑DB dataset show that feature‑based models outperform textualization in the wild, though textual models can excel when rich transcripts are available.
By Nicolas Richet, Soufiane Belharbi, Haseeb Aslam, Meike Emilie Schadt, Manuela Gonz\'alez-Gonz\'alez, Gustave Cortal, Alessandro Lameiras Koerich, Marco Pedersoli, Alain Finkel, Simon Bacon, Eric Granger
The paper demonstrates that a pretrained symbolic music transformer already encodes jazz pianist identity sufficiently for accurate classification across two benchmarks. By adding cross‑attention over learned pianist embeddings, the model can generate music conditioned on a specific artist’s style, and evaluation protocols confirm that the generated continuations are correctly attributed to the intended pianist. Additionally, the classifier is repurposed to identify the most characteristic moments in a performance, revealing the musical gestures that distinguish each pianist’s voice.
By Drew Edwards, Akira Maezawa, Simon Dixon
The paper introduces a hardware-software co‑design framework that compresses Mixture‑of‑Experts (MoE) model weights into low‑precision, hardware‑native sparse representations, enabling efficient execution on Sparse Tensor Cores (SpTCs). By relaxing discrete support selection through continuous reparameterization, the method jointly optimizes quantized weights and supports a router‑weighted reconstruction objective, achieving up to 4.35 percentage‑point gains in joint sparse‑quantization accuracy while retaining 96.09% of the original model’s performance. A custom grouped sparse GEMM kernel further boosts inference speed, outperforming NVIDIA’s baseline by up to 1.65× and reducing latency by up to 4.03× on B200 GPUs.
By Kwanhee Lee, Namhoon Lee, Dan Alistarh