Hugging Face Trending Papers

frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study

For more than 20 years, the Model-RB benchmark frb100-40 remained an open challenge; since 2014, its public record had stood at 99 of 100 variables. We give a directly checkable 100-vertex independent set for its 4,000-vertex graph.

arXiv AI
Sep 3

frb100-40 After Two Decades: An Optimality Certificate and a Preregistered Search Study

The paper presents a verified solution to the long‑standing frb100‑40 benchmark, providing a 100‑vertex independent set and a partition into 100 cliques of size 40, thereby proving the maximum independent‑set size is 100 and the minimum vertex‑cover size is 3,900. A preregistered experimental campaign of 8,668 runs evaluated new repair operators, finding no performance improvement over the baseline ULSA algorithm. Additional comparisons with other solvers (LibMVC‑NuMVC, group‑aware CSP pipeline) and exhaustive enumeration confirm the optimality certificate and characterize the search barrier for this instance.

By Onur U\u{g}urlu (\.Izmir Bak{\i}r\c{c}ay University)
arXiv AI
3d ago

Who Verifies the Graph? Misspecification Attacks on Causal Action Verification for Language Agents

The paper investigates how causal action verifiers, which guard language agents’ tool calls by checking identifiability against a committed action‑state graph, can be compromised through small graph misspecifications. By removing a single bidirected edge or reversing an arrowhead, the authors demonstrate that a verifier (CIVeX) that originally had zero false executions can suffer false execution rates up to 48.9%, with most of those executions being harmful and overall utility dropping dramatically. An additional attestation step that samples executions can detect these attacks with few false alarms, but it also leads to many wrongful rejections that reduce beneficial actions and incur significant experimental costs. whyItMatters":"The study shows that even minor errors in the verifier’s underlying graph can drastically undermine safety and performance, highlighting the need for robust auditing mechanisms."

By Fabio Rovai
arXiv Machine Learning
Sep 24

Does Graph Structure Earn Its Place in Microservice Root-Cause Analysis? A Controlled Study on RCAEval, and What the Benchmark Was Really Measuring

The paper investigates whether graph structure improves microservice root‑cause analysis by conducting a controlled study on the RCAEval benchmark. Using identical features, optimizers, and evaluation protocols across three model variants, the authors find no consistent advantage for graph‑based models over flat models, with a negligible Avg@5 difference (0.003, p=0.844). They identify two benchmark properties—limited fault injection and a non‑uniform telemetry schema—that bias results, and propose a new model, PSC‑GRCA, which achieves higher Avg@5 mainly through a system prior rather than graph information.

By Imad Bulji\'c
arXiv Machine Learning
Sep 2

Deterministic LLM Inference Across GPU Kernels: Power-of-Two INT8 Quantization Scales and the Limits of Tolerance-Based Conformance

The paper evaluates the effectiveness of tolerance‑based conformance tests for INT8 quantized GEMM kernels used in large language models. By injecting nine faults into a Qwen3‑1.7B reference pipeline, the authors show that most faults shift outputs by at most one bfloat16 spacing, rendering a tolerance of one spacing blind to these errors. They further demonstrate that requantizing weight scales to the nearest power of two aligns CUTLASS and Triton implementations bit‑for‑bit and produces identical token sequences, with only minor perplexity changes.

By Teng-Ruei Chen
arXiv Machine Learning
Aug 11

Quality-Diversity Stress Tests for Process Reward Models:What Archive Coverage Can and Cannot Certify

arXiv:2608. 08008v1 Announce Type: new Abstract: Process reward models (PRMs) score intermediate reasoning steps and are widely used for search, ranking, and training, but optimization can exploit these learned proxies by increasing reward while turning correct reasoning into incorrect reasoning.

By Ibne Farabi Shihab, Fariya Afrin