Hugging Face Trending Papers

GraphDecide: Benchmarking System One Models on Graph Tasks

Read the original on Hugging Face Trending Papers →

GraphDecide is a model‑independent benchmark designed to evaluate System One models—such as Jev—that make decisions directly from supplied options on graph‑related tasks. The benchmark combines structural task profiles, matched graph‑text input contrasts, and heuristic‑proposal controls to diagnose graph decision performance. In testing fourteen model‑interface configurations, GraphDecide shows that accurate adjacency recognition does not guarantee broader structural correctness, joint graph‑text input does not consistently improve prediction, and feasible construction does not ensure high solution quality.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Aug 14

Unified Multi-Dimensional Benchmark for Complex Graph Reasoning in Large Language Models

arXiv:2608. 12391v1 Announce Type: cross Abstract: Graph reasoning provides a promising testbed for evaluating the reasoning ability of large language models (LLMs), as graph instances can be programmatically generated, structurally controlled, and naturally scaled to long-input settings.

By Fali Wang, Ali Al-Lawati, Iliyas Bektas, Jinxuan Fang, Alek Melenski, Tianxiang Zhao, Yao Ma, Suhang Wang
arXiv Machine Learning
Jun 11

GraphInfer-Bench: Benchmarking LLM's Inference Capability on Graphs

arXiv:2606. 11562v1 Announce Type: new Abstract: Graph analysis underlies many applications whose answers cannot be looked up in a single record or retrieved along a path: laundering rings, drug repurposing, user preference, and scientific theme are all inferred from a node together with its neighbourhood.

By Zhuoyi Peng, Jingzhou Jiang, Hanlin Gu, Lixin Fan, Yi Yang