Benchmarks and evaluation

Leaderboards, eval harnesses and ablations — the contested business of deciding which model is actually better.

16,065 stories · RSS feed

arXiv AI
Jul 23

Logic-Guided Data Extraction with Answer Set Programming and Large Language Models

arXiv:2607. 19365v1 Announce Type: new Abstract: When Large Language Models (LLMs) are used for semantic data extraction from unstructured text, producing candidate relational facts from natural language, they may remain unreliable for tasks requiring complex combinatorial reasoning and global consistency.

By Mario Alviano, Lorenzo Grillo, Nicola Leone, Fabrizio Lo Scudo
arXiv AI
Jul 23

Recovering Clinical Utility Under Differential Privacy: Empirical Validation of Adaptive Federated Aggregation on Heterogeneous Cardiovascular Datasets

arXiv:2607. 19403v1 Announce Type: cross Abstract: Validating federated learning frameworks on real clinical data is an essential step between proof-of-concept demonstrations in controlled synthetic environments and deployment in real multicenter healthcare settings.

By Rodrigo Tertulino, Laercio Alencar, Ricardo Almeida