Hugging Face Trending Papers

An LLM-Based Automatic Sportscast Solution for Robot Soccer Matches

RoboCup has always been a scenario to develop systems that solve real-world problems. Driven by the main goal of playing against the 2050 FIFA World Cup champions, the RoboCup Soccer leagues need to constantly measure how the research community is progressing.

arXiv Computation and Language
Aug 21

StreamSoccer: Event-Driven Memory for Streaming Soccer Commentary

arXiv:2608. 19723v1 Announce Type: cross Abstract: Streaming video understanding requires models to causally update state as video arrives and organize growing history into semantic units that can evolve, persist, and be recalled under bounded computation and memory.

By Chenxi Shao, Bozhong Wang, Jiaxin Huang, Zhao Liu, Sunwei Zhu, Tianxin Hang, Gaoqi He, Yang Li, Changbo Wang
arXiv AI
Sep 18

Semantic Action Graph: A Shared Representation for Agent Grounding and Human Interpretation of Sports Highlights

The paper introduces the Semantic Action Graph, a lightweight domain schema that models a sports match using performer, action, recipient, moment, and state nodes linked by role, temporal, and outcome edges. This structure supports both an agentic pipeline for generating narrated highlights and a visual interface that lets viewers query and inspect the same representation. In a prototype called SportSAGE, 12 soccer fans reported satisfaction with the generated highlights and used the graph interface to search, navigate, and interpret match moments.

By Tica Lin, Deepak Chandran, Gauri Jagatap, Chen Chen, Andrea Fanelli, David Gunawan, Josh Kimball
arXiv AI
Aug 28

Temporally-Grounded Language Generation: Towards Real-Time Vision-Language Models

The paper introduces Temporally-Grounded Language Generation (TGLG), a benchmark that tests vision‑language models on their ability to produce semantically accurate and temporally precise utterances in real‑time settings. It identifies perceptual updating and contingency awareness as key capabilities, curates datasets from sports broadcasting and egocentric interactions, and proposes the TRACE metric to jointly evaluate semantic similarity and temporal alignment. The authors also present VLM‑TSI, a model that interleaves visual and linguistic tokens in a time‑synchronized manner, achieving better performance than a strong baseline yet still showing modest overall results, underscoring the challenge of real‑time VLMs.

By Keunwoo Peter Yu, Joyce Chai