Hugging Face Trending Papers

Structured Data Extraction from Real Estate Documents using Clustering, Classification, and Large Language Models

Read the original on Hugging Face Trending Papers →

Real estate property listings expose structured metadata through the API. Still, the richest property-level information (i.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Sep 10

Scaling E-Commerce Attribute Extraction with Parallel Decoding

The paper presents a two‑stage large language model pipeline for e‑commerce attribute extraction. First, it discovers a compact, ranked set of purchase‑discriminative attributes for each product category; second, it extracts those attribute values from catalog text using a fine‑tuned Qwen3‑4B model with Hyper‑Parallel Decoding. The approach attains 85% extraction accuracy while cutting inference costs by 92%, enabling scalable production use and producing consistent, comparable product knowledge bases.

By Nikhita Vedula, Dushyanta Dhyani, Bryan Wang, Shervin Malmasi
arXiv AI
Aug 7

Schema-Guided Hierarchical Information Extraction and Semantic Evaluation Using Generative AI

arXiv:2608. 06167v1 Announce Type: new Abstract: We present a schema-based framework for extracting complex, structured information from unstructured text documents using generative AI, followed by automated semantic evaluation of the extracted information against a gold standard.

By Modhurita Mitra, Jan-Willem Versteeg, Maarten D. Schermer, Shiva Nadi Najafabadi, Marie L. De Bruin, Lourens T. Bloem
arXiv AI
Aug 3

An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous Documents

arXiv:2607. 28662v1 Announce Type: new Abstract: Large language models extract entities and relationships from unstructured documents fluently but inconsistently: type vocabularies fracture across documents, the same person surfaces under several name variants, relationships duplicate, and distinct individuals who share a name risk silent conflation.

By Vaibhav Dangaich, Kevin Lewis, Kundeshwar Pundalik