Beyond Scale and Generation: Understanding Language Model-based Entity Matching
arXiv:2607. 24688v1 Announce Type: cross Abstract: Entity matching identifies records that refer to the same real-world entity.
arXiv:2607. 24688v1 Announce Type: cross Abstract: Entity matching identifies records that refer to the same real-world entity.
arXiv:2608. 05163v1 Announce Type: cross Abstract: A common assumption holds that switching to a non-English language makes a multilingual RAG system easier to attack for personal information.
arXiv:2605. 29738v2 Announce Type: replace-cross Abstract: Legal NLP benchmarks overwhelmingly evaluate a single language or aggregate tasks that differ fundamentally across jurisdictions, making cross-lingual comparison impossible.
arXiv:2608. 02616v2 Announce Type: replace-cross Abstract: We present what is, to our knowledge, the first systematic evaluation of OpenAI's Privacy Filter (OPF), a 1.
arXiv:2608.23507v1 Announce Type: cross Abstract: Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or...
arXiv:2607. 25579v1 Announce Type: cross Abstract: Entity alignment (EA) identifies entities across knowledge graphs (KGs) that refer to the same real-world object.
Historical people may appear under different languages, scripts, and transcription traditions, while distinct individuals may share highly similar or even identical names. This makes historical identi...
arXiv:2608. 16394v1 Announce Type: new Abstract: Generating regulation-compliant test scenarios is essential for validating safety-critical automotive systems, yet Large Language Models (LLMs) struggle to ground outputs in long, hierarchical standards.
arXiv:2603. 15510v2 Announce Type: replace Abstract: The synthesis of inductive loop invariants remains a critical bottleneck in automated program verification.
arXiv:2606. 31718v1 Announce Type: cross Abstract: Relation extraction (RE) for low-resource languages is typically constrained by the lack of annotated corpora.
arXiv:2606. 05781v1 Announce Type: new Abstract: Deploying frontier large language models (LLMs) for domain-specific structured evaluation tasks often incurs substantial latency, cost, and data privacy overhead.
arXiv:2607. 09424v1 Announce Type: cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English.