arXiv Computation and Language By To Duy Hinh, Nguyen Le Quoc Anh, Phan Van Tri, Khuong Nguyen-An

Automated Extraction of Records of Processing Activities (RoPA) Using Hybrid RAG and Locally Deployed Large Language Models

Read the original on arXiv Computation and Language →

The paper introduces RoPA Manager, a system that automates the extraction of Records of Processing Activities (RoPA) required by Vietnam’s new Personal Data Protection Law. It combines hybrid retrieval techniques—lexical ranking, dense‑vector search, and Reciprocal Rank Fusion—with locally deployed large language models to avoid data‑sovereignty issues. A Vietnamese RoPA benchmark of 32 organizations and 77 processing activities was created, and the system achieved robust scorer performance (F1 ≈ 0.95) and moderate end‑to‑end token coverage (≈ 50‑55%).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Jul 21

RAGAL: A Frugal, Fully Local Retrieval-Augmented Assistant for Technical Support at a Government Agency

Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premises, ruling out cloud-hosted language models entirely. We report on RAGAL, a retrieval-augmented assistant for the technical-support team of AFIR, the Romanian Agency for Financing Rural Investments, built and operated under three hard constraints: zero data egress (no external API calls, even for synthetic data), a read-only mandate (the assistant drafts, humans execute), and a single 8 GB consumer laptop as the only development and training machine.

arXiv Computation and Language
Sep 1

Cloud and On-Premises Deployment of Uzbek Legal RAG via Targeted Retriever Fine-Tuning

The paper reports on building a retrieval‑augmented legal assistant for Uzbek that operates in both a managed cloud service and an on‑premises deployment. It introduces two new domain benchmarks—one for retrieval and one for end‑to‑end QA—and shows that fine‑tuning an open‑weight text embedder (UTE‑1) can close the performance gap with proprietary models under tight cost and latency constraints. The authors also provide negative results for a QLoRA experiment and release the benchmarks, evaluation code, and the fine‑tuned embedder for future low‑resource legal NLP work.

By Tatul Danielyan, Mariam Avetisyan, Hrant Davtyan
arXiv AI
Sep 15

Empirical Evaluation of Open-Source Large Language Models for Retrieval-Augmented Generation in ESG Domain

The paper evaluates seven open‑source large language models for retrieval‑augmented generation in the ESG reporting domain, using 498 real‑world ESG reports from EU‑listed companies and 100 synthetic QA pairs. Performance is measured with RAGAS metrics, showing strong retrieval scores but variable generation quality, especially in faithfulness and factual correctness. The results highlight significant differences across model architectures and underscore the need for domain‑specific fine‑tuning to improve factual accuracy.

By Motaz Saad, Anna Borrelli, Ivan Gentile, Kianna Kazemi, Francesco Piccialli, Antonella Longo
arXiv AI
Sep 3

Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search

Hybrid Retrieval-Augmented Generation with Knowledge Graph Expansion, RRF Fusion, and Per-Chunk Grounded Evaluation for Enterprise Document Search describes DocuSearch, an offline multi‑agent system designed for telecom network operations. The system combines semantic vector search, BM25 full‑text search, and knowledge‑graph neighbor expansion, merges the results via Reciprocal Rank Fusion, and reranks with a cross‑encoder before pruning with Maximal Marginal Relevance. A per‑chunk evaluation loop ensures only grounded answers are returned, achieving Precision@10 of 0.69, Recall@10 of 0.79, and an 89.6% grounding rate—improvements of 15, 16, and 18.4 percentage points over a dense‑only baseline.

By Harish Saragadam, Sudhanshu Sharma, Meghana Pujari