TSWAP: A Multilingual Retrieval-Augmented Thai Wellness Advisor
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606. 24200v1 Announce Type: cross Abstract: Retrieval-augmented generation (RAG) in clinical settings increasingly requires multilingual retrieval against predominantly English evidence corpora.
arXiv:2608. 12138v1 Announce Type: cross Abstract: General-purpose large language models (LLMs) have recently been reported to match or exceed specialized clinical AI tools on medical benchmarks, but such comparisons draw on a narrow set of systems and on benchmarks developed largely in high-income settings.
arXiv:2608.21365v1 Announce Type: cross Abstract: As a low-resource language, Khmer presents several retrieval challenges, including limited annotated data, ambiguous word boundaries, weak support in...
The paper presents a smartphone‑compatible, retrieval‑augmented language model tailored to Bangladeshi statutory law. By distilling a 9‑billion‑parameter Gemma‑2 teacher into a 2‑billion‑parameter student using supervised fine‑tuning and QLoRA, the authors achieve significant gains in ROUGE‑L and BERTScore on an English benchmark while keeping the model lightweight (1.6 GB) and operable offline on a Pixel 6. The system retrieves from 36,029 statutory passages using a hybrid dense/BM25 approach, and cross‑lingual evaluation shows effective Bangla query handling against an English‑only corpus, with a practicing lawyer rating the responses highly in a pilot study.
The paper introduces RoPA Manager, a system that automates the extraction of Records of Processing Activities (RoPA) required by Vietnam’s new Personal Data Protection Law. It combines hybrid retrieval techniques—lexical ranking, dense‑vector search, and Reciprocal Rank Fusion—with locally deployed large language models to avoid data‑sovereignty issues. A Vietnamese RoPA benchmark of 32 organizations and 77 processing activities was created, and the system achieved robust scorer performance (F1 ≈ 0.95) and moderate end‑to‑end token coverage (≈ 50‑55%).
Public institutions hold large volumes of sensitive documents and support tickets that cannot leave the premises, ruling out cloud-hosted language models entirely. We report on RAGAL, a retrieval-augmented assistant for the technical-support team of AFIR, the Romanian Agency for Financing Rural Investments, built and operated under three hard constraints: zero data egress (no external API calls, even for synthetic data), a read-only mandate (the assistant drafts, humans execute), and a single 8 GB consumer laptop as the only development and training machine.