arXiv AI By Aarya Bodhankar, Aditya Joshi, Bao Gia Doan, Thomas Marchant, Oscar Leslie, Flora Salim

Didact: A Cross-Domain Capability Discovery System for Defence

Read the original on arXiv AI →

arXiv:2606. 06942v1 Announce Type: cross Abstract: Policymakers in defence and defence-aligned sectors must monitor rapidly evolving research alongside sector priorities relevant to operational and strategic needs.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 25

W-RAG: Source-Aware Retrieval for Enterprise Document Generation from Heterogeneous Knowledge Bases

W-RAG is a source-aware retrieval framework designed for enterprise document generation from heterogeneous knowledge bases. It uses ontology-guided retrieval, local ranking within each knowledge base, and source-level weighting to balance evidence from diverse sources. A new dataset covering multiple document types and industry domains demonstrates that W-RAG improves document coverage and generation quality compared to standard RAG pipelines.

By Hridya Dhulipala, Rajesh Ombase, Michael Wang, Tien N. Nguyen
arXiv AI
Sep 10

Evidence-Grounded Retrieval for Investigation Hunt Lead Generation from CTI Reports

The paper introduces AHLERT, a system that automatically extracts environment-aware hunt leads from Cyber Threat Intelligence reports. It combines a hybrid retriever—dense vector search plus multi-hop knowledge‑graph traversal seeded with MITRE ATT&CK—with ontology‑grounded retrieval‑augmented generation to constrain leads to a defender’s assets. Evaluations on public CTI reports show that AHLERT doubles mean F1 scores and achieves an effectiveness score of ~86.95% compared to off‑the‑shelf LLM models.

By Akash Prakash, Boubakr Nour, Makan Pourzandi, Chadi Assi, Mourad Debbabi
arXiv AI
Sep 2

Verifiable Disaster Storylines and Causal Knowledge Graphs: A Citation-Grounded Pipeline from Heterogeneous Humanitarian Sources

The paper introduces a pipeline that merges structured disaster records from EM‑DAT with unstructured documents from ReliefWeb and the European Media Monitor to generate source‑grounded disaster storylines and causal knowledge graphs. Using Retrieval‑Augmented Generation, it produces tabular event profiles covering 17 fields and builds causal graphs enriched with citation‑grounded explanatory narratives, allowing traceability to primary sources. Human evaluation across three crisis cases shows high retrieval precision, strong faithfulness of causal relations, and a clear expert preference for citation‑grounded components over ungrounded ones.

By Ivan Decostanzi, Michele Ronco, Sergio Consoli, Christina Corbane, Lorenzo Bertolini, Indaco Biazzo, Daria Mihaila, Manuel Garcia-Herranz, Felix Schwebel, Yelena Mejova, Kyriaki Kalimeri
arXiv AI
Oct 2

Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning

The paper surveys recent advances in Retrieval Augmented Generation (RAG), a technique that integrates external retrieval into language model generation to reduce hallucinations and keep knowledge current. It introduces a four‑axis taxonomy—efficiency, defense, interactivity, and reasoning—to organize contemporary RAG research, covering retrieval methods, fusion strategies, embedding optimizations, and reinforcement learning policies. The survey also reviews evaluation practices, domain‑specific applications, and architectural variants, while highlighting ongoing challenges such as retrieval quality, reliability, domain adaptation, scalability, and explainability.

By Meghana Sunil, Shravya V, Shravan Venkatraman, Joe Dhanith PR