arXiv Computation and Language
Sep 14

Extracting Dataset Mentions in Forced Displacement and FCV Documents: A Weakly Supervised Framework with LLM-Based Label Refinement

The paper introduces a weakly supervised framework for extracting dataset mentions from forced displacement and Fragile, Conflict, and Violence (FCV) documents. It uses a lightweight model trained on general research literature to generate candidate mentions, which are then refined by a large language model that validates or rejects them and corrects boundaries. The refined annotations are augmented with synthetic and contrastive examples to fine‑tune the model, achieving 74.1% precision and 70.5% recall on a benchmark of 1,706 passages, with higher precision (89.5%) on passages that contain dataset references.

By Rafael Macalaba, Aivin V. Solatorio, Patrick Michael Brock, Olivier Dupriez
arXiv AI
6d ago

MiST: Mid-Training LLMs for Cybersecurity

MiST (Mid-trained Security Transformer) is a suite of 8B and 32B language models tailored for cybersecurity, achieving strong performance on public benchmarks. The approach uses a mid-training stage that adapts general pre-trained models to the domain by curating a compact, expert-vetted seed corpus and generating high-quality synthetic training data, rather than continual pre-training on large raw text. MiST checkpoints improve mean cybersecurity accuracy by +13.1 and +8.6 absolute percentage points over Qwen baselines for 8B and 32B models, respectively, and provide a stronger initialization for downstream task-specific fine-tuning and reinforcement learning.

By Oded Ovadia, Elad Ben Zaken, Elad Guttman, Orly Moreno Kadosh
arXiv Machine Learning
Sep 10

Strategic Doctrine Language Models (sdLM): A Learning-System Framework for Doctrinal Consistency and Geopolitical Forecasting

Strategic Doctrine Language Models (sdLM) are a learning‑system framework that integrates multi‑document attention, temporal encoding, and a doctrine‑consistency layer to enable multi‑document strategic reasoning with doctrinal consistency constraints and calibrated uncertainty. The authors evaluate sdLM on expert‑panel scoring of 47 strategic scenarios, doctrine consistency across 336 doctrine publications (12,847 statements), and geopolitical forecasting on 127 historical counterfactuals spanning 1945‑2020 over 12‑60 month horizons, showing higher strategic quality and better calibration than strong general‑purpose LLM baselines and competitiveness with human experts on long‑horizon judgments. Ablation studies, scaling trends, and deployment‑oriented performance/latency characteristics are reported to identify which components drive improvements and how they translate to operational settings.

By Olaf Yunus Laitinen Imanov, Taner Yilmaz, Derya Umut Kulali