arXiv Machine Learning

Data Quality Rule Generation with LLMs

The paper introduces LeDQeR, an approach that uses large language models to automatically generate data quality rules for rule‑based enterprise tools. It follows a generate‑filter framework where the LLM proposes candidate rules from dirty data, and four filters ensure the rules are executable, correct, generalizable, and non‑redundant. Experiments show that LeDQeR produces effective, compact rule sets across diverse datasets and error types.

arXiv AI
Aug 20

From Threat Intelligence to Detection: Knowledge-driven Enrichment and Template-based Rule Grounding for Automated Sigma Rule Generation

The paper introduces AUTOSIGMA, an automated system that converts unstructured cyber threat intelligence reports into Sigma detection rules. It enriches input data with a structured knowledge base, matches it against existing Sigma rule repositories, and uses a large language model as a judge to validate the generated rules. Experiments on real-world APT reports and security blogs show that AUTOSIGMA outperforms other methods in rule validity, relevancy, MITRE ATT&CK coverage, and robustness to input quality.

By Sepehr Ghaffarzadegan, Boubakr Nour, Makan Pourzandi, Mourad Debbabi, Chadi Assi
arXiv AI
Jun 29

POTracker: Optimizing Large Language Models for Standard-Compliant Power Outage Report Generation

arXiv:2606. 23533v2 Announce Type: replace Abstract: Recent large language models (LLMs) are good at general text generation, but it is still hard to use them for domain-specific data generation because the output must follow strict formatting and structural rules.

By Hung Phan, Aniroop Naladala, Dubey Avanindra, Supryia Chinthavali, Lunga Dalton, Ali Jannesari
arXiv AI
Jun 6

PSEBench: A Controllable and Verifiable Benchmark for Evaluating LLMs in Patient Safety Event Triage

arXiv:2606. 05463v1 Announce Type: new Abstract: Patient safety event triage, determining whether a clinical event is reportable under jurisdiction-specific policy, is a high-stakes task typically performed manually by patient safety experts.

By Keqi Han, Ryan Young, Annabel Strauss, Lindsey Hughes, Katharine M. Nesbitt, Nicole Schueler, Che Ngufor, Carl Yang, Yuan Xue, Zhijun Yin