arXiv AI By Alexandre Cristov\~ao Maiorano

Isolating LLM Alignment from Regex: Zero Coverage and Metric-Dependent Divergence Under Adversarial Mutation

Read the original on arXiv AI →

arXiv:2607. 20494v1 Announce Type: new Abstract: Production LLM applications commonly stack a regex filter in front of model-side alignment; prior work found no measurable coverage gain from adding a live Gemini backend behind an active regex filter.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 3

Which Defense Closes Which Threat? Attributing OWASP-LLM-Top-10 Coverage and Its Brittleness Under Paraphrasing

arXiv:2606. 02822v1 Announce Type: cross Abstract: Production LLM applications stack several defense families -- refusal-phrase filters, token-budget controls, model allowlists, rate limits, tool-registry authentication -- yet existing breach-and-attack-simulation (BAS) benchmarks report a single aggregate coverage number, hiding which family closes which threat.

By Alexandre Cristov\~ao Maiorano
arXiv Machine Learning
Aug 27

OpenSanctions Pairs: Large-Scale Entity Matching with LLMs

OpenSanctions Pairs is the first large‑scale public benchmark for entity matching on sanctions and OSINT data, comprising 755,540 expert‑labeled pairs drawn from over 1 million entities across 293 source datasets and 45 jurisdictions. The dataset spans multiple languages and writing systems, inconsistent structures, and time‑varying provenance, making it far more heterogeneous than prior benchmarks. Baseline experiments show a rule‑based matcher achieving 91.3 % F1, GPT‑4o reaching 99.0 % F1, and a locally deployable open‑source model scoring 98.2 % F1, with complementary failure modes that highlight the need to focus on downstream pipeline components.

By Chandler Smith, Magnus Sesodia, Friedrich Lindenberg, Christian Schroeder de Witt
arXiv AI
2d ago

On-Device Named-Entity Recognition: A Deployability Study of Accuracy, Cost, Reliability, and Confidence

The paper evaluates nine on‑device named‑entity recognition models ranging from classical taggers to large language models, measuring not only accuracy but also latency and output validity. Using a silver‑gold benchmark derived from an LLM judge panel and a human‑validated corpus, the study shows that encoder‑based models achieve comparable accuracy to a 4 B instruct LLM while being much smaller, faster, and producing no malformed output. Confidence calibration of GLiNER is analyzed, revealing over‑confidence but improved reliability after temperature scaling and thresholding.

By Vinay Kumar Chaganti