CrackedPDFs: A Controlled Benchmark for Hidden Prompt Injection in PDFs
arXiv:2607. 19396v1 Announce Type: new Abstract: Document-based LLM systems often flatten a PDF before guardrails inspect it.
arXiv:2607. 09290v1 Announce Type: cross Abstract: In the digital era, Portable Document Format (PDF) is one of the most widely used file formats for storing and exchanging digital documents due to its platform independence and rich functionality.
arXiv:2607. 19396v1 Announce Type: new Abstract: Document-based LLM systems often flatten a PDF before guardrails inspect it.
Delphi Scanner is a static malware detection system for Windows PE files that balances efficiency and interpretability. It employs a convolutional neural network to model Windows API sequences and a rule‑based interpretation layer to map APIs to high‑level malicious capabilities. Tested on over 190,000 PE files, it achieves 95.35% accuracy with a 1.53 MB model, and demonstrates robustness against out‑of‑distribution samples and adversarial manipulations.
The paper introduces a neural network model that identifies text content file types, especially source code, with higher accuracy and speed than existing tools. Experiments on open-source files show the model is more accurate on average, runs about four times faster than Magika, and is 28% smaller in size.
arXiv:2606. 30586v1 Announce Type: cross Abstract: Most corporate workplace environments enforce policies and technical controls that limit the storage of sensitive data on client endpoints.
arXiv:2606. 30819v1 Announce Type: cross Abstract: Generative AI has emerged as a significant cybersecurity threat, with several recent attack campaigns leveraging LLMs to generate code for malicious purposes via scripting languages such as PowerShell.
Most corporate workplace environments enforce policies and technical controls that limit the storage of sensitive data on client endpoints. Consequently, ransomware operators have evolved variants that expand their attack surface from local systems to network drives and shared storage resources.
arXiv:2606. 30572v1 Announce Type: cross Abstract: Malware classification remains a challenging problem due to its inherent heterogeneity, the presence of packed binaries, and the diverse distribution of malware families.
SCRIPTIOC-BENCH is a benchmark designed to evaluate how well large language models can statically extract indicators of compromise (IOCs) from script-based malware. It contains 634 manually verified JavaScript, PowerShell, and VBScript samples and covers four IOC types—URLs, domains, IP addresses, and filesystem artifacts—while distinguishing between directly exposed and encoded indicators. Experiments show that even the best models achieve only 65.4 F1, and a false‑positive taxonomy is introduced to analyze error patterns, with two mitigations (deterministic string utilities and task‑specific adaptation) improving precision and shifting errors toward sample‑grounded mismatches.
arXiv:2606. 02834v1 Announce Type: cross Abstract: Malware analysis starts with the raw bytes of an executable program, and tools to "lift" these to higher-level representations, such as assembly, are expensive and subject to error.
arXiv:2606. 20436v1 Announce Type: cross Abstract: Malware analysts often inspect compiled binaries through decompiled pseudo-C, when source code is unavailable.
Script-based malware remains a prevalent attack technique. These scripts often contain indicators of compromise (IOCs) that provide actionable threat intelligence. However, statically recovering such...
arXiv:2509. 16749v1 Announce Type: cross Abstract: LLMs are increasingly pervasive in the security environment, with limited measures of their effectiveness, which limits trust and usefulness to security practitioners.