Multi-Source Cybersecurity Logs: An ATT&CK-Labeled Dataset and SLM Evaluation
arXiv:2606. 18190v1 Announce Type: cross Abstract: Multi-stage cyberattacks span system, network, and browser logs.
arXiv:2507. 13505v2 Announce Type: replace-cross Abstract: Cybersecurity simulation environments, such as cyber ranges, honeypots, and sandboxes, require realistic human behavior to be effective, yet no quantitative method exists to assess the behavioral fidelity of synthetic user personas.
arXiv:2606. 18190v1 Announce Type: cross Abstract: Multi-stage cyberattacks span system, network, and browser logs.
arXiv:2608. 13575v1 Announce Type: cross Abstract: Recent machine learning (ML) advances have demonstrated that deep learning (DL) achieves impressive results in different application domains, including the classification of computer network traffic to corresponding applications.
arXiv:2609.38397v1 Announce Type: new Abstract: Virtual clients offer a cost-effective approach to support applications such as A/B testing, recommender system development, and interface evaluation....
arXiv:2606. 30801v1 Announce Type: cross Abstract: Personalization algorithms determine what content users encounter on online platforms.
arXiv:2609.01257v1 Announce Type: new Abstract: As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. Whil...
arXiv:2604. 04611v2 Announce Type: replace Abstract: Federated learning (FL) enables multiple clients to collaboratively train a global model by aggregating local updates without sharing private data.
arXiv:2608. 08245v1 Announce Type: cross Abstract: LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representative evaluation dataset or to track the ongoing evolution of production traffic.
arXiv:2607. 00763v1 Announce Type: cross Abstract: Digital forensic investigations of network intrusions require analytical outputs that are traceable, reproducible, and court-defensible - requirements existing machine learning pipelines do not satisfy, since they treat original evidence as training data and produce opaque classifications without instance-level justification.
arXiv:2607. 03334v1 Announce Type: cross Abstract: The federated learning (FL) paradigm fosters distributed pervasive computing combined with artificial intelligence techniques, allowing for optimized data usage and improved mitigation of privacy concerns.
ZeroHAT is a new framework for generating synthetic human activity traces (HATs) in a target region without any real data from that region. It transfers behavioral patterns learned from real HATs in source regions and adapts them using publicly available contextual information about the target region. The system includes a consistency-aware intent extractor, a cross-region behavioral cloning module, and a behavior-conditioned activity realization module, and it outperforms the strongest baseline by 4.5–6.4× in downstream utility and improves fidelity by 15.6–40.8% across ten cities.
arXiv:2607. 13123v1 Announce Type: cross Abstract: Cybersecurity is the practice of protecting systems, networks, and data from digital attacks.
arXiv:2602. 02838v2 Announce Type: replace-cross Abstract: The detection of online influence operations -- coordinated campaigns by malicious actors to spread narratives -- has traditionally depended on content analysis or network features.