arXiv Machine Learning By Chongru Fan, Wentao Huang, Wei Wang, Zhenquan Ding, Jinqiao Shi, Wei Cai, Zhiyu Hao

FlowAtom: Atom-Based Evidence Aggregation for Multi-Label Website Fingerprinting

Read the original on arXiv Machine Learning →

FlowAtom is a new method for multi‑label website fingerprinting that aggregates evidence from encrypted traffic flows into shared prototypes called Atoms. It pre‑trains a flow encoder on unlabeled traffic, then combines Atom responses across flows within each observation window to produce a fixed‑dimensional, permutation‑invariant representation for predicting which monitored websites were visited. In closed‑world tests on Direct HTTPS, Trojan, and VMess traffic, FlowAtom achieves micro‑F1 scores of 97.82%, 94.43%, and 93.92% respectively, and consistently outperforms baseline approaches in open‑world scenarios involving monitored visits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 11

Privacy-Preserving Data Drift Detection and Recovery for Large-Scale LLM Applications via Proxy Representations

arXiv:2608. 08245v1 Announce Type: cross Abstract: LLM applications deployed at scale face a fundamental challenge: privacy constraints prevent direct inspection of user interactions, making it difficult to obtain any representative evaluation dataset or to track the ongoing evolution of production traffic.

By Michael Levit, Josh Ledgard, Haoyu Dong, Vishwas Suryanarayanan, Eyal Kolman, Sharon Tan, Qiang Gan, Vishal Chowdhary
arXiv Machine Learning
Sep 3

SoK: Where Do Flow Labels Come From? Auditing Label Provenance in Encrypted Traffic Benchmarks

The paper audits 14 encrypted traffic classification benchmarks to investigate how flow labels are generated. It finds two common labeling strategies—coarse inheritance, which may mislabel flows, and overstrict filtering, which may discard useful flows—leading to inconsistencies between benchmark labels and actual traffic records. The study also quantifies the impact of these labeling practices on classifier accuracy, showing that inherited labels limit balanced accuracy to 0.56–0.76, while filtering can raise macro accuracy from 0.44 to 0.65.

By Sizhe Huang, Shujie Yang
arXiv AI
Sep 4

Identifying AI Web Scrapers Using Canary Tokens

The paper introduces a method to detect which web scrapers feed data to large language models (LLMs) by deploying dynamic websites that issue unique canary tokens to each scraper. By querying LLMs for information about these sites, the authors can identify when an LLM consistently outputs the unique tokens, indicating exposure to a specific scraper. Experiments on 22 production LLM systems show the technique reliably uncovers both known and undisclosed scrapers, offering a tool for third parties to monitor and control unwanted web scraping.

By Steven Seiden, Triss Ren, Caroline Zhang, Taein Kim, Enze Liu, Emily Wenger
arXiv Computation and Language
Sep 16

RiskChainBench: A Benchmark for Obfuscated Platform Message Restoration and Evidence-Grounded Web Investigation

RiskChainBench is a new benchmark that pairs 3,600 synthetic token‑text restoration inputs with 600 human‑labeled local web environments to evaluate how well models can recover obfuscated platform messages and then investigate the associated websites. The benchmark measures both message restoration accuracy and the subsequent web‑investigation decision, using a fixed multimodal evidence judge to assess faithfulness, sufficiency, completeness, and consistency. Across ten models, performance varies widely, with entry recovery ranging from 35.2% to 95.2% and web decision accuracy from 26.3% to 62.8%, highlighting execution failures as the main bottleneck.

By ZhuoXin Liu, Zhiming Ma, Ying Zhang, Mengzheng Yang, Yifan Wang, Zhengqi Huang, Yanhan Zhou, Zekun Lin, Jun Zhang, Shun Zhang, Yue Chen, Qiao Zhao, Peng Chen