arXiv Machine Learning

Do LLMs Make Neural Distinguishers Wise?

The paper investigates whether large language models (LLMs) can enhance neural distinguishers, a cryptanalysis technique that uses machine learning to recover secret keys from plaintext–ciphertext pairs. Experiments on SPECK-32/64 show that LLM-based distinguishers do not outperform traditional ResNet models, that difference choice loses effectiveness at higher rounds, and that incorporating XOR operation results into the prompt significantly boosts LLM performance.

arXiv Machine Learning
Jun 10

Do LLMsMakeNeural Distinguishers Wise?

arXiv:2606. 10692v1 Announce Type: cross Abstract: Neural distinguishers are a cryptanalysis method for symmetric-key cryptography that trains machine learning models on pairs of plaintexts and ciphertexts with specific differences in order to recover a secret key.

By Tatsuya Sakagami, Masashi Hisai, Naoto Yanai
arXiv AI
Sep 15

SpliTEE: Improving LLM Inference on Trusted Hardware with Differentially Private GPU Outsourcing

SpliTEE extends the split‑inference architecture of Slalom to large language models by protecting intermediate GPU computations with differential privacy rather than encryption. The authors show that masking intermediate representations is essential, as a prompt‑reconstruction attack can recover prompts with about 80% accuracy. Their global sensitivity analysis bounds the noise needed, and they demonstrate that SpliTEE on Intel TDX achieves near‑double the speed of fully CPU‑based inference and outperforms encryption‑based Slalom while maintaining higher accuracy.

By Shashie Dilhara Batan Arachchige, Robin Carpentier, Hassan Jameel Asghar, Dali Kaafar
arXiv AI
Sep 11

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning

The paper reports that large language models can acquire cipher-based covert communication skills without fine‑tuning, using prompting or in‑context learning instead. This enables new jailbreak attacks that bypass alignment safeguards by encrypting harmful requests, making them appear as nonsensical text to harmfulness classifiers. The authors demonstrate successful attacks against frontier models from Anthropic, Google, and OpenAI.

By Thomas Rivasseau