arXiv Machine Learning By Manu Nandan, Michael Brautbar, Edward Raff

Fast And Accurate Text Content File Type Identification

Read the original on arXiv Machine Learning →

The paper introduces a neural network model that identifies text content file types, especially source code, with higher accuracy and speed than existing tools. Experiments on open-source files show the model is more accurate on average, runs about four times faster than Magika, and is 28% smaller in size.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 18

Delphi Scanner: efficient and interpretable static malware detection via API sequence modeling

Delphi Scanner is a static malware detection system for Windows PE files that balances efficiency and interpretability. It employs a convolutional neural network to model Windows API sequences and a rule‑based interpretation layer to map APIs to high‑level malicious capabilities. Tested on over 190,000 PE files, it achieves 95.35% accuracy with a 1.53 MB model, and demonstrates robustness against out‑of‑distribution samples and adversarial manipulations.

By Bijied Brahimi, Vincent Cohadon, Gabriel Glazman, Rayan Al Mohaize, Omran Berjawi, Rida Khatoun
arXiv Computation and Language
Aug 27

VietAIDetector: An Open-Source Zero-Shot Detector for Vietnamese AI-Generated Text

VietAIDetector is an open‑source, zero‑shot tool for detecting Vietnamese AI‑generated text. It offers a Gradio web interface that accepts raw Vietnamese text, common file formats, scanned documents, and very long texts beyond typical LLM context limits. Built on a Vietnamese‑specific language model, it outperforms existing English‑centric methods on out‑of‑domain datasets and lets users choose detection thresholds based on F1, accuracy, or TPR@0.05FPR, with results viewable or downloadable as a PDF report.

By Trieu Hai Nguyen, Van-Dung Hoang