[1]
PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark
T. Dalton, H. Gowda, G. Rao, S. Pargi, A. Hadj Khodabakhshi, J. Rombs, S. Jou, M. Marwah
Under review · arXiv:2507.10854, 2026
The largest publicly available phishing website dataset and benchmark, with temporal splits and leakage
control via locality-sensitive hashing (LSH). Includes baseline evaluations of an encoder-only
transformer, a feed-forward neural network, and a linear support vector machine. Downloaded 25,000+
times on Hugging Face.
[2]
Classifying Malware Using Function Representations in a Static Call Graph
T. Dalton, M. Schmidtler, A. Hadj Khodabakhshi
CSoNet 2020, Springer LNCS · arXiv:2012.01939
RNN-based seq2seq autoencoder function embeddings combined with Weisfeiler–Lehman graph kernels over
static call graphs for malware family classification.
[3]
Cloud Services Intelligence: ML Classification of Cloud Service HTTP(S) Traffic
T. Dalton, J. Rombs
US20250365339A1 & US20250365340A1, 2025
A proxy-deployed classification system that infers cloud service actions (e.g., upload, download,
login, edit) from HTTP(S) request and response metadata, enabling real-time monitoring of cloud service
usage and potential data exfiltration.
[4]
Methods to Improve Quality of Collected Web Data
T. Dalton, M. Marwah, A. Hadj Khodabakhshi, J. Rombs
U.S. patent pending, 2026
Cleaning web-scraped phishing data by grouping pages with LSH, manually reviewing one prototype per
group, and removing near-neighbors of rejected prototypes (e.g., takedown notices, cloaked or redirected
pages).
[5]
Methods for Training a Website Classifier with Incomplete Data
T. Dalton, M. Marwah
U.S. patent pending, 2026
Training website classifiers on incomplete web data by imputing missing components (e.g., HTML from
URLs) with a generative model trained jointly with the classifier.