Thomas Dalton

## publications

[1]

PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark

T. Dalton, H. Gowda, G. Rao, S. Pargi, A. Hadj Khodabakhshi, J. Rombs, S. Jou, M. Marwah
Under review · arXiv:2507.10854, 2026
The largest publicly available phishing website dataset and benchmark, with temporal splits and leakage control via locality-sensitive hashing (LSH). Includes baseline evaluations of an encoder-only transformer, a feed-forward neural network, and a linear support vector machine. Downloaded 25,000+ times on Hugging Face.
[arxiv] [dataset]
[2]

Classifying Malware Using Function Representations in a Static Call Graph

T. Dalton, M. Schmidtler, A. Hadj Khodabakhshi
CSoNet 2020, Springer LNCS · arXiv:2012.01939
RNN-based seq2seq autoencoder function embeddings combined with Weisfeiler–Lehman graph kernels over static call graphs for malware family classification.
[arxiv]

## patents

[3]

Cloud Services Intelligence: ML Classification of Cloud Service HTTP(S) Traffic

T. Dalton, J. Rombs
US20250365339A1 & US20250365340A1, 2025
A proxy-deployed classification system that infers cloud service actions (e.g., upload, download, login, edit) from HTTP(S) request and response metadata, enabling real-time monitoring of cloud service usage and potential data exfiltration.
[US20250365339A1] [US20250365340A1]
[4]

Methods to Improve Quality of Collected Web Data

T. Dalton, M. Marwah, A. Hadj Khodabakhshi, J. Rombs
U.S. patent pending, 2026
Cleaning web-scraped phishing data by grouping pages with LSH, manually reviewing one prototype per group, and removing near-neighbors of rejected prototypes (e.g., takedown notices, cloaked or redirected pages).
[5]

Methods for Training a Website Classifier with Incomplete Data

T. Dalton, M. Marwah
U.S. patent pending, 2026
Training website classifiers on incomplete web data by imputing missing components (e.g., HTML from URLs) with a generative model trained jointly with the classifier.