Thomas Dalton

Photo of Thomas Dalton

## about

I am a principal data scientist and cybersecurity machine learning engineer focused on applying AI/ML to cybersecurity systems.

## research interests

My interests center on applying machine learning and AI to cybersecurity, across the full model development lifecycle: from problem formulation and data collection to model evaluation, deployment, and monitoring.

Topics I've focused on recently include representation learning, adversarial and imbalanced data, phishing and web threat detection, malware and malicious code analysis, network traffic, and data loss prevention.

## news

2026Started as Principal Data Scientist at OpenText.
2026Filed U.S. patent applications on improving the quality of collected web data and on training website classifiers with incomplete data.
2025Completed an M.S. in Statistics at The University of Texas at Austin.
2025Patent applications on ML classification of cloud service HTTP(S) traffic published (US20250365339A1, US20250365340A1).
2025Released the PhreshPhish preprint (arXiv:2507.10854) and dataset.

## selected publications

[1]

PhreshPhish: A Real-World, High-Quality, Large-Scale Phishing Website Dataset and Benchmark

T. Dalton, H. Gowda, G. Rao, S. Pargi, A. Hadj Khodabakhshi, J. Rombs, S. Jou, M. Marwah
Under review · arXiv:2507.10854, 2026
[arxiv] [dataset]
[2]

Classifying Malware Using Function Representations in a Static Call Graph

T. Dalton, M. Schmidtler, A. Hadj Khodabakhshi
CSoNet 2020, Springer LNCS · arXiv:2012.01939
[arxiv]

[ all publications & patents ]