Hello there, I'm

Phillip

Data, AI & Machine Learning

I build machine learning systems end to end, from data pipelines and model training to APIs and deployment.

01 About

I'm a developer focused on machine learning and data-driven software. My recent work spans retrieval-augmented generation, NLP classification, LLM security research, and predictive modelling, usually taken all the way from raw data to a deployed, containerised service. I'm currently looking for data science, software and AI/ML engineering opportunities.

02 Projects

RAG Research Assistant

Locally hosted retrieval-augmented generation system for querying research papers through a chat interface. PDFs are converted to markdown, chunked, and embedded in a vector database; answers are grounded with per-page source citations. Runs fully offline with local models and Docker Compose.

  • Python
  • FastAPI
  • ChromaDB
  • Ollama
  • Docker

Powerlifting Total Predictor

Predicts a powerlifter's next competition total from 150,000+ IPF records in the OpenPowerlifting database. An XGBoost model cuts prediction error by 17% versus a persistence baseline (19.1 kg MAE, R² 0.959), served through both a REST API and a Streamlit UI.

  • Python
  • pandas
  • XGBoost
  • scikit-learn
  • FastAPI
  • Streamlit
  • Docker

Prompt-Injection Detector

Detects prompt-injection attacks against open-source LLMs by classifying statistical features of attention weights (entropy, variance, TF-IDF). The best configuration (LLaMA 3.2 1B + Support Vector Machine) reaches 71% accuracy on unseen data, outperforming competing detection tools.

  • Python
  • PyTorch
  • scikit-learn
  • pandas
  • NumPy
  • Hugging Face
  • LLaMA
  • LLM security

Political Ideology Analysis WIP

Places news articles on the political compass using a two-stage transformer pipeline: one model first filters for political content, then a fine-tuned DeBERTa classifier scores the article's political leaning, a task general-purpose MNLI classifiers handle poorly.

  • Python
  • PostgreSQL
  • Transformers
  • DeBERTa
  • NLP

UK Sanctions Data Pipeline

Cleans the UK financial sanctions list into an analysis-ready dataset: standardising names and aliases, consolidating scattered address fields, and preserving partial dates across ~4,000 individual records, prioritising data integrity over aggressive automation.

  • Python
  • pandas
  • NumPy
  • Data engineering

03 Skills

Languages

  • Python
  • SQL
  • Java
  • JavaScript
  • C

ML & Data

  • PyTorch
  • Hugging Face Transformers
  • XGBoost
  • scikit-learn
  • pandas
  • NumPy

Tools

  • Docker
  • FastAPI
  • Streamlit
  • ChromaDB
  • Ollama
  • Git
  • Jupyter Notebook

04 Contact

I'm open to software and machine learning roles — feel free to reach out or look through my work.