Theme: Dark
Patrick Poon

Patrick Poon

AI/ML Software and Data Engineer

LLMs, RAG Pipelines & Gemini (Vertex AI)

Applying AI/ML to solve real-world problems. Solving data engineering challenges. Exploring the intersection of technology and creativity. Enthusiast of global travel & street gastronomy.

NEW YORK, NYAI/ML • COMPUTER VISION • DATA
01 / Technical Highlights & Systems

LLM, Vision & ML Infrastructure

Architecting generative RAG pipelines, foundation vision models (ViT/DINOv3), and high-throughput data engines.

Production RAG & Multimodal vLLM Pipelines

Engineered an enterprise Retrieval-Augmented Generation (RAG) system converting 10,000+ structured XML/HTML documents into vector embeddings, utilizing Google Gemini (Vertex AI) for high-accuracy grounded responses. Developed automated feature extraction pipelines for 100+ multimodal PDFs (text & graphs) using vLLMs and Milvus Vector DB.

Vector Database & LLM ArchitectureMilvus + Vertex AI
Corpus
10,000+ Docs
XML / HTML Embeddings
Multimodal
100+ Complex PDFs
vLLM Graph & Text
GenAI Engine
Gemini Vertex AI
Sub-second Latency

Image Embeddings Analysis

Leveraged Vision Transformers (ViT) and DINOv3 foundation models to generate and index 5M+ image embeddings, identifying data clusters linked to weak model performance and correcting train/test splits.

5M+
Embeddings
DINOv3
Clustering

YOLO-NAS Pose & Active Learning

Reduced YOLO-NAS Pose training time by 25% and boosted pose estimation accuracy from 95% to 97%. Built active learning loops benchmarking label quality between YOLO and DETR models.

Pose Accuracy95% → 97%
Training Speedup-25% Time

Large-Scale ML Infrastructure & Data Engineering

Engineered high-throughput CI/CD pipelines handling >1TB data ingestion. Achieved an 85% batch ingestion speedup for 50M+ synthetic rows. Auto-audited 2D ground-truth labels on 100k+ images and analyzed >100GB clinical EHR time-series on AWS EMR and Apache Spark (+33% F-score).

Ingestion Speedup
85% Faster
50M+ Synthetic Rows
Label Auto-Audit
100k+ Images
Rule-Based QA
EHR Clinical ML
+33% F-Score
Spark & AWS EMR
MLOps & Infra
Docker & K8s
CI/CD Automation
02 / Technical Repertoire

Tools & Stack

From foundation models and RAG pipelines to high-throughput data engineering at terabyte scale.

PyTorch
PyTorch
Google Cloud
Vertex AI / Gemini
Hugging Face
vLLMs / Hugging Face
Apache Spark
Apache Spark
Python
Python
NVIDIA
NVIDIA CUDA
OpenCV
OpenCV
Docker
Docker
Kubernetes
Kubernetes
FastAPI
FastAPI
PostgreSQL
PostgreSQL / Vector DB
Git
Git & CI/CD
PyTorch
PyTorch
Google Cloud
Vertex AI / Gemini
Hugging Face
vLLMs / Hugging Face
Apache Spark
Apache Spark
Python
Python
NVIDIA
NVIDIA CUDA
OpenCV
OpenCV
Docker
Docker
Kubernetes
Kubernetes
FastAPI
FastAPI
PostgreSQL
PostgreSQL / Vector DB
Git
Git & CI/CD

LLMs, RAG & Generative AI

Enterprise RAG Systems (10k+ Docs)Google Gemini (Vertex AI)vLLMs Multimodal PDF ExtractionMilvus Vector DatabaseVector Similarity & EmbeddingsLangChain & Prompt Engineering

Computer Vision & Active Learning

Vision Transformers (ViT & DINOv3)5M+ Image Embedding IndexingYOLO-NAS Pose Estimation (97%)DETR & YOLO Object DetectionActive Learning Label OptimizationData-Centric AI Development

MLOps & Big Data Infrastructure

High-Throughput Pipelines (>1TB Data)Synthetic Data Generation (50M+ Rows)Apache Spark & AWS EMR (>100GB EHR)Automated Label Auditing (100k+ Images)Docker, Kubernetes & CI/CDData Quality & Validation
03 / Career Experience

Work & Experience

Proven track record leading AI/ML systems, computer vision models, and engineering infrastructure.

Acubed by Airbus

Senior AI/ML Engineer

Promoted from ML & Data Engineer (July 2023)
Acubed by AirbusSunnyvale, CA
Sept 2021 — Present
Led a cross-functional team of 3 engineers & 2 analysts; architected an enterprise RAG system embedding 10,000+ documents using Google Gemini (Vertex AI) and built multimodal PDF extraction with vLLMs & Milvus Vector DB.
Leveraged Vision Transformers (ViT) and DINOv3 to generate and index 5 Million image embeddings for performance clustering; optimized YOLO-NAS Pose (-25% training time, 95% → 97% accuracy) and active learning with DETR.
Engineered high-throughput CI/CD data pipelines (>1TB), cut batch ingestion time by 85% for 50M+ synthetic rows, and automated 2D label auditing across 100k+ images.
UCSF

Data Science Intern

UCSFSan Francisco, CA
Jan 2021 — July 2021
Boosted ML model prediction accuracy and F-score by 33% for patient antibiotic administration 2 days prior using feature engineering, importance analysis, and resampling on >100GB EHR time-series data.
Engineered statistical features using Apache Spark and SQL on AWS EMR clusters in collaboration with the Director of Clinical Informatics.
HDR

Engineering Consultant

HDR Inc.San Diego, CA
Aug 2016 — Aug 2020
Delivered water resources engineering solutions focusing on hydraulic flood modeling and hydrological rainfall forecasting.
Developed spatial data analyses and computational simulations to assess environmental impact and optimize watershed infrastructure.
04 / Selected Writings

Articles & Notes

Reflections on AI research, campus visits, and engineering notes.