Di Wu (Woody)
I build and evaluate AI systems. Currently: the validation stage of a legal-AI citation-verification pipeline at Georgia Tech’s Center for Scientific Software Engineering, and a benchmark for multi-agent AI in clinical decision-making at Emory’s Madabhushi Lab. My background is in ML systems, multimodal retrieval, and data engineering.
Current Research
🔍 [Legal AI / LLM Verification] Case Search — Citation Verification (Georgia Tech CSSE)
The validation stage of mellea-lrc, an open-source system that checks whether the case citations in a legal filing are real or AI-fabricated.
- Description: Built the validation stage of an end-to-end pipeline that checks every citation in a legal filing against the public court archive: a groundedness gate that accepts a verdict only when the model quotes the cited page and the quote can be relocated in the retrieved text, and an explicit error-attribution vocabulary that keeps what the system established separate from what it could not determine. 89.1% validation accuracy over 423 citations with zero false positives; citation extraction improved from 88.6% to 94.8% recall at 100% precision.
- Live Demo: mellea-lrc-visualizer.woodygoodenough.com
- Benchmark: false-citation-bench on HuggingFace — 26 real legal documents with annotated AI fabrications and a public evaluation protocol
- Project Page: Details
🏥 [Clinical AI / Agentic Systems] Tumor Board Agents Benchmark (Emory, Madabhushi Lab)
A benchmarking framework for evaluating multi-agent AI systems in tumor-board (cancer-care) clinical decision-making.
- Description: In progress — surveying clinical-AI datasets and agent methods (MIMIC-IV, MIRA) to shape a benchmark that measures how multi-agent clinical decision systems actually behave on real tumor-board tasks.
ML Systems Projects
👗 [Multimodal Retrieval / ML Systems] Fashion Search Demo
An end-to-end cross-modal retrieval demo that turns natural-language fashion queries into relevant product image results.
- Description: Built a production-style retrieval system that combines a fine-tuned CLIP-based model in PyTorch, a Python backend, FastAPI JSON endpoints, and a Next.js frontend. The system uses an ANN index for efficient retrieval over the image corpus and includes interactive model analysis so users can inspect how representation quality affects search behavior.
- Live Demo: fashionsearch.woodygoodenough.com
- Project Page: Details
📐 [Numerical Linear Algebra / Interactive Visualization] Numerical Linear Algebra Explorer
An interactive project for building intuition around the numerical foundations behind machine learning and scientific computing.
- Description: Designed a visualization-focused learning environment for core numerical linear algebra ideas such as sensitivity, conditioning, geometric behavior in 3D, and algorithm progression over time. The project emphasizes intuition-building through interactive numerical analytics and visual demonstrations that connect low-dimensional examples to high-dimensional ML practice.
- Live Demo: nla.woodygoodenough.com
- Project Page: Details
🔬 [Deep Learning / Multimodality] OpenCLIP Fine-Tuning for FashionGen Retrieval
Fine-tuned OpenCLIP variants for large-scale image-text retrieval and comparative multimodal learning analysis.
- Description: Fine-tuned multiple OpenCLIP variants (ViT-B/32, ViT-B/16, SigLIP2) on the FashionGen dataset (image–caption pairs), comparing InfoNCE vs. BCE objectives and studying architectural trade-offs, training dynamics, and caption augmentation for improved multimodal retrieval performance.
- Project Page: Details
- Source Code: GitHub
Data Engineering Projects
📊 [Dashboarding / Data Visualization] Finance Data Analytics — End-to-End Automation
- Description: Automated financial data analytics system covering API data ingestion, ETL pipeline design, feature engineering, aggregation, and analytics-ready data publishing. Built a fully reproducible workflow with scheduled execution, versioned outputs, and downstream visualization support, enabling rapid experimentation and multi-dashboard consumption.
- Demo (Streamlit): Live Demo (AWS EC2 + Cloudflare; offline 2am–7am EST for cost control)
- Project Page: Details
- Source Code: