A few projects that show how I build and evaluate AI and data systems — with an emphasis on measurement, correctness, and responsible data handling.
A reusable, config-driven framework for turning any course's materials into an evaluated, privacy-safe RAG assistant — with human-in-the-loop feedback tooling and an evaluation loop. Proven end-to-end on a university capstone.
RAGLLM evaluationprivacy-by-designPython
Monte-Carlo existence proofs of how a single aggregate metric can mislead across populations — a literal Simpson's paradox, throughput-driven metric deflation, and non-uniform optimal effort allocation, with symbolic verification.
measurementstatisticsevaluationPython
A Windows desktop app that automates the repetitive Salesforce/Outlook workflow of running a course caseload — batch actions across filtered students, templated email/text, at-rest encryption, a test suite, and a packaged build.
applied softwarePythondesktopautomation
Open-source research software for the creation, manipulation, and display of circle packings (in the sense of Thurston) — tied to my published mathematics on generalized branching.
math researchJavaopen source