Production AIthat ships.
Range and depth,
by deliberate design.
Clinical AI
Twenty-one years of clinical, education, and admissions work — anesthesia, psychiatry, special education. Domain depth that informs every system I design.
Agentic AI
Production agents on Claude's tool-use API and MCP. ReAct orchestration, multi-step planning, and evaluation harnesses with adversarial test sets.
Design-engineering
Working prototypes wired to real systems. Iterative bake-offs with Claude and Cursor, with the visual and interaction craft pushed until it ships.
Responsible AI
Pre-declared evaluations with published failures, SHAP interpretability, and programmatic labeling and weak supervision as production patterns — humans in the loop.
Flagship systems first.
A patent-pending fall-detection research prototype, a hackathon-shipped agent-governance platform, citation-grounded RAG, and DOI-published research — each solving a specific problem at the intersection of expert domain knowledge and machine intelligence.
Four questions every flagship system here is built to answer.
Who is allowed to change what, and who approved this change?
Compass-BlackBox IQ
What did it do, why, and can I reconstruct that later?
GUSS
How do I know it is right, and what is the pass condition?
Provenance
What are the parts, and are they swappable?
Governance Drift Researcher
Prime Radiant
Epidemic forecasting · vintage-honest ML · best work to date
Live · dashboard + PyPI · CDC hub registration open
A CDC FluSight forecaster that predicts weekly flu hospital admissions for every US state and territory, 23 calibrated quantiles at a time, built so the numbers can be trusted: the model never sees data dated after the forecast moment — training data is checked out of the hub's git history as it existed on that date — and the code path that submits live is the same one backtested across 55 historical weeks. LightGBM ensembled with a replica of the CDC baseline; in the 2025-26 season backtest it is ahead of UMass-flusion on natural-scale relative WIS (0.609 vs 0.625) and behind it on the log scale, and it prints the two earlier seasons it loses just as plainly. 100% test coverage with the CI gate set there, every workflow gate mutant-tested, and the off-season runs itself: a shadow forecaster that arms automatically the week CDC's data feed wakes, and a watcher that files the go-live checklist the day the new season's config lands.
PREVERA Guardian+AI
Floor 2D LIDAR + cameras · on-device RF-DETR
Public research prototype · Patent pending
A privacy-first fall-detection research prototype for senior-care rooms, on a Jetson Orin Nano. A floor-level 2D LIDAR finds a person lying down from geometry alone, except end-on. For that blind spot, RF-DETR is served on the device by Roboflow Inference. On recorded camera frames the stock model found the person and could not tell lying from standing, so I fine-tuned one on Roboflow; it passed its pre-declared bars on public data, then on recorded room frames from the counter camera, while the floor camera failed one pose. Detection today, prediction next; nothing predicts falls yet. One subject, one room. The failures are published beside the passes.
Governance Drift Researcher
Agent governance · evidence-verified findings · CI gate
Live · PyPI + WeaveMind Cloud
Vendor model lifecycles move; the approved agent inventory doesn't. This scans an Azure AI Foundry catalog against a human-owned baseline and reports the drift under rules that make the report worth trusting: every finding carries its source URI, JSON field path, and content hash, and any finding whose citation no longer resolves is dropped and counted rather than published. Coverage gaps are stated, not implied. It exits nonzero on surviving drift so it works as a CI gate, and emits SARIF for Defender and code scanning. Also ported to Weft and run end-to-end on WeaveMind Cloud, where the human approval gate is structural — there is no API by which the agent can approve its own report.
Triton Kernel Lab
GPU kernels + compiler IR · study artifacts
Public · CI green on every kernel
Triton kernels — elementwise, softmax, matmul — written to be read as much as run. Each one commits its own lowering artifacts: TTIR, TTGIR, LLVM IR, and PTX, so the compiler's choices are inspectable rather than asserted. Benchmarks put the matmul against a hand-written CUDA C++ baseline and cuBLAS, with the correctness gate bounded by cuBLAS's own error against fp64 rather than an arbitrary tolerance, and softmax against torch.compile's steady state with a committed Dynamo explain artifact. The Jetson bring-up log records six failures in the order they happened, including the one an external cold read caught in the repo's own message.
GUSS
Edge inference appliance · headless + autonomous · solo build
Running · public case study
A governed autonomous agent on 7 watts. A hand-assembled, permanently headless Jetson Orin Nano serving a local Qwen3-8B at 12 tok/s that monitors its own health, restarts its own services, paper-trades a benchmarked portfolio inside code-enforced guardrails, and publishes its own dashboard — twice daily, unattended. The build's core lesson, earned across four documented model failures: code decides, the model narrates. Architecture, ADR-style decision log, and an honest build log in the public case study.
Compass-BlackBox IQ
Microsoft Foundry + MCP · solo build
Agents League · AI Skills Fest 2026
A git-backed, Markdown-native blackbox flight recorder and skill auditor for autonomous agents. It governs the layer permissions and orchestration don't — memory and competence. Every action is an immutable decision record; audit heuristics turn uncited decisions into draft skill proposals a human approves. Agents propose; humans promote. Grounds on Microsoft Foundry IQ, kept orthogonal to governed memory.
Provenance
Citation-grounded multimodal RAG · no GPU
Live
Ask a textbook a question; get an answer where every claim is checked against the exact page image that proves it. Retrieval is Cohere Embed v4 over 1,347 page images; answering and verification are Claude vision calls in a LangGraph verify→repair loop. Faithfulness 0.985, recall@5 0.864 — reported as measured, not tuned. Fully API-based and serverless.
Clinical AI Agent
SMART on FHIR + transparent RAG
Live demo
Patient-grounded, citation-traceable clinical decision support on open standards. A 5-agent pipeline fuses a patient's live SMART on FHIR record with cited evidence, emits dual citations (patient data + literature), and scores its own faithfulness with a built-in Clinical Work IQ harness. Built to fill the unoccupied square in a scan of nine frontier healthcare-AI companies: open, transparent, patient-grounded, self-proving.
FetchMerck AI
Published RAG · clinical decision support
DOI 10.57967/hf/8101
An end-to-end RAG pipeline published on Hugging Face with a DOI: ingestion, chunking, embeddings, semantic retrieval, re-ranking, prompt engineering, deployment. LLM-as-a-Judge scoring rubrics across multiple clinical query classes — the evaluation infrastructure is the asset, not the model itself.
Also shipped — earlier work & explorations
Calibrated, falsifiable exam-readiness with a 60-second reliability-diagram check. Gradio on Hugging Face — free tier, so it sleeps when idle and wakes in ~25s.
A LICENSE-style AI-USE.md declaring how AI touched a project — plus a CI checker that fails when the declaration goes missing or stale.
Offline-first EEG neurofeedback for the Mind Media NeXus-10: EDF/EDF+ verifier, montage configs, web dashboard. PyPI + live app.
Interactive in-browser Python data-structures tutorial — real Python executes client-side via Pyodide, zero backend. Live on GitHub Pages.
US chronic disease prevalence — diabetes, obesity, heart disease, inactivity — carried from data prep through interactive visualization. Live dashboard.
A 280K-line third-party WiFi-CSI pose codebase I audited: built the eval gate that exposed the flagship pose metric as unmeasured, and the ADR recording it. Fork · 66 ADRs · archived.
Thin client and runnable examples for the Semantic Scholar Academic Graph, Recommendations, and Datasets APIs.
Local-first macOS menu bar + web dashboard for a self-hosted Monero P2Pool rig. Reads your own node and miner; the wallet address never leaves the machine. BSD-3.
Multi-model deliberation (Karpathy-style, 3-stage) as an MCP server. MIT · PyPI.
scikit-learn Random Forest · RMSE 277.28 · R² 0.9326 · model + Streamlit + Flask.
V-JEPA 2 + Claude analyzing medical procedure videos into teaching content.
30-day readmission prediction · XGBoost · SHAP · 0.76 AUC-ROC.
Predictive maintenance for wind turbines · TensorFlow.
Safety-first helmet detection — VGG-16 transfer learning with threshold analysis tuned for Zero-Harm deployment.
Smart-contract security learning log: CTF writeups, contest findings, and vuln-class notes as they land.
Plain-language guide to spotting AI-made photos, video, and voice — and talking about them with your teen. No tech background needed.
From bedside
to build.
The methodology is the same across every chapter: design the system that lets a real person do something hard, with humans in the loop and the stakes visible.
1// patient context, pulled live via SMART on FHIR2GET /Observation?patient=demo-134[o-a1c] HbA1c = 8.2 % -> flag5[o-egfr] eGFR = 52 -> flag6[o-ldl] LDL = 110 mg/dL
Proof, not
promises.
Responsible by
construction.
Interpretability and honest evaluation aren't a compliance afterthought. They're design constraints I build in from the first commit.
Pre-declared evaluation
Pass bars are committed before the first result, and failures are published beside the passes: PREVERA Guardian+AI's LIDAR prototype reports one subject in one room as exactly that.
Honest evaluation
LLM-as-a-Judge rubrics, a reproducible Clinical Work IQ harness, and scope-and-saturation reporting. Numbers as measured, not tuned.
Interpretability
SHAP attributions, append-only decision records, and dual citations you can verify against the source by eye.
Human-in-the-loop
Agents propose; humans promote. Programmatic labeling and weak supervision with a person in the loop, and revocable memory.
A practitioner's
toolkit.
Verified through a pending patent, published research, and shipped systems — not résumé decoration.
Built in public.
Runnable, reproducible.
Several systems ship as public, runnable artifacts — MCP servers, Hugging Face Spaces, and DOI-published pipelines. Reproducible from the repo and gated by offline tests.
MCP-native
Two MCP servers — llm-council and Compass — callable from Claude Code.
Reproducible evals
Offline CI gates, golden sets, and honest scorecards in every repo.
Live demos
Runnable Hugging Face Spaces and Vercel deployments, not screenshots.
DOI-published
Citable research — DOI 10.57967/hf/8101.
# multi-model deliberation, from Claude Codeclaude mcp add llm-council \--env OPENROUTER_API_KEY=sk-or-... \-- llm-council-mcp> council_deliberate("ship v1 Friday?")# 3 YES / 1 NO -> chairman synthesis
A non-linear path with through-line.
Every role taught something that compounds in the AI work: clinical reasoning under pressure, mixed-methods research, multi-stakeholder delivery, and consultative communication with expert and non-expert audiences alike.
Technical Founder · PREVERA Guardian+AI
Independent · Remote
Shipped AI applications zero-to-one, with one patent pending. Full ownership: research, architecture, deployment, evaluation, patent prosecution. Co-founder and technical lead of The Gerontechnology Group, coaching two domain-expert collaborators new to coding. Volunteer coach and mentor for CodePath engineering students. Most recently: PREVERA Guardian's public LIDAR fall-detection prototype (v0.1.0, pre-declared evaluations, RF-DETR fine-tuned on Roboflow and served on-device on a Jetson Orin Nano) and Prime Radiant. Earlier: Governance Drift Researcher (published to PyPI with trusted publishing and a clean-room wheel test), Triton Kernel Lab, and Compass-BlackBox IQ (Agents League @ AI Skills Fest 2026).
Special Education Teacher
Hawai'i Department of Education
Supervised 4–6 educational assistants and 2–3 paraprofessionals each year and wrote their performance evaluations, backed by student progress data. Designed data-driven experimental interventions for 20+ students per year — mixed-methods research with iterative calibration under tight timelines and real stakes. Coordinated cross-functional IEP teams and authored 40+ federally compliant IEPs per year.
Associate Director of Admission & Adjunct Faculty
Hawai'i Pacific University
Consultative delivery at scale: understood objectives, designed academic solutions across programs, navigated multi-stakeholder decisions, and drove outcomes against quarterly targets. Created a new admission counselor position, then hired, onboarded, and wrote the SOPs for it. Taught six undergraduate psychology courses as adjunct faculty.
Substitute Special Education Teacher
Hawai'i Department of Education
First five-year stretch of IEP portfolio delivery across diverse classroom contexts, coordinating cross-functional teams under legally binding targets.
Anesthesia & Psychiatric Technician
Queen's Medical Center · Level I Trauma Center
Eight years of frontline clinical reasoning under uncertainty — perioperative workflows, malignant hyperthermia response, psychiatric stabilization. The domain foundation under every AI system I now build.
The training behind the work.
Graduate research methodology, frontier AI training, and cloud fundamentals — the foundation for every system I ship.
Post Graduate Program in AI/ML
Great Learning · McCombs School of Business · GPA 3.95
Anthropic Academy — 13/13 Courses
Claude API, prompt engineering, agentic systems
AI Developer Path — Agents 101 & 201
Agent architectures on accelerated inference
Azure AZ-900 & AI-900
Cloud fundamentals · ML, computer vision, NLP, generative AI
Advanced MLOps & SQL Analytics
Machine learning operations and SQL analytics badges
Patent Pending — PREVERA Guardian+AI
Aspects of the system, including a proprietary V-JEPA verification stage
PhD ABD — Integrative Mind-Body Medicine
Classic Grounded Theory methodology
MS Counseling Psychology · BA Psychology
BA Magna Cum Laude
Biofeedback Certification
Biofeedback Certification International Alliance · applied psychophysiology
Let's build something
that matters.
Open to roles at frontier AI labs and healthcare AI companies — remote, or relocating to SF or Seattle. Also open to advisory and consulting at the intersection of clinical AI, evaluation, and human-in-the-loop product design.
Based in Seattle, WA · Honolulu & SF · UTC−7