Selected work · 8 projects

Every project, with the case study behind it.

Agent platforms, research artifacts, retrieval systems, clinical AI and neuroimaging. 3 have public repositories, 2 are on arXiv.

The shape of the work

8 projects. 5 domains. One case study each.

3
Agent platforms
2
Research artifacts
1
Retrieval systems
1
Clinical AI
1
Neuroscience

Filter by domain or theme

Domain
Theme

8 of 8 projects shown.

Platform · Human-in-the-loop · MCP

01

Zasti Agentic Automation Platform

A shared agent harness (LangGraph through MCP) with four automation pipelines built on it — sales, resource tracking, marketing, and internal software engineering — adopted company-wide including leadership, with a human approval gate before any consequential action.

Zasti Inc · Design and implementation

4
verticals automated
1
shared agent harness
10
person company, adopted end to end
0
consequential actions without human approval
LangGraphMCPTask-dependency planningValidation gatesHuman approval gates
Case study

Open source · Agent safety · MCP

02

Unwind

A reversibility layer for agentic tool use. A transparent MCP proxy that works out which agent actions can be taken back, quietly reverses the ones that go wrong, and interrupts you only for the ones that genuinely can't be undone.

Sole author · v0.1.0 · early beta · public since Aug 2026

R0–R4
reversibility classes
3
distribution channels
v0.1.0
early beta release
Apache-2.0
license
PythonTypeScriptMCPPyPI & npm packagingDockerGitHub Actions CI+1
Case studysourcedocs

Research artifact · arXiv:2608.14639

03

verifydoc — per-field selective risk control

Reference implementation for “Valid Per-Field Selective Risk Control for Document Extraction: Three Failure Modes, a Validity Ladder, and When Conditioning Pays.” It shows that the natural per-field accept/review procedure silently violates its own selective-risk guarantee on real documents.

Sole author · Submitted 28 July 2026 · announced August 2026 · arXiv preprint · 14 pages

13,859
fields evaluated
800
CORD receipts
3
failure modes diagnosed
49.0%
base field accuracy
PythonSelective risk controlCalibrationclaude-sonnet-5CORD datasetReproducible benchmark harness
Case studysourcepaper

Research artifact · arXiv:2604.16706

04

agenthallu-bench — auditing agent evaluation

Code and data for “Auditing Automated Evaluation, Error Propagation, and Runtime Mitigation in Tool-Using Language Agents.” It audits the widely-assumed reliability of automated evaluation for tool-using LLM agents against human annotation.

Sole author · v1 April 2026 · revised 9 August 2026 · arXiv preprint · 11 pages, 4 figures, 8 tables

14,750
execution traces
13
LLM agents audited
4
task domains
9 + 4
proprietary + open-weight models
PythonLLM agent evaluationExecution-trace analysisHuman annotationBenchmark harness design
Case studysourcepaper

Agentic AI · RLHF

05

Autonomous Multi-Agent Research Platform

A Planner–Retriever–Reasoner–Validator agent loop with persistent memory and tool use across PubMed, ArXiv, and Semantic Scholar APIs.

+41%
synthesis quality
−68%
hallucinated citations
5K+
concurrent sessions
<200ms
median latency
PythonLangGraphGPT-4oPPOFAISSNeo4j+4
Case study

Clinical AI · Knowledge Graphs

06

Multimodal Clinical AI Assistant

GPT-4V vision, FHIR parsing, Neo4j patient graphs, and vector search combined for knowledge-graph patient profiling.

94%
retrieval accuracy
15+
biomarker categories
87%
anomaly sensitivity
3
pilot healthcare teams
PythonGPT-4VLangChainFHIRNeo4jChromaDB+4
Case study

Real-Time Retrieval · Source Verification

07

Conversational AI Search Engine

QLoRA-fine-tuned Llama 3.1 with hybrid dense + sparse retrieval and automated citation and fact-verification.

−40%
misinformation
150ms
P95 time-to-first-token
1K+
concurrent users
+35%
NPS
PythonLlama 3.1QLoRAKafkaElasticsearchQdrant+3
Case study

M.S. Thesis · Biomedical AI

08

EEG–fMRI Fusion for Neural Decoding

A multimodal neuroimaging framework predicting visual stimuli by fusing 70-channel EEG with 3T fMRI (Wakeman–Henson dataset) using temporal convolutional networks.

University of Cincinnati · M.S. thesis · advisor Prof. Vikram Ravindra · 2022 — 2024

84.8%
within-subject accuracy
81.1%
cross-subject (LOSO)
0.93
ROC-AUC (within)
0.90
ROC-AUC (cross)
PythonPyTorchTemporal CNNsEEG/fMRI preprocessingMNENilearn
Case study

45 distinct tools across the 8

  • LangGraph
  • MCP
  • Task-dependency planning
  • Validation gates
  • Human approval gates
  • Python
  • TypeScript
  • PyPI & npm packaging
  • Docker
  • GitHub Actions CI
  • OpenSSF Scorecard
  • Selective risk control
  • Calibration
  • claude-sonnet-5
  • CORD dataset
  • Reproducible benchmark harness
  • LLM agent evaluation
  • Execution-trace analysis
  • Human annotation
  • Benchmark harness design
  • GPT-4o
  • PPO
  • FAISS
  • Neo4j
  • FastAPI
  • PostgreSQL
  • Redis
  • Kubernetes
  • GPT-4V
  • LangChain
  • FHIR
  • ChromaDB
  • Streamlit
  • AWS EKS
  • Llama 3.1
  • QLoRA
  • Kafka
  • Elasticsearch
  • Qdrant
  • React
  • PyTorch
  • Temporal CNNs
  • EEG/fMRI preprocessing
  • MNE
  • Nilearn

Keep reading

The research behind these →Experience & teaching →