QA Lead @ Lucid Motors Β· AI QA Architect
I build the missing QA layer for AI systems β evaluation harnesses, guardrails, and validation frameworks that take LLM apps from "demo works" to "production ready." 14+ years in software QA, now focused on making AI testable, measurable, and shippable.
- π LLM evaluation β golden datasets, LLM-as-judge scoring, CI release gates, production drift monitoring
- π€ Agent validation β trajectory scoring, tool-call correctness, multi-agent orchestration testing
- π‘οΈ AI safety β guardrails, adversarial red-teaming, PII protection
- π LLM observability β tracing, cost attribution, prompt-version comparison
- π§ Model lifecycle β fine-tuning, prompt optimization, cost-aware routing
Evaluation & quality
- ai-qa-eval-harness β LLM-as-judge harness: golden datasets, rubric scoring, CI release gates
- rag-eval-toolkit β RAG eval: chunking comparisons, retrieval precision/recall, drift detection
- deepeval-starter β DeepEval framework integration with CI threshold gates
- agent-validation-suite β Validate AI agents: tool-call correctness, plan adherence, release gates
Agents
- ai-agent-lab β ReAct agent built from scratch: tools, trajectory logging
- multi-agent-orchestrator β Planner β worker β critic architecture
- langchain-qa-agent β LangChain tool-calling agent with offline eval
- mcp-server-lab β Model Context Protocol: QA-tools server + client
Observability & ops
- langsmith-evalops β Datasets β experiments β regression detection
- langfuse-observability β LLM tracing, scores, cost tracking, dashboard
- llm-router β Cost/latency-aware routing + semantic caching
Safety
- llm-guardrails β PII redaction, jailbreak defense, policy pipeline
- llm-red-teaming β Prompt-injection attack library + defense scoring
Model lifecycle
- llm-finetune-lab β LoRA fine-tuning with before/after evals
- dspy-prompt-optimizer β Metric-driven prompt optimization, versioned registry
Open to AI QA Architect / Senior SDET / QA Lead roles β Bay Area & remote.