Skip to content
View sureshbujji's full-sized avatar

Block or report sureshbujji

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
sureshbujji/README.md

Hi, I'm Suresh Itha πŸ‘‹

QA Lead @ Lucid Motors Β· AI QA Architect

I build the missing QA layer for AI systems β€” evaluation harnesses, guardrails, and validation frameworks that take LLM apps from "demo works" to "production ready." 14+ years in software QA, now focused on making AI testable, measurable, and shippable.

What I work on

  • πŸ” LLM evaluation β€” golden datasets, LLM-as-judge scoring, CI release gates, production drift monitoring
  • πŸ€– Agent validation β€” trajectory scoring, tool-call correctness, multi-agent orchestration testing
  • πŸ›‘οΈ AI safety β€” guardrails, adversarial red-teaming, PII protection
  • πŸ“Š LLM observability β€” tracing, cost attribution, prompt-version comparison
  • πŸ”§ Model lifecycle β€” fine-tuning, prompt optimization, cost-aware routing

Featured projects

Evaluation & quality

  • ai-qa-eval-harness β€” LLM-as-judge harness: golden datasets, rubric scoring, CI release gates
  • rag-eval-toolkit β€” RAG eval: chunking comparisons, retrieval precision/recall, drift detection
  • deepeval-starter β€” DeepEval framework integration with CI threshold gates
  • agent-validation-suite β€” Validate AI agents: tool-call correctness, plan adherence, release gates

Agents

Observability & ops

Safety

  • llm-guardrails β€” PII redaction, jailbreak defense, policy pipeline
  • llm-red-teaming β€” Prompt-injection attack library + defense scoring

Model lifecycle

Stack

Python LangChain PyTorch GitHub Actions Docker

GitHub stats

Suresh's GitHub stats Top Langs

Connect

LinkedIn

Open to AI QA Architect / Senior SDET / QA Lead roles β€” Bay Area & remote.

Popular repositories Loading

  1. DataScienceAndAI DataScienceAndAI Public

  2. ai-qa-eval-harness ai-qa-eval-harness Public

    AI QA evaluation harness: LLM-as-judge scoring, golden datasets, CI release gates, production drift monitoring

    Python

  3. rag-eval-toolkit rag-eval-toolkit Public

    RAG pipeline evaluation: chunking strategy comparison, retrieval precision/recall, embedding-drift detection

    Python

  4. ai-agent-lab ai-agent-lab Public

    Build an AI agent from scratch: ReAct loop, tools, trajectory logging, bug-triage demo

    Python

  5. agent-validation-suite agent-validation-suite Public

    Validate AI agents: tool-call correctness, plan adherence, loop detection, release gates

    Python

  6. multi-agent-orchestrator multi-agent-orchestrator Public

    Multi-agent architecture: planner, worker, critic with bounded rework and blackboard state

    Python