I build AI agents and the tooling that keeps them safe and measurable.
MCP Β· tool-use guardrails Β· agent evaluation Β· open source
A small MCP-first agent framework where every tool comes from an MCP server, with guardrails at the tool layer: dry-run, human approval, argument rules and tool allow-lists.
- Own agent loop with tracing and a step limit; tool errors go back to the model so it can recover
- Works with Gemini and free OpenRouter models, with automatic retry and model fallback
- Ships two example agents: issue triage (GitHub MCP server plus its own MCP server) and a web research agent
- Tested with 47 automated tests running in CI, and honest about what is and isn't verified in its README
Chat-driven API testing: a local LLM agent (LangGraph + Ollama) routes requests, while deterministic pytest checks do the actual testing, so the LLM steers and the assertions stay reliable.
| Project | Merged contribution |
|---|---|
| HelpCode-ai/anythingmcp | #698: added an adapter:new scaffolder so new connector adapters start from a template (closes #585) |
| agentevals-dev/agentevals | #224: CI fix so autofixable ruff violations fail the build instead of passing silently |
- Measuring agent accuracy across models, not just demoing them
- Agent safety: prompt injection and tool permissions
- Fixing bugs and improving tooling in projects I use
Feedback and collaboration welcome, reach me on LinkedIn.
