I build production MCP servers β the kind where every write is previewed before it happens, confirmed when it matters, and recorded in a chain you can verify.
Not as a demo. I run one over live invoices, client prices and statutory filing deadlines every working day, which is how I learned most of what's below the hard way.
π housewarden β a guarded household-operations MCP server
31 tools. Reads are free; every mutating tool goes through one guard β a dry-run preview of exactly what will change, a human ask when the action is risky, execute-exactly-once, then an append to a hash-chained audit log. The web console uses the same path, so an assistant can never reach a weaker one than a person can.
npm run e2e drives it with a real MCP client and prints a pass/fail table, so you don't have to take my word for any of it. Streamable HTTP Β· self-hosted Β· MIT
Three-minute demo β real console, real MCP client, nothing mocked.
Because running these in production surfaces problems nobody has tools for yet.
| π‘οΈ mcp-shield | Security scanner for MCP servers β finds what your tool surface exposes before someone else does |
| π§ͺ mcp-testkit | Testing framework for MCP servers Β· npm |
| πͺ mcpgate | Open-source MCP gateway β reverse proxy for MCP servers |
| π mcp-audit | Python security scanner for MCP servers |
| π shopify-mcp | MCP server for the Shopify Admin API |
Every 21 minutes: my self-healing monitor took my business phone line down for two days
A repair loop that couldn't tell broken from a human is part-way through fixing it, a monitor that trusted its own cache over the service, and active (running) answering a question I wasn't asking. Four bugs, one shape: a system confidently answering something slightly different from what was asked.
Tool surface reviews, production builds, and keeping them running afterwards.
Scope and fixed prices: The Write Path
π¬ support@bizfilo.com
LLM developer tooling β llm-cost-profiler (spend visibility in two lines, pip install llm-spend-profiler), llm-bench (race providers in the terminal), ai-stability (measure output consistency, pipx install ai-stability).
Python TypeScript MCP Claude Code Anthropic API OpenAI API CLI tooling
