I build the trust layer around AI agents.
Memory with provenance · independent outcome checks · executable claims · behavior compatibility
NexTool · Product docs · npm · Contact
Most agent systems make it easy to call a tool. My open-source work focuses on the harder questions that come next: Why did the agent decide this? Did the outside world actually change? Is the public claim supported? Did an upgrade silently alter behavior? Did a critical rule remain true?
| Failure to prevent | Project | Public evidence |
|---|---|---|
| Lost context and provenance | Soul MCP | npm · 23 MCP tools · 373 tests |
| Tool success without real outcome | Postcondition | npm · MCP + SDK + CLI · 53 tests |
| Unsupported README or release claims | Proofspec | source release · five reports · 70 tests |
| Silent behavior drift after an upgrade | Behaviorlock | source release · nine matchers · 73 tests |
| Critical runtime conditions going unproved | Agent Invariants | source release · fail-closed policies · 36 tests |
| Opaque, disposable cognitive state | ANIMA Kernel | source release · zero runtime dependencies · 446 tests |
intent
↓
Soul ── remembers decisions, provenance, corrections, and outcomes
↓
Agent Invariants ── checks what must remain true during execution
↓
Postcondition ── observes whether the intended world state exists
↓
Proofspec ── binds public claims to narrow evidence contracts
Behaviorlock ── compares observable behavior before and after upgrades
Soul is the long-running flagship (on npm since February 2026, seven releases). Postcondition, Proofspec, Behaviorlock and Agent Invariants are young v0.1 releases from July 2026: small, tested, and early.
These projects can work together, but none pretends to be the whole agent
stack. Each has a narrow promise, tests for that promise, a security or evidence
boundary, and an honest unknown state where the software cannot prove enough.
Local-first agent memory with durable runs, receipts, episodes, conflicts, calibration, deliberation, and a governed skill registry.
claude mcp add soul -- npx -y soul-mcpProduct guide · Source · npm
Declare an intended result, observe constrained file, HTTP, Git, npm, or manual
evidence, and retain an honestly classified receipt instead of trusting
success: true.
npx -y postcondition-mcp serveProduct guide · Source · npm
Turn README, product, benchmark, and release claims into versioned evidence contracts with dependencies, policies, receipts, CI gates, and JSON, Markdown, HTML, SARIF, and JUnit reports.
Product guide · Source · v0.1.0 release
Compare portable before-and-after traces with deterministic contracts for tool sequences, permissions, output structure, outcomes, ranks, sets, and bounded metrics — without provider calls or an LLM judge.
Product guide · Source · v0.1.0 release
I do not use fake adoption numbers or turn test counts into a claim of product quality. These are narrow, inspectable facts about the current public surfaces:
| Surface | Current public fact |
|---|---|
| Soul MCP | 373 automated tests; CI on Node 20, 22, and 24 |
| Postcondition | 53 automated tests; clean-install and MCP handshake checks |
| Proofspec | 70 automated tests; 98.05% line coverage; five report formats |
| Behaviorlock | 73 automated tests; 99.06% line coverage; nine deterministic matchers |
| Agent Invariants | 36 automated tests; source release and CI |
| ANIMA Kernel | 446 automated tests; zero runtime dependencies |
| NexTool | 269 browser-tool pages; 131 technical guides; static site audit in CI |
Passing tests establish only what they cover. Hash chains are not signatures. Portable traces are not provenance attestations. A green compatibility gate covers only declared scenarios. Those boundaries live beside the product claims, not in fine print after them.
I also run Miguel, a private personal AI operating layer that connects memory, voice, multiple model providers, tools, and long-running local services. It is the laboratory from which reusable mechanisms are extracted — not a 50,000-line personal monolith I intend to dump online. Private conversations, identity state, credentials, machine paths, and raw session history stay private.
The public products are rebuilt as small, documented, testable systems with portable schemas and explicit security boundaries.
I am a self-taught developer in Vienna. I have a full-time job and build these projects in the evenings and on weekends, without funding so far. In earlier jobs I dispatched a fleet of 50+ vehicles; I bring the same habit to software: deliver what was promised and say plainly what is not done yet. No degree yet; I plan to start university (business informatics) in autumn 2027.
I use capable models as research, coding, review, and testing collaborators. I choose the architecture, resolve conflicts, validate results, manage releases, and remain responsible for every public claim.
My default loop is:
define the failure → design the evidence → build → falsify → ship → observe
Inspect the source. Run the tests. Judge the boundary.

