Skip to content

Latest commit

 

History

65 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Smoke Monkey Harness

Build production-ready AI agents and coding agents in TypeScript.

An embeddable, framework-agnostic agent runtime for building AI coding assistants, autonomous developer tools, desktop agents, and MCP-powered applications — with zero runtime dependencies.

Docs & Live Demo NPM Downloads npm version GitHub Stars MCP Ready license

Smoke Monkey Harness — TypeScript agent runtime, MCP client, SKILL.md skills, permissions, and context management

Smoke Monkey Harness is not an AI model. It is the runtime that turns an LLM into an agent capable of planning, calling tools, editing files, interacting with MCP servers, managing context, asking for permission, recovering from failures, and resuming work.

No NestJS. No database. Zero dependencies. Just the agent runtime.

docs & live simulator: https://smoke-monkey-harness.vercel.app/ · npm: @smoke-monkey/harness · dependencies: 0 · license: MIT · CI: passing · types: TypeScript

🌐 Interactive Documentation & Live Simulator: Visit https://smoke-monkey-harness.vercel.app/ to explore interactive quickstarts, test the agent loop live in your browser, view the 6-phase state machine architecture, and try the drop-in React chat UI.


Install

pnpm add @smoke-monkey/harness
# or: npm install @smoke-monkey/harness
# or: yarn add smoke-monkey-harness

GitHub Packages (same package, scoped):

# .npmrc
@rajdeepdevelopment:registry=https://npm.pkg.github.com
//npm.pkg.github.com/:_authToken=GITHUB_PAT
pnpm add @rajdeepdevelopment/smoke-monkey-harness

GitHub Packages requires auth even for public packages — create a fine-grained PAT with read:packages permission on this repository.


Quickstart

import { createAgent } from '@smoke-monkey/harness';

const agent = createAgent({
  provider: 'nvidia',
  model: 'nvidia/nemotron-3-super-120b-a12b',
  apiKey: process.env.NVIDIA_API_KEY, // or pass a resolver: (provider) => key
  workspacePath: process.cwd(),
});

const result = await agent.run('Refactor the auth middleware to use JWTs, then run its tests.');
console.log(result.status);

That's it. Smoke Monkey handles the agent loop, tool execution, planning, verification, recovery, and context management — no framework, no database.

Add human-in-the-loop permissions

Route permission prompts and questions to your UI (or set autoApprove: true):

agent.on('permission.required', (e) => {
  const { toolCallId, toolName } = e.data;
  agent.resolvePermission(toolCallId, /* allow | deny */ 'allow');
});

agent.on('ask_user.required', (e) => {
  agent.respond(e.data.toolCallId, await promptUser(e.data.question));
});

Any OpenAI-compatible endpoint works — openrouter, gemini, xai, … or run locally with Ollama:

const agent = createAgent({ provider: 'ollama', model: 'qwen3:8b', workspacePath: process.cwd() });

Package names

New installs should use the scoped names below. The older unscoped smoke-monkey-harness and smoke-monkey-harness-mcp still work and are still published — they are kept for existing installs, not deprecated.

Package Use this Replaces Docs
Agent runtime @smoke-monkey/harness smoke-monkey-harness Docs & Live Simulator
MCP server @smoke-monkey/mcp smoke-monkey-harness-mcp MCP Server Guide
Chat UI @smoke-monkey/ui — Chat UI Showcase

Use it over MCP (connect this repository)

Smoke Monkey ships its own stdio MCP server that exposes the harness itself to any MCP client — so you can build a full agentic, looping workflow in seconds without writing any library code yourself. Point Claude Code, Codex, opencode, Cursor, or any MCP-capable editor at the server and your agent can plan (harness_plan), scaffold (harness_scaffold), wire the loop (harness_guide / harness_api), verify (harness_verify), and apply the 25 bundled engineering skills category-wise (harness_skills_by_category / harness_skill_content) — directly through MCP.

💡 Interactive Guide: See the MCP server documentation & architecture walkthrough.

No install in your project needed:

{
  "mcpServers": {
    "smoke-monkey-harness": {
      "command": "npx",
      "args": ["-y", "@smoke-monkey/mcp"]
    }
  }
}

The harness toolset appears directly — the looping agent workflow (run loop, tool execution, permissions, context management, MCP integration, sessions, recovery) is built into the library and driven by these tools. Inside your own harness, register it like any other MCP server:

const agent = createAgent({
  provider: 'nvidia',
  model: 'nvidia/nemotron-3-super-120b-a12b',
  apiKey: process.env.NVIDIA_API_KEY,
  workspacePath: process.cwd(),
  mcp: [
    {
      id: 'smoke-monkey',
      name: '@smoke-monkey/mcp',
      description: 'Build agents on Smoke Monkey',
      command: 'npx',
      args: ['-y', '@smoke-monkey/mcp'],
      enabled: true,
    },
  ],
});

The npx command works on npm and the GitHub Package registry alike. The MCP server (in plugin/mcp/) is dependency-free and speaks JSON-RPC 2.0 over stdio, so it connects anywhere MCP stdio servers work.


Why an AI Agent Harness? (Harness Engineering)

"Harness engineering is the discipline of creating deterministic, observable, and safely bounded runtimes around non-deterministic LLMs." — Martin Fowler, Software Architecture

Raw prompt chaining and naive while loops fail when building production autonomous agents: they hallucinate tool calls, exhaust token budgets, get trapped in infinite loops, and execute destructive operations without guardrails.

Smoke Monkey Harness is an enterprise-grade agent harness that bridges unstructured LLM outputs with deterministic execution guarantees:

  1. Deterministic 6-Phase State Machine: Replaces fragile ReAct loops with structured phases (Plan → Tool Call → Execute → Verify → Compact → Recover).
  2. Zero-Dependency Runtime: Completely standalone TypeScript with 0 external dependencies. Runs seamlessly in Node.js, Electron desktop apps, CLI tools, VS Code/Cursor extensions, or serverless workers.
  3. First-Class Human-in-the-Loop: Asynchronous pause-and-resume mechanisms for command approval, interactive questions, and token compaction thresholds.
  4. Native MCP (Model Context Protocol): Expose tools and connect to existing MCP ecosystem tools without protocol wrappers.

Framework Benchmark Comparison

Feature / Architecture Smoke Monkey Harness LangChain / LangGraph CrewAI Mastra From Scratch
Runtime Dependencies 0 (Zero) 50+ packages 30+ packages 20+ packages 0
Agent Execution Loop Deterministic 6-Phase Machine Graph / StateGraph Role Agents Workflow Graph Fragile while(true)
Human-in-the-Loop Pauses Native Async Event Bus Complex Checkpointers Limited Partial Custom implementation
Model Context Protocol (MCP) Built-in Client + Server Plugin / External No native support Wrapper Manual JSON-RPC
Context Window Compaction Automatic Token Budgeting Manual message trims Context window errors KV / Vector sync Unhandled context overflow
Engineering Skills System 25+ Universal SKILL.md Prompt templates Role definitions Action tools Raw system prompts
Embeddability CLI, Desktop, Extension, Web Node / Python heavy Python runtime Node.js backend Custom
Database Requirement None (In-Memory or File) Vector / SQL DB SQLite / ChromaDB PostgreSQL Any

Core capabilities

🤖 Agent Runtime Multi-step planning through a phase machine (explore → plan → edit → verify → recover → complete), with step / no-progress / runaway / spin guards, finish_task detection, interruption and resume.

🛠️ 24 Built-in Tools Filesystem, terminal, search, Git, and agent-management tools in five groups. Disable groups with tools or register your own.

Group Tools
filesystem read_file, write_file, edit_file, line_edit, replace_lines, apply_patch, delete_file, list_directory, inspect
terminal run_command, run_test
search glob, grep
git git_status, git_diff, git_log
agent ask_user, context_manage, todo_write, finish_task, list_skills, use_skill

🔌 MCP Native Connect local stdio or remote Streamable HTTP MCP servers (<server>__<tool> tool names, lazy connect, close at run end). A curated stock catalog (flattenStock / stockToMcpConfig) provisions well-known servers. Discovery is opt-in (mcpStockSearch: true): inspect_mcp_stock then returns a compact ranked inventory and never stops the run — the model decides for itself, and only request_mcp_approval pauses for the user. The stock-search guidance is omitted from the system prompt when the tool is off.

🧠 Skills Claude Code / Codex / AniGravity / opencode-style SKILL.md folders, loaded just-in-time: the system prompt carries only a one-line catalog; the model pulls full instructions with use_skill when the task matches — no context bloat from skills that don't apply. Point anywhere with skillsDir, or let it scan .opencode/skills, .claude/skills, .codex/skills in the workspace and home dirs.

🔐 Permissions allow-all / deny-all / ask-default, or your own resolver:

permissions: ({ toolName, args }) => {
  // your policy
  return 'allow'; // 'deny' | 'ask'
};

📦 Context Management Automatic compaction summarizes the conversation when it crosses the token budget, so long-running tasks keep going without blowing the context window.

💾 Resumable Sessions Continue work across runs with a persistent sessionId:

const agent = createAgent({ /* ... */, sessionId: 'project-123' })

🌐 Multi-provider LLMs openai, openrouter, nvidia, xai, gemini, opencode, omniroute, ollama (REST + SSE streaming), or any OpenAI-compatible endpoint via baseUrl override.


💬 Chat UI — @smoke-monkey/ui A published browser package built for this harness: streaming markdown, tool cards with icons and live progress, inline pause dialogs, artifacts, sources, and charts. Drop-in (ChatPanel, SmokeMonkeyChat) or headless (ChatRuntime). A help chat, a support widget, or a whole agent app.

Built for

  • AI coding assistants
  • Autonomous code editors
  • Desktop AI applications
  • Developer copilots
  • MCP-powered agents
  • Internal engineering agents
  • Agentic automation tools
  • Research and experimentation platforms

Architecture

┌──────────────────┐
│ Your Application  │
└────────┬─────────┘
         ▼
┌──────────────────┐
│  Agent Harness   │
└────────┬─────────┘
         ▼
┌──────────────────┐
│    Agent Loop    │
└────────┬─────────┘
         ▼
 Tools · MCP · Skills · Permissions · Context · Sessions · Events
         │
   ┌─────┴─────┐
   ▼           ▼
 LLM        MCP Servers
Providers   + Custom Tools

Visual Architecture Overviews

1. Trio Architecture: Core Runtime, MCP Protocol, and React UI Bridge

Smoke Monkey Trio Architecture: Core Runtime, Model Context Protocol, and UI Bridge

2. The 6-Phase Deterministic Execution Loop

Smoke Monkey 6-Phase Agent Loop: Plan, Tool Call, Execute, Verify, Compact, Recover

3. Interactive Pauses & Human-in-the-Loop Permissions

Interactive Pauses and Human-in-the-Loop Permissions
---

Connect the chat UI

@smoke-monkey/harness (Node) and @smoke-monkey/ui (browser) ship as separate packages and share no interface, so something has to translate. createHarnessBridge is that translation, and it ships with the UI.

💡 Try it live: Test the streaming chat UI, tool call execution cards, charts, and reasoning stream in the Live Simulator.

// server
import { createHarnessBridge } from '@smoke-monkey/ui';

const bridge = createHarnessBridge({ agent, messageId });
for await (const event of bridge.events()) socket.send(JSON.stringify(event));

// the two paused-run answers, coming back from the browser
socket.on('message', (raw) => {
  const { type, data } = JSON.parse(raw);
  if (type === 'resolve_ask_user') {
    bridge.answer({ toolCallId: data.toolCallId, kind: 'ask', answer: data.response });
  } else if (type === 'resolve_permission') {
    bridge.answer({ toolCallId: data.toolCallId, kind: 'permission', answer: data.decision });
  }
});
// browser
import { SmokeMonkeyChat, WebSocketTransport } from '@smoke-monkey/ui';
import '@smoke-monkey/ui/ui.css';

<SmokeMonkeyChat
  transport={new WebSocketTransport({ url: 'wss://api.example.com/ws' })}
  toolPresentations={agent.getToolPresentations()}
/>

permission.required and ask_user.required pause the run. Nothing resolves them on their own. Render them and route the answer back, or the run deadlocks with nothing logged — it looks like a hung request, not a bug.

A runnable example lives in examples/chat-demo — it consumes both packages from npm. The full mapping table, the component-by-component breakdown, and a wiring checklist are in ui/README.md and in the MCP server's harness_guide_ui_bridge_and_components().

MCP configuration

Connect your agent to the outside world. Smoke Monkey supports MCP servers over stdio and Streamable HTTP. Servers connect lazily and can require explicit user approval before activation.

import { createAgent, stockToMcpConfig, findStockEntry } from '@smoke-monkey/harness';

const agent = createAgent({
  provider: 'nvidia',
  model: 'nvidia/nemotron-3-super-120b-a12b',
  apiKey: process.env.NVIDIA_API_KEY,
  workspacePath: process.cwd(),
  mcp: [
    // stdio server:
    {
      id: 'filesystem',
      name: 'filesystem-mcp',
      description: 'Local filesystem tools',
      command: 'npx',
      args: ['-y', '@modelcontextprotocol/server-filesystem', '/tmp'],
      enabled: true,
    },
    // streamable-HTTP server:
    {
      id: 'miro',
      name: 'miro',
      description: 'Miro boards',
      url: 'https://mcp.miro.com/',
      headers: { authorization: 'Bearer ' + process.env.MIRO_TOKEN },
      enabled: false, // disabled servers need user approval before use
    },
    // ...or pull a config from the stock catalog:
    stockToMcpConfig(findStockEntry('playwright-mcp')!),
    // bundled Agent Skills (25 engineering SKILL.md bundles shipped in the package):
    stockToMcpConfig(findStockEntry('agent-skills-backend')!),
    // ...or add your own server from the desktop-style stock catalog:
    stockToMcpConfig(findStockEntry('agent-skills-qa')!),
  ],
});

Supported connection types

  • Local stdio servers
  • Remote Streamable HTTP servers
  • Authentication headers
  • Lazy connection
  • Approval gates
  • Runtime server management (addMcpServer / removeMcpServer / listMcpServers)
  • Stock server catalog

To expose the Smoke Monkey harness itself as an MCP server, see Use it over MCP.


Skills

Reusable instruction bundles in the same SKILL.md folder format used by Claude Code, Codex, AniGravity, and opencode — loaded just-in-time so the system prompt stays small.

const agent = createAgent({
  provider: 'nvidia',
  model: 'nvidia/nemotron-3-super-120b-a12b',
  apiKey: process.env.NVIDIA_API_KEY,
  workspacePath: process.cwd(),
  skillsDir: ['examples/skills'], // scans for <dir>/<skill>/SKILL.md + <dir>/<skill>.md
  autoApprove: true,
});
---
name: Commit Message
description: Write conventional, concise git commit messages for the uncommitted changes.
---

# Conventional Commit Message

…instructions the agent follows while the task matches…

Discovery defaults to .opencode/skills, .claude/skills, .codex/skills under the workspace plus ~/.claude/skills, ~/.codex/skills, ~/.opencode/skills, ~/.config/opencode/skills. The tools list_skills (browse catalog) and use_skill (load instructions) register automatically when at least one skill is present. Lower-level pieces: loadSkillsFromDirs(), defaultSkillDirs(), and SkillRegistry (all exported from the package root).

Bundled Agent Skills (category-wise)

The package ships the 25 agent-skills engineering skills (SKILL.md folders in plugin/agent-skills/skills) — loaded directly as harness skills, without spawning the MCP server. Categories mirror the stock MCP entries, so a run can register just the backend (or frontend/devops/qa) skill set:

import { loadAgentSkills, buildAgentSkillRegistry, AGENT_SKILL_CATEGORIES } from '@smoke-monkey/harness';

// catalog of the 4 categories: agent-skills-backend / -frontend / -devops / -qa
console.log(AGENT_SKILL_CATEGORIES.map((c) => `${c.id} → ${c.domain}`));

// all 25 skills, each tagged with its catalog `domains` and lifecycle `phase`
const all = loadAgentSkills();

// just the backend domain set (api-and-interface-design, tdd, security, …)
const backendSkills = loadAgentSkills({ category: 'agent-skills-backend' });

// or a ready-to-use SkillRegistry for one category
const qa = buildAgentSkillRegistry({ category: 'agent-skills-qa' });

Each loaded skill carries domains: string[] and phase: 'define' | 'plan' | 'build' | 'verify' | 'review' | 'ship' | 'meta' metadata. The phase/domain mapping is shared with the bundled MCP server via plugin/agent-skills/catalog.json, so list-wise behavior never drifts between the skill loader and the server.


API

createAgent(options) → AgentHarness

  • agent.run(task, opts?) — run the agent to completion, pausing on questions / permission prompts.
  • agent.respond(toolCallId, text) — answer a pending ask_user.
  • agent.getToolPresentations() — host-provided { name: Presentation } map for custom tools (pass to UI before first events)
  • agent.resolvePermission(toolCallId, 'allow' | 'deny') — resolve a permission.required pause.
  • agent.resolveMcpDecision(toolCallId, { action: 'enable' | 'add' | 'skip', names }) — resolve an mcp.approval_required pause.
  • agent.addMcpServer(config) / agent.removeMcpServer(id) / agent.listMcpServers() — manage MCP servers at runtime.
  • agent.skills — the live SkillRegistry (.all(), .get(id), .count).
  • agent.abort() — stop the current run.
  • agent.on(type, cb) / agent.onAny(cb) — subscribe to events.
  • agent.events / agent.store — the raw emitter and store, for advanced wiring.

Events: run.started, step.started/ended, tool.started/output/progress/completed/failed, text.delta/thought/end, phase.changed, agent.state, context.updated, tool.started/tool.completed include presentation when declared (precedence: per-call event → host map → inferred), ask_user.required, permission.required, mcp.approval_required, mcp.resolved, compaction.started/completed, llm.thinking, todo.updated, run.completed/failed/interrupted.

For the full sliced-by-area surface (tools, providers, subcontexts, loop, permissions, hooks) see docs/api.md, and start with docs/getting-started.md.

Hooks

hooks are the seam for logging, metrics, tracing, cost accounting, authorization, and custom policy — no fork required:

const agent = createAgent({
  workspacePath: '/repo',
  hooks: {
    async beforeToolCall({ toolName, input, userId }) {
      if (toolName === 'write_file' && !(await isAllowed(userId ?? 'anonymous', input.path))) {
        return { block: true, reason: 'outside the writable allowlist' };
      }
      return { input: { ...input, content: redact(input.content) } };
    },
    async afterToolCall({ toolName, result, durationMs }) {
      metrics.timing('tool.call', { tool: toolName, durationMs, ok: result?.success !== false });
    },
    async afterModelCall({ usage, error }) { cost.record({ usage, failed: Boolean(error) }); },
  },
});

beforeModelCall / beforeToolCall can rewrite the payload or block the call; after* are observability only and fire on every exit path. A throwing before* hook fails closed — a crashing authz check never allows the call. Details in docs/api.md.

Errors

Every failure is classified before it leaves the loop, so a client can tell the cases apart instead of pattern-matching prose:

interface AgentErrorInfo {
  code: string;                   // 'provider_rate_limited'
  layer: 'provider' | 'tool' | 'run' | 'hook' | 'permission' | 'transport';
  severity: 'info' | 'warning' | 'error' | 'fatal';
  message: string;                // user-facing, no stack traces
  retryable: boolean;             // is a retry worth offering?
  hint?: string;                  // the actionable next step
}

run.warning is the non-terminal one: the loop is still retrying, so a rate limit can be surfaced the moment it happens instead of only after the run gives up. retryable: false on a fatal error means the UI should not offer a Retry button that cannot work. Full table in docs/api.md.


Use it from Claude Code, Codex, opencode, Antigravity, Copilot (or any agent)

This repo ships as a plugin (at plugin/): a SKILL.md (the universal skill format every tool above reads) plus a dependency-free MCP server that guides any agent to build a new looping agent on this library. The plugin dir carries native manifests (.claude-plugin/plugin.json, .codex-plugin/plugin.json, and a strict agent-plugins.org plugin.json), and the repo root carries the distribution files — a Claude marketplace, a Codex marketplace, an Antigravity workspace plugin (.agents/plugins/), a Copilot project skill (.github/skills/), opencode's .opencode/skills/, and an AGENTS.md all agents read.

Claude Code:

claude plugin marketplace add https://github.com/RajdeepDevelopment/smoke-monkey-harness
claude plugin install smoke-monkey-harness@smoke-monkey-harness

Any agent (~70 supported):

pnpm run plugin:install                 # skill → every agent's global skills dir
pnpm run plugin:install -- --list       # see the supported-agent table
pnpm run plugin:install -- --agent cursor  # install for one agent by id
pnpm run plugin:install -- --local      # + project install + .mcp.json + opencode.json
pnpm run plugin:install -- --help       # see options (--force, --repo)

The universal installer reads plugin/agents.json (same cross-agent SKILL.md path table as the AGENTS ecosystem) and drops the skill where each tool natively reads skills — aider-desk, cline, cursor, windsurf, gemini-cli, goose, copilot, opencode, Claude Code, Codex, and more. The portable .agents/skills/ path covers a dozen agents with one --local install.

Every install also drops the 25 bundled agent-skills under <skills-dir>/agent-skills/<skill>/ (the category-wise skills from plugin/agent-skills/skills), so any agent can apply backend/frontend/devops/qa engineering methodology natively — the same set the library loads via loadAgentSkills({ category }) and the MCP server serves via the stock agent-skills-* entries.

Once installed, ask your agent to "build me an agent that …" — it will load the smoke-monkey-harness skill, read harness_guide, and harness_scaffold a starter project on disk. Any agent that reads AGENTS.md at the repo root gets the same full workflow. Details in plugin/README.md.


Development

pnpm install
pnpm run build   # tsc ESM (dist/) + CJS (dist/cjs/)
pnpm run typecheck
pnpm test                  # unit tests (node:test)
pnpm run test:fixtures     # offline MCP client + skill + plugin manifest checks (no LLM)
pnpm lint
NVIDIA_API_KEY=nvapi-... pnpm run example     # examples/basic-agent.ts
NVIDIA_API_KEY=nvapi-... pnpm run demo        # examples/mcp-demo.ts (MCP + custom sub-contexts)
NVIDIA_API_KEY=nvapi-... pnpm run demo:skills # examples/skills-demo.ts (SKILL.md just-in-time)

Deeper material lives in docs/; see CONTRIBUTING.md before opening a PR.

Star History

If Smoke Monkey Harness helps you build more reliable AI agents and agentic developer tools, please consider starring the repository! 🌟

Star History Chart


Community & Ecosystem


License

MIT — free to use, modify, and distribute, including commercially. See LICENSE. Contributions are welcome under the same terms (CONTRIBUTING.md, CODE_OF_CONDUCT.md).

About

Open source AI agent harness & agentic AI runtime in TypeScript. 6-phase state machine, 24 tools, human-in-the-loop permissions, MCP server, and composable agent APIs. Zero dependencies.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages