Skip to content

Repository files navigation

kpopper. An ink-drawn robot works at a drafting surface labelled GROUNDING.yaml. Blue arrows connect Evidence and Assumptions to a Decision, then to What would change it. The lowercase kpopper wordmark has a blue k. The caption reads: The canvas your AI didn't know it needed.

Phone layout

CI kpopper checks its own record Latest release License: MIT native-runtime: Rust reasoning-runtime: Lean 4 history-engine: versioned records

Your project, more self-aware.

Your agents reason. kpopper makes that reasoning explicit, persistent, and deterministically checkable.

Important

TL;DR: kpopper makes your AI sessions less forgetful and your work easier to pick up, check, and build on.

Try it and see for yourself →

kpopper connects decisions to the evidence, assumptions and earlier decisions they depend on, and records what would make them worth revisiting. When a recorded premise changes, kpopper traces its reach through the record and surfaces what needs another look.

Agent reasoning costs time and tokens. kpopper runs calculations and dependency checks deterministically, aiming to turn seconds of model work into milliseconds of computation—and focus the agent on decisions that need judgment.

A changed recorded assumption is highlighted amber in a dependency graph. Deterministic local checks follow its connections in blue while unrelated nodes fade. Decisions A, B and C are returned to the agent for review. A prominent badge reads: A check that takes an agent 10 seconds runs locally in 4.5 milliseconds. When a recorded assumption changes, deterministic checks show the agent which decisions need another look.

Phone layout

Get started · Choose your path · Capabilities · Record format · Commands · Contributing · Why the name?

Meet GROUNDING.yaml, with a handwritten your new friend note pointing to the filename. A large code editor displays the judgment, wrong_if condition, earlier 30/30 review snapshot, and current seven-day retention reading. The earlier session saves the reason supporting a 30-day download email; the later session updates the recorded draft-policy value from 30 to 7. Short blue arrows connect those actions to the relevant lines. A teal connector leads from the condition to kpop check below the code. The pink failure shows 7 less than 30 and FAIL downloads.availability, prompting the agent to review the email or policy. The closing line reads Deterministic Reasoning that outlives the conversation.

Phone layout · Example record

Contents

Where would you like to start?

Choose your path. Click a banner to open its section, or a title to open or close it. Each section includes workflows, examples and the records behind them. You can also go straight to installation.

PR A enables private projects. PR B adds a cache keyed only by query, assuming all results are public. Both changes point to the merged cache, where a private project and a red failure mark show private results exposed.

Coding agents
Claude Code · Codex · Cursor · and more
Explore workflows and examples

Private data exposure

Two agents start from a search service whose results are all public. PR A adds private projects and filters results for each user. PR B adds a shared cache keyed only by the query, relying on those results being public and identical for everyone.

PR A adds private projects in search.py. PR B caches by query in cache.py because all results were public. Tests pass separately and Git merges cleanly. The combination risks serving Alice's private result to Bob. kpopper flags the failed recorded condition and points back to the cache decision. The closing line reads Two green PRs. One data leak.

Phone layout

The files merge cleanly. The reason for sharing the cache no longer holds.

The same record retains the surrounding design: query matching, cache behavior, test coverage, operational limits and open questions.

A richer cache record connects the public-results premise to the query-only cache and its hit-path authorization assumption. The minimap is generated from the complete record.

Phone layout · Full design record

A selected excerpt from PR B:

known:
  search.results_public:
    v: true
    from: s.search
    at: INCLUDE_PRIVATE_PROJECTS is false
    measure: search_results_public
  cache.key_fields:
    v: query
    from: s.cache
    at: 'lookup: query not in _CACHE; _CACHE[query]'
  cache.hit_reuses_result:
    v: true
    from: s.cache
    at: 'lookup: return _CACHE[query]'
  cache.test_data:
    v: One mocked public project; no private-result fixture in that test.
    from: s.cache_tests
    at: test_public_query_is_served_once_across_users
judgments:
  search.shared_cache:
    rests_on:
    - search.results_public
    - cache.key_fields
    - cache.goal
    verdict: Search responses can share a cache keyed only by query because all results are public.
    wrong_if: search.results_public == false
    seen:
      search.results_public: true
      cache.key_fields: query
      cache.goal: Avoid a second search for the same public query across users.
  cache.hit_authorization:
    rests_on:
    - search.shared_cache
    - cache.hit_reuses_result
    - cache.miss_delegate
    verdict: Treat cross-user cache hits as depending on the public-results decision, not as a fresh
      authorization check.
    reopened_by: The hit path, key scope or public-results decision changes; review the combined
      path.

known holds values taken from a source or derived from other entries; judgments holds the decisions that rest on them. rests_on names the premise. seen keeps its value at the last review. When the measured value becomes false, wrong_if fires and identifies the cache decision. Source locators and the complete record are in the example.

How the example measures the change

In the runnable example, a reviewed measurement recipe reads the explicit visibility switch in search.py. After PR A enables private projects, kpop remeasure --run tests that new reading against the cache decision and fails its condition. It leaves the canonical record unchanged; check alone would still see the old recorded value.

The example creates two local Git branches, runs their tests, merges them and checks the combination. A separate integration probe also catches the privacy problem. kpopper preserves the declared reason and connects the change to it; it does not infer arbitrary security properties from code or replace behavioral tests.

The opening overview uses a download promise and a retention policy. The complete merge example shows the same 7-versus-30-day conflict across two branches, with executable code, measurement recipes and a separate integration probe.

These are executable examples. Checks cover the assumptions the record declares and the inputs deliberately measured or recorded. The examples on this page retain their compact legacy records, which the native runtime continues to read without migration. New records also preserve immutable history.

Two working modes

The difference is whether sessions share one project context or work against different versions of the code. Both modes support several sessions and competing hypotheses.

GROUNDING.yaml is the readable project record shown below: one shared context in Simple, and a version kept with each branch, including main, in Advanced. New history-backed records also keep their supporting history in .kpopper/; carry that directory with the YAML when sharing or versioning the record.

Simple: sessions share one sourced project record labeled GROUNDING.yaml, with competing hypotheses beside it; consolidation compares and checks proposals, then folds them into the record, refutes them with a reason, or leaves them pending. Advanced (experimental): Branch A, Branch B and main each have a GROUNDING.yaml record for their version of the code. Branch records pass through consolidation before integration into main. A continuing kpopper Board, shared findings pending review path captures feature-independent findings with their source, scope and status. Dashed arrows show local reading before merge. With permission, findings enter a knowledge PR, where consolidation reconciles them with the target record before acceptance. Both paths reach the same main record; the shared path continues for the next batch.

Phone layout

View the full-size illustration · View the vertical version

Simple — one shared context. Sessions read and contribute to the same project graph. Competing ideas are named hypotheses beside it: several sessions can examine the same proposal, and one session can work on several. For a research project, for example, sessions can read different papers and compare explanations against the same sourced record. Consolidation is how those proposals become part of the shared record: compare them with what it already holds, check the combined dependencies and resolve conflicting claims. A proposal can be folded in, refuted with its reason retained, or left pending.

Advanced (experimental) — branch contexts and shared findings. Each branch's record describes its version of the code. A cache decision on one branch may depend on results being public, while another branch introduces private results. Keeping those premises with their branches lets review check whether the reasoning still holds when the changes are combined. Consolidation reconciles those records: it identifies overlapping subjects and conflicting claims, and checks which decisions need another look under the combined premises. Resolve what needs judgment before folding a proposal in; unresolved hypotheses remain explicit.

[!WARNING] Advanced mode is experimental. Parallel branches can add different entries to the same part of GROUNDING.yaml, leaving an already-checked PR with a merge conflict after another PR lands. Resolving it can trigger another CI run. Automatic conflict-free integration is not implemented. We are keeping the readable file while this remains an open design question: the problem, alternatives and their tradeoffs.

For a supported local record conflict, preview a resolution with kpop consolidate --resolve --dry-run. Competing claims still need a decision.

Knowledge then follows two paths:

  • Feature knowledge travels with its branch. Its assumptions, measurements and hypotheses stay attached to the code they describe. Consolidation tests their combination with the target record as part of reviewing the change.
  • Shared findings have a continuing path of their own. Shareable findings that apply independently of the feature enter pending_grounding, with their sources and scope. They remain available even if the originating session closes or its worktree is removed.

Suppose a session building an integration discovers a documented change to the vendor's API limit. The feature may be abandoned; the finding can still help the project. Other local worktrees can read it immediately, marked as pending, without waiting for the feature to merge. It appears alongside their branch record; reading it does not adopt it or replace their recorded premises. Their consolidation dry run tests it against that record too, and turns red when accepting it would break a recorded decision, until a person decides the finding.

kpopper Board is this continuing shared-findings area. At the start of meaningful work in an unconfigured Git project, kpopper offers Simple (recommended) with a chosen external shared record, or Advanced (experimental) with branch-specific knowledge. The illustration above explains both. Advanced can keep Board findings local or publish them to a named repository and target through one continuing knowledge PR. A publication choice enables branch and PR updates, not merging; it is remembered across local worktrees. kpop board shows the current mode, selection and publication state.

With the project's publication permission, Board findings accumulate in one knowledge PR. Consolidation reconciles the proposed knowledge with the target record before acceptance; conflicts and changed premises need resolution. Accepted contributions are then verified in the target branch, shown as main above. The same publication branch is reused for the next batch, while the local contribution history persists across review cycles. Local capture and reading also work without publication.

A measurement of an unmerged commit can be a fact about that commit. A proposed conclusion remains a hypothesis. Merging means the team accepted the contribution; it does not prove the claim, increase confidence or refresh its last review. Private material and information whose sharing permission is unclear stay in a structured private draft.

Projects without Git start in Simple; new Git projects currently start in experimental Advanced mode, including those with only one checkout. Simple can also be configured for a Git project with one external shared record. Existing registered shared records keep their current location and behavior; changing mode requires explicit reconciliation.

See project modes and publication for routing, reproducible reads and the publication lifecycle, and consolidation for the dry run, folding and refutation commands.

These project modes also apply to document and research work.


A plan built around one venue kitchen meets workshop participants cooking remotely from their own homes.

Claude Cowork / ChatGPT Work
Keep plans current when the brief changes.
Explore workflows and examples

Outdated planning assumptions

A cooking workshop is planned around the venue's shared kitchen, equipment and ingredients. A later client brief moves it entirely online. For ChatGPT Work, see the runtime requirements and current limits.

A later session records the client's change from an onsite cooking workshop to an online event. The saved plan still assumes one shared kitchen with equipment and ingredients provided. kpopper reports that the workshop format moved from onsite to remote. The agent needs to revisit equipment, ingredients and activities; this is a review notice, not an automatically failed conclusion.

Phone layout

Once the agent records the new format, kpopper flags the saved plan for review. It does not decide whether the activities can work remotely. The next session can recover the old reason, read the updated brief, and work out what participants need. Try the Cowork example.

The planning record connects 13 readings to five saved plans and four open questions. The format change reaches the agenda, equipment, ingredients, group arrangement and supervision; the earlier review snapshots remain visible.

A complete planning record shows the remote format beside five plans last reviewed as onsite, with unchanged client constraints and unresolved home requirements.

Phone layout · Full planning record

Selected entries from the after record:

known:
  workshop.format:
    v: remote
    from: s.updated_brief
    at: Participants join online from home
  workshop.participants:
    v: 18
    from: s.session_plan
    at: Client constraints
  workshop.duration_minutes:
    v: 90
    from: s.session_plan
    at: Client constraints
  venue.workstations:
    v: 6
    from: s.logistics_plan
    at: Venue provision
judgments:
  workshop.agenda:
    rests_on:
    - workshop.format
    - venue.shared_kitchen
    - workshop.duration_minutes
    - workshop.recipe
    verdict: Use the shared-kitchen agenda and provide ingredients at the venue.
    reopened_by: The format or access to the shared kitchen changes; review activities, equipment
      and preparation.
    seen:
      workshop.format: onsite
      venue.shared_kitchen: true
      workshop.duration_minutes: 90
      workshop.recipe: Fresh pasta with tomato sauce.
  workshop.equipment_plan:
    rests_on:
    - workshop.format
    - venue.equipment_supplied
    - venue.workstations
    verdict: Plan equipment around the six venue workstations.
    reopened_by: The format, supplied equipment or access to the workstations changes.
    seen:
      workshop.format: onsite
      venue.equipment_supplied: true
      venue.workstations: 6
  workshop.ingredient_plan:
    rests_on:
    - workshop.format
    - venue.ingredients_supplied
    - workshop.participants
    verdict: Prepare ingredient portions at the venue for the 18 participants.
    reopened_by: The format, ingredient provision or participant count changes; review purchasing,
      portions and distribution.
    seen:
      workshop.format: onsite
      venue.ingredients_supplied: true
      workshop.participants: 18

The current value is remote; the decision was reviewed against onsite. That is a MOVED notice. The prose in reopened_by tells the agent what deserves attention; it is not an executable predicate. Read the complete after record.

Before the first answer

The next session opens with the saved context: all five plans need review because workshop.format changed from onsite to remote. Then the user asks:

Prepare the ingredient portions for the 18 participants.

On a host with the prompt hook enabled, this matching prompt produces a focused pointer (actual hook output, shown here with Codex skill syntax):

kpopper: the record holds workshop.ingredient_plan on this - $ground workshop.ingredient_plan before answering from memory.

The agent follows that pointer with pull workshop.ingredient_plan. It retrieves remote, the old venue-based plan and its onsite review snapshot before answering. The next decision is how ingredients will reach participants at home. See the opening and retrieval output.

Keep your existing documents, notes and task tools. The record links back to relevant evidence; there is no need to migrate your knowledge system. For another example that needs judgment, explore a replacement offer with an uncertain deadline.


Galaxy rotation, gravitational lensing and the cosmic microwave background introduce the research evidence.

Research
Connect evidence and revisit conclusions as findings change.
Explore workflows and examples

Dark matter across studies

The original three-paper view shows the shared-record idea:

Three research tasks connect galaxy rotation, Bullet Cluster lensing and the CMB in a shared record, with assumptions and open questions attached.

Phone layout

A research record can connect observations across scales and retain the tensions between them. This worked example follows six papers: Rubin's galaxy rotation, SPARC's galaxy catalog, the radial acceleration relation, Bullet Cluster lensing, Planck's cosmological fit and the first LZ particle search.

Six papers connect galaxy rotation and baryonic structure, cluster mass location, the CMB fit and a particle-search limit. Their assumptions and shared inputs remain attached to the synthesis.

Phone layout

The argument includes 36 source readings, two scope readings, one calculation, seven linked judgments and five open questions. Its depth comes from the links: SPARC and the acceleration paper share data; the Planck density ratio comes from one model fit; a null WIMP search constrains specific interactions without settling the identity of the astronomical mass component.

The research GROUNDING.yaml includes six paper sources, detailed readings, an exact density-ratio calculation, intermediate judgments and the final synthesis. A framework change reaches multiple interpretations.

Phone layout · Complete research record · Sources and exact locations

A source-grounded excerpt:

known:
  sparc.sample_size:
    v: 175
    from: paper.sparc
    at: Abstract
    unit: galaxies
  rar.sample_size:
    v: 153
    from: paper.rar
    at: p. 1, Data / Galaxy Sample
    unit: galaxies
  rar.points:
    v: 2693
    from: paper.rar
    at: 'p. 1, Galaxy Sample: velocity-precision cut'
    unit: points
  cmb.omega_c_h2:
    v: 0.12
    from: paper.cmb
    at: Abstract, combined analysis
    uncertainty: 0.001
    confidence: 68%
  cmb.omega_b_h2:
    v: 0.0224
    from: paper.cmb
    at: Abstract, combined analysis
    uncertainty: 0.0001
    confidence: 68%
  lz.si_limit:
    v: 9.2e-48
    from: paper.lz
    at: 'Abstract, version 4: limit at 36 GeV/c^2'
    unit: cm^2
    confidence: 90%
  lz.mass_at_limit:
    v: 36
    from: paper.lz
    at: Abstract, version 4
    unit: GeV/c^2
  cmb.dark_to_baryon_density:
    rule:
      expr: cmb.omega_c_h2 / cmb.omega_b_h2
judgments:
  evidence.shared_catalog:
    rests_on:
    - sparc.sample_size
    - sparc.inputs
    - rar.catalog
    - rar.selection
    verdict: Treat SPARC and the RAR analysis as related evidence, not two independent observational
      votes.
    reopened_by: A different catalog, sample selection or independently calibrated replication changes
      the dependence between these findings.
  synthesis.dark_matter:
    rests_on:
    - galaxy.extended_mass
    - galaxy.baryon_coupling
    - cluster.mass_location
    - cosmology.cold_component
    - particle.search_scope
    - research.framework
    - research.selection
    verdict: Under the stated models, the selected evidence supports a dark-matter account across
      scales, while baryonic regularities and direct-search limits constrain its explanation.
    reopened_by: A source finding or shared assumption is revised, or a worked alternative jointly
      addresses the galaxy, cluster, cosmological and particle-search constraints. Reassess the
      affected paths and the synthesis.

The next session inherits the argument, including its unresolved questions.

Run the checker on that same record:

kpop check examples/dark-matter/GROUNDING.yaml

Its final summary (exit 0; the full response also lists the seven prose review conditions):

7 judgments, 58 entries, 0 problems, 7 declared

For a starting overview use open; to inspect a subject use pull; to trace a premise use affects. The ground skill guides that workflow for an agent.

Read the context, inspect the calculation, and trace a change

Current context — selected lines from the actual open response:

kpop open examples/dark-matter/GROUNDING.yaml --chars 1000
Six-paper worked example; selected source readings and an authored synthesis, not a complete review or live research-agent evaluation.
holds: rar (8) · research (8) · cmb (7) · lz (6) · paper (6) · sparc (6) · lensing (5) · rotation (5)
58 entries, 7 judgments, 5 open questions, updated 2026-09-19

A computed reading and the judgment that uses it:

kpop pull cmb.dark_to_baryon_density examples/dark-matter/GROUNDING.yaml --budget 1100
cmb.dark_to_baryon_density: 75/14 = cmb.omega_c_h2 / cmb.omega_b_h2 (Ratio of the abstract central physical de ...
+ cosmology.cold_component: Within base Lambda-CDM, the fitted physical cold-dark-matter density exceeds the b ...
    holds
    because: The density ratio connects two central values from the same fit. The model assumptions and
             retained lensing-amplitude tension limit the interpretation.
    reopened by: The likelihood, data combination or cosmological model changes enough to alter the inferred c ...

affects <entry> shows what a change reaches

The reach of the adopted framework:

kpop affects research.framework examples/dark-matter/GROUNDING.yaml
cluster.mass_location
    via research.framework -> flagged only
cosmology.cold_component
    via research.framework -> flagged only
galaxy.extended_mass
    via research.framework -> flagged only
synthesis.dark_matter
    via research.framework -> flagged only

4 judgments reached

See all captured commands and responses, including pull synthesis.dark_matter. Here check reports record consistency; the scientific review conditions still require judgment.

The illustrations are schematics drawn from the record. The full example distinguishes source readings, authored interpretations, shared assumptions and the computed 75/14 density ratio. That arithmetic does not establish the scientific synthesis.

Run the shared-record example: three concurrent CLI writers retain the prepared six-paper contributions, and the writer records each judgment's review snapshot. A proposed change of framework then reopens several interpretations and the synthesis. No model call or claim of live independent research is part of that replay.

Get started

You can use this request with any of the agents listed below. The agent can carry out the steps its tools allow and guide you through any clicks, approvals or administrator steps that need your input.

Copy and paste the following text to your agent:

Dear agent,
Please help me install kpopper in this environment. Open the guide below and
follow the installation instructions for your environment:
https://github.com/ilanbm/kpopper#get-started
Complete the setup, including the required runtime. Guide me through any
steps that need my input, and tell me when to start a new session.

The sections below explain each app's setup and provide manual alternatives.

Claude Code

In the Claude Code chat: paste the installation request above. It covers both the plugin and its required runtime.

Manual alternative: slash commands or terminal commands

Inside Claude Code, enter these slash commands one at a time:

/plugin marketplace add ilanbm/kpopper
/plugin install kpopper@kpopper

Or, in a terminal, run the equivalent commands:

claude plugin marketplace add ilanbm/kpopper
claude plugin install kpopper@kpopper --scope user

Both routes install the plugin. The native runtime is a separate required step; ask Claude Code to finish it, or follow the manual runtime setup. That guide explains how to find the installed plugin directory and includes Windows instructions. Start a new session after setup is complete.

Codex

In a Codex task: paste the installation request above. The agent can use its terminal tools for setup when available; it will guide you through any client prompts that need your input.

Manual alternative: terminal commands and hook approval

In a terminal:

codex plugin marketplace add ilanbm/kpopper
codex plugin add kpopper@kpopper

Install the native runtime in the plugin root printed by codex plugin add using the native runtime setup guide. Then start a new task and review and trust kpopper's hooks when Codex asks. You can also inspect the hook definitions inside Codex with:

/hooks

See Codex setup and behavior.

Claude Cowork

In Cowork: paste the installation request above for guided setup. Claude guides the setup while you complete the marketplace and Install clicks in Customize → Plugins: select Add marketplace, enter ilanbm/kpopper, then install kpopper.

Cowork runs tasks in its own environment, where kpopper has not yet been verified end to end; see the Cowork notes.

ChatGPT Work

In Work: paste the installation request above for guidance through the workspace setup. If the plugin is not available, a workspace administrator must import it from Admin → Plugins → Add → Import marketplace, with https://github.com/ilanbm/kpopper as the source and Path left empty. Members can then install it from Plugins.

kpopper's runtime in Work is not yet validated; see Work setup and current limits.

Grok Bot

In Grok Bot: paste the installation request above. Follow the Grok Bot guide to install the runtime on its cloud computer, choose a persistent project folder and save a skill that reads kpopper's shared instructions.

This route has not yet been verified in a live Grok Bot session. It uses explicit commands; automatic hooks and plugin import remain unverified. Grok Build has a separate plugin system, so its Claude Code compatibility claim does not establish Grok Bot support.

Other agents

In Cursor, Gemini CLI, Windsurf, GitHub Copilot, OpenClaw or OpenCode: paste the installation request above. The agent should follow the adapter for the app and environment you're using:

Cursor · Gemini CLI · Windsurf · GitHub Copilot · OpenClaw · OpenCode.

Manual alternative: local checkout and adapter setup

In a terminal, clone the repository where it can stay:

git clone https://github.com/ilanbm/kpopper.git

Follow the adapter linked above. Add to the agent's existing configuration rather than replacing it. What runs automatically differs by agent; see the capability matrix and verification status.

After installing

In Claude Code and Codex, kpopper opens with each new session. Installing creates no record; your agent starts GROUNDING.yaml when it records the first finding worth keeping (on native Windows, under WSL). To start from what already exists, ask your agent to map the project. The native runtime needs no Python, Node, Rust or Lean toolchain.

Use the command line without an agent · Runtime setup and troubleshooting

Installed? See it in your work: Coding agents · Claude Cowork / ChatGPT Work · Research

Quick reference

Use the CLI in a terminal, or ask your agent to follow a plugin skill. A skill guides a workflow and may use several CLI commands. You can also describe what you need in ordinary language.

CLI

These commands assume an installed kpop. From a source checkout, build the native crate with cargo build --manifest-path native/Cargo.toml --release and use the resulting kpop binary. For plugin work, use the command supplied by the session's KPOPPER_AGENT_CONTEXT. <id> names a record entry, such as workshop.ingredient_plan; kpop --help lists command groups.

Command What you need
kpop where Locate this project's record
kpop open Open the current context and attention items
kpop pull <id> Retrieve a subject, its sources and reasons
kpop context <id> Read records with their declared dependencies and checks; use --direction impact for dependents
kpop search "terms" Find matching claims and local source passages
kpop affects <id> Trace what depends on a premise
kpop check Check recorded conditions and changed premises
kpop assess <id> Inspect findings and scoped attention as JSON
kpop export <id> Share a focused Markdown excerpt
Writing, history and background commands

These commands can record decisions, update project state or configure work. The linked guides describe their arguments and review steps.

Command or command group What you need
kpop add <id> field=value ... Add a finding or judgment
kpop set <id> <value> --why "reason" --as-of YYYY-MM-DD Update a reading with its reason and date
kpop update --file report.json Apply one prepared source report
kpop review <id> Record a completed review
kpop same <a> <b> · kpop distinct <a> <b> "reason" Resolve whether two IDs name the same subject
kpop answer <question> <id> · kpop answer <question> --dropped "reason" Close an open question with what answered it, or with why it no longer matters
kpop correct <id> field=value ... Fix an entry that no commit holds yet
kpop consolidate --dry-run · kpop consolidate --from <ref> --dry-run Test hypotheses or another branch before folding
kpop consolidate · kpop consolidate --refute <name> "reason" Fold eligible hypotheses or retain a refutation
kpop history status Inspect committed history acceptance
kpop remeasure --run Rerun the record's configured measurement recipes
kpop map Begin a guided mapping of existing material
kpop ingest Capture and process reports in the background
kpop watch Configure branch checks or inspect shared findings
kpop followups Manage deferred work and its recorded outcomes
kpop config Inspect or change workspace guidance

See the command reference, history commands, background capture and followups. review records your assessment; it does not make that assessment for you.

Plugin skills

Command Skill What happens in practice · possible CLI calls
/kpopper:kpopper kpopper Explains the method and chooses the workflow that fits the task. CLI calls follow the selected workflow.
/kpopper:ground ground Finds relevant IDs, then prefers kpop context <id> for records with dependencies and checks.
Uses pull for concise readings or when the checked reader is unavailable, and affects for changed inputs.
/kpopper:record record Saves findings, their sources and reasons; records decisions, open questions and completed reviews.
May use kpop update --file report.json, kpop add, kpop set or kpop review.
/kpopper:map map Examines the agreed materials, builds a sourced record and reports coverage and gaps.
Starts with kpop map --json or kpop map --deep --json, then follows the returned workflow.
/kpopper:consolidate consolidate Compares proposals with the record, surfaces disagreements and guides folding or refuting them.
May use kpop consolidate --dry-run, kpop consolidate or kpop remeasure --run.
/kpopper:watch watch Inspects branch checks and shared findings; configures background checks or daily review when requested.
May use kpop watch status, kpop watch setup, kpop watch shared or kpop followups daily install, plus the host's scheduler.

The agent chooses the calls for the task and the record's state, using its source-reading tools as needed.

For example, /kpopper:ground workshop.ingredient_plan asks the agent to retrieve that plan and its basis. ground is a skill; the CLI reads use open, pull, affects and check. Other hosts expose skills through their adapters.

Experimental — optional HTML applications

These skills require the optional HTML setup and explicit selection. Their interfaces and artifact formats may change.

Command Application What happens in practice · possible CLI calls
/kpopper:hub kpopper Hub Builds a browsable snapshot of the record and checks its layout and coverage.
Uses kpop experimental hub, for example with --open or --verify.
/kpopper:annotated-doc Annotated Documents Authors or refreshes a standalone HTML document with selected evidence and reviewable copy updates.
Uses kpop experimental annotated-doc with operations such as guide, build and refresh.

The CLI entry points are kpop experimental hub and kpop experimental annotated-doc. page and document remain compatibility aliases.

Context during a conversation

Codex opens a canonical view by default: original bodies for the selected evidence, navigation for the rest, and automatic anchors for follow-up questions. Full selected views remain the default; incremental delta delivery is experimental and off. Other hosts retain their existing opening. The persistent record needs no migration for this upgrade.

Follow the conversation lifecycle, try the worked CLI example, or use the checked-session API for explicit service setup.

What you can do with kpopper

Carry the reasoning into the next session, check it against recorded inputs, and keep the earlier evidence available when the work changes.

What you want to do What kpopper provides Example
Resume work with its context Open the project's standing decisions and attention items, then retrieve the facts and reasons relevant to a question. Before the first answer. Resume a job search and recover why three roles were shortlisted.
Trace why a decision was made Follow its sources, declared dependencies and the values used at its last review. The knowledge record. Trace the upload queue decision back to the test that exposed request timeouts.
Keep earlier decisions inspectable History-backed records retain immutable claim versions and explicit acceptance, review, correction and refutation acts as the current record evolves. History. See why a trip moved from July to August, without losing the original constraints.
Calculate and check explicit rules Evaluate exact arithmetic, compound Boolean conditions and conditional expressions with the packaged native reasoning runtime. Missing inputs and execution errors remain visible. Deterministic reasoning. Check whether 24 guests fit a venue with 18 seats.
Ask questions over a recorded collection Filter, select, count or sum within a declared scope, or test whether all/any members meet a condition. The result retains scope evidence and diagnostics. Collection queries. Find apartments below $2,000 with an elevator and a lease that allows pets.
Notice which decisions need another look Trace changed premises, evaluate declared breaking conditions and explicitly rerun configured measurement recipes. Checks and measurements. Record a babysitter's cancellation and surface the evening plans that depended on it.
Test alternatives and reconcile branch work Keep named hypotheses, inspect the proposed combination and retain conflicts or refutations. Worktrees keep their code-specific context, with checks available before a merge and in CI. Two working modes. One branch removes password login; another adds a feature that still requires it.
Keep useful findings across sessions and branches Capture explicit reports in the background and retain scoped project contributions with their sources and pending/accepted status. Background capture · Shared contributions. A discarded prototype's documented API limit remains available to the next integration task.
Return to work when its conditions change Tie followups to dates, changed recorded inputs or earlier work, with scheduling and delivery configured in the host. Followups. Resume tax preparation once the missing bank statement has been recorded.
Share a focused piece of the reasoning Export selected entries as Markdown, with earlier/current readings, omitted values marked and optional Mermaid diagrams. Focused exports. Share why you chose a school, including the commute times and fee comparisons.
Build a tool on structured findings Read versioned assessment JSON with computational results, review comparisons, contention, integrity and history evidence kept distinct. Assessment contract. Build a grant dashboard that separates budget overruns from missing receipts.
Let the record fit the work Add domain-specific subjects and vocabulary while preserving explicit sources, dependencies and review conditions. Evolving structure. Organize a garden plan around plants, watering schedules and frost precautions.

New records use core/v1 and immutable history by default; existing legacy records require explicit adoption. Collection queries use the declared query/v1 capability. The checks cover recorded inputs and supported rules. The agent still interprets sources and makes judgments. Your existing documents, tools and memory stay where they are.

Advanced mode's branch and shared-findings workflow is experimental. Parallel PRs can still conflict on the record and need another CI run after resolution. See the open merge design question before relying on this workflow.

Optional experimental tools

What you want to do Tool and scope Example
Browse the record visually kpopper Hub: a rendered snapshot with project layouts, source links and an interactive graph. Included in native bundles; the application remains optional and experimental. Applications. Open a visual overview of a renovation's quotes, decisions and unresolved questions.
Share a document with inspectable evidence Annotated Documents: standalone HTML with selected source snapshots and reviewable copy updates. Included in native bundles; the application remains optional and experimental. Document workflow. Produce a client report with the source invoices beside each expense total.
Bind an agent's reads to a known revision Checked sessions: a revision-bound view and optional MCP transport, with their own setup and session checks. Checked-session integration. An agent refreshes its view after another session changes the recorded API contract.

Keep the conversation moving

New information often arrives halfway through another task. kpopper can retain an explicit report and process a supported update in a separate worker. Routine results stay quiet; important unresolved findings are available for delivery back to the conversation.

The main agent captures an explicit venue-cancellation report and continues with the set list. A software worker in a separate process saves the dated source, records the change from confirmed to cancelled and checks the venue-to-announcement dependency. The announcement needs review; important findings return through the configured delivery route while routine updates stay quiet. The closing line reads Get notified only when something needs attention.

Phone layout

Captured, applied and checked are different states. If the current answer depends on an update, inspect its outcome before relying on it. Background work is useful where the conversation can safely continue without that result.

Background updates and delivery

Supported reports can update stored values and add new readings, rules and judgments. Related changes from one source can be applied together. The agent supplies the source quotation and explicit updates; background processing does not guess what an ambiguous message means. Existing judgments still require deliberate review.

Supported record layouts and authoring limits are described in the background capture guide. Delivery after an answer requires the host capabilities described in the native delivery guide.

Keep the knowledge system already in use

Your Markdown files, Obsidian vault, project wiki and agent memory can stay where they are. Your agent reads relevant material through its available tools and keeps sourced claims and decisions in a project record.

How existing knowledge connects to the record

GROUNDING.yaml presents the claims the work relies on, with links back to their sources; new records retain their history in .kpopper/. There is no need to migrate the existing notes or replace the agent's memory system. This applies to coding, planning and research.

Dense clusters of notes and memory, documents and research, conversations, plans and commitments, and code and data fill the left side. An agent selects relevant evidence. On the right, kpopper arranges claims, decisions and review conditions in GROUNDING.yaml. Sources stay put; reasoning stays connected.

What “deterministically checkable” means

A recorded decision can name the values it relied on and a condition that would make it wrong. kpopper evaluates that condition itself, without asking a model, so the same values always give the same answer.

Here is the example from the top of this page, step by step.

1. The agent records the decision and what it relies on. A download email promises a working link for 30 days, because files are kept for 30 days. The agent records both numbers, the decision, and the comparison that would break the promise:

known:
  files.days: {v: 30}
  link.days: {v: 30}
judgments:
  downloads.availability:
    rests_on: [files.days, link.days]
    verdict: "Retention supports the promised window"
    wrong_if: "files.days < link.days"
    seen: {files.days: 30, link.days: 30}

known holds the recorded values, and judgments holds the decisions that rest on them. rests_on names the values the decision depends on. wrong_if is the condition that would make it wrong. seen keeps the values the decision was made with.

2. A later session changes one value. A draft policy keeps files for 7 days, and the agent records files.days: 7.

3. kpopper evaluates the condition. 7 < 30 is true, so the promise no longer holds. The agent sees the broken decision as soon as it records the new value, the next session opens with it under “needs a person”, and kpop check fails:

$ kpop check
FAIL downloads.availability: wrong_if holds (files.days < link.days) - broken by its own condition

1 judgments, 3 entries, 1 problems

The agent then reviews the email or the policy. Nobody had to remember that the two were connected. kpop check exits with status 1 when a condition fires, so a CI job can fail on it as well: add reasoning checks to CI.

What this asks of you. You don't write the YAML. The agent writes these entries when it records a decision, and you can ask for one directly: “Record this decision and what would make it wrong.” The entries stay readable in GROUNDING.yaml. When a check fails, you and the agent decide what changes; kpopper points to the decision that needs another look.

What the program checks, and what stays judgment:

  • Checked by the program: comparisons and calculations over recorded values, such as files.days < link.days or guests > seats. It also compares each input with the value the decision was made with: a change the condition allows stays quiet, and a change no condition covers is flagged for review.
  • Left to the agent or a person: conditions that need reading or judgment, such as “Legal asks us to keep customer files for less time.” The record keeps them as text and shows them for review; the program never evaluates them.
  • Not checked: whether a recorded value is still true in the world. kpopper does not watch your documents. A value changes when someone records a new one, or when a measurement you set up re-reads it.

A passing check means that none of the recorded conditions fired. It does not mean the decision is right.

Supported calculations and conditions · What check reports · The reasoning core's formal scope

Go deeper

Open a topic when you need its details. The examples, commands and explanations remain here for reference.

One project, across your existing tools

A project is work around a goal. Its materials may span documents, conversations, calendars, task systems, files and earlier sessions. In software, they also include code, commits and pull requests. One project can cross several tools; one source can serve several projects.

kpopper presents a picture of the project's reasoning in GROUNDING.yaml: a readable record that connects claims to sources and decisions to their premises, backed by immutable history in .kpopper/ for new records. Your documents and tools keep their own content. The record makes the reasoning between them available to the next person or agent working on the goal.

When you return to… The useful thing to recover
A software project in Codex or Claude Code Why a design was chosen, the code or test supporting it, and the changes that could invalidate it.
Financial analysis in Claude Cowork Which sources support a forecast's assumptions, when they were checked, and which decisions depend on them.
A product launch Which commitments support the launch plan, and which assumptions changed.
Your weekly plan Why a task has priority, the deadline behind it, and the availability it assumes.
Research with an agent and an Obsidian vault The evidence for an explanation, competing accounts, and the observation that would challenge it.
A mortgage application Which lender offer and documents support the plan, when the offer expires, and which conditions still need confirmation.

The same five questions orient the work:

  1. What are we trying to achieve? Goals, outcomes and priorities.
  2. What is known now? Facts, commitments, deadlines and constraints.
  3. What was decided, and why? Decisions, assumptions and alternatives.
  4. Where is the evidence? Sources and enough detail to find the relevant passage again.
  5. What needs another look? Open questions, conflicting reports and changed premises.

These questions guide what to record. Learn from findings during ordinary work, or ask for an initial map or deeper investigation of selected materials. The agent uses the sources available in your context; access to a file alone does not make it part of the project.

Past, present, future

Keep the work connected across time: the sources and decisions behind it, what needs attention now, and the checks or actions to return to later.

Three ink panels connect past sources, evidence, decisions, reasons and last-review snapshots; present claims, changes, contradictions and new information; and future followups, daily reviews and time or event triggers when configured. A return loop explicitly says the agent records outcomes as evidence.

Phone layout

The Future panel shows deferred work tracked by followups and scheduled reviews when configured in the host. The return arrow is the agent's step of recording useful outcomes as evidence; marking a followup complete is a separate operation and does not automatically rewrite the knowledge record.

How it works — record format and review

How a decision gets checked is explained above.

The knowledge record

The technical term is an epistemic record: a record of what is known and how it is grounded. These are roles in the method, not six mandatory YAML sections. Start with what the work needs; a source and one finding can be enough.

The native runtime preserves immutable versions of new claims and recorded acts. GROUNDING.yaml presents the current readable record; .kpopper/ holds the history and its authority metadata. Retain both together. Supported CLI writes update the record through that history, and ordinary reads automatically use core/v1. An existing legacy YAML record is not migrated by reading it. See the history contract for adoption and editing rules.

Piece What it preserves In a new record
Source The document, conversation, observation or other origin of a claim, with dates and locators. sources
Reading A value or quotation taken from that source. known
Derivation A structured rule relating inputs to a result. Its references supply graph dependencies and the local Lean core computes its value. Legacy text rules remain readable and unevaluated until explicitly converted. known, as a rule
Judgment A conclusion, its reasoning, declared dependencies and condition for reconsideration. judgments
Review snapshot What those dependencies held when the judgment was last reviewed: seen. seen, inside each judgment
Open question Something unresolved, retained without inventing an answer. open

The method is opinionated about grounding conclusions, declaring dependencies and preserving a basis for review. A judgment needs the values it was reviewed against to make drift detectable, and a meaningful condition for reconsideration. A record with no judgments yet does not need invented conclusions, snapshots or derivations just to fill a template. Field names and project-specific categories are described below.

Change is compared with the last review. When a recorded scalar differs from a judgment's seen snapshot, the reader identifies the movement. A supported wrong_if comparison says whether the change crosses the judgment's stated boundary. Movement can call for review without refuting the conclusion; a movement inside its declared threshold can stay quiet. affects traces the wider reach through declared judgments and rule references.

The record does not observe the world on its own. A changed document matters only after someone or an authorized tool supplies a new reading. Optional measurement recipes can re-read selected values when explicitly run. These are reviewed commands, not general source monitoring.

History and the present belong together, with their dates intact. An old conversation can explain why a deadline was chosen; it does not establish that the deadline still holds. A session's proposed action does not establish that anyone performed it. Reusing a recommendation requires evidence for its conditions now.

Competing claims can stay separate. A conflicting write can become a hypothesis beside the base record. Consolidation checks the proposed combination before an explicit fold; it can also retain a refuted hypothesis as a negative finding. A signal never grants permission to change a decision or act outside the record.

On return, open gives a bounded orientation and attention report; pull retrieves the subject you need. The optional checked session mode below adds a complete, navigable view within a token budget. In both cases, the aim is to spend the next session's context on the work at hand.

A structure that grows with the project

Start with a finding, not a database design. A research project can name claims and experiments; a workshop can name participants and supplies; a codebase can name interfaces and deployment assumptions. Add subjects, categories and views when the work creates a reason for them.

The readable record is YAML, and Git is optional. Keep it and its .kpopper/ directory with the project or in a deliberately configured external location; your source documents stay in their existing tools. Storage and location.

The vocabulary is flexible. The ordinary reader recognizes dependency, predicate and snapshot roles by their shape; facts/claims can serve the same purpose as known/judgments. The documented names are the easiest starting point. When two fields fit the same role, an explicit schema mapping resolves the ambiguity. Some control fields, including reopened_by and blocked_on, are recognized by name; flexible vocabulary does not mean every keyword can be renamed.

The relationships stay explicit: where a reading came from, what a decision depends on, what it was reviewed against, and what would bring it back for review. Supported checks enforce their structural and comparison rules; arbitrary prose still needs interpretation.

This is the useful sense of an evolving record: the person or agent can adapt its structure as needs emerge, and later checks recompute what moved from the current values and saved snapshots. Conclusions are not silently rewritten, and a structural change should migrate the existing record and pass its checks in the same change. See the shape and how structure grows.

When a review needs judgment

Some conditions can be compared mechanically; others require reading a source and making a judgment. reopened_by preserves a prose condition for reconsideration. blocked_on records why a condition cannot currently be checked. Neither field calls a model or schedules work.

When a declared, comparable premise changes, the write response and subsequent open or check can surface the affected decision. During the task, the agent reads the relevant sources and decides whether to retain, revise or question it. A review explicitly updates seen; the checker never does that on its own.

For a review that must happen later, attach a followup to a date or recorded change and connect it to an available host schedule. The optional daily review can also select one flagged decision for attention. A prose condition alone is not an automatic background review of every judgment. See followups for triggers, work budgets and scheduling.

Followups and background checks

Deferred work can become ready on a date, after a recorded value changes, when a threshold is crossed, or after another task finishes. Keep the task in your existing system; kpopper connects it to the knowledge it depends on and has a private local fallback when needed.

A configured daily review can revisit due work and a flagged decision within a small budget, surfacing meaningful results. During active sessions, supported host hooks also surface changed readiness. A saved followup alone does not start an agent or create a schedule.

With local watch enabled, worktree graph changes are checked in a separate background process against the selected main ref. Results identify the exact versions checked; comparison rewrites neither graph. Scoped external observations can live in one canonical record outside the branches, while branch experiments remain local.

Run /kpopper:watch in Claude Code or $watch in Codex to inspect and configure these routines. The host supplies scheduling within your authorization. Recording a followup's outcome and reviewing a decision are explicit actions.

See setup and host limits, followup routing and review budgets, and branch compatibility and shared observations.

Coding: check the reasoning behind a merge

Two branches can pass their own tests and undermine each other's decisions when merged. Git checks whether their text can be combined. kpopper adds a check on the recorded premises and conditions behind the work.

The two executable merge stories reproduce the opening illustrations in disposable local Git repositories:

Story Independent changes The declared condition that fails together
Search cache Add private projects; cache results by query because all results are public. search.results_public == false
Download promise Retain files for seven days; promise downloads for 30 days. exports.retention_days < downloads.promised_days

Run them with Git and a native kpop on PATH:

sh examples/merge-assumptions/run.sh

Each branch's tests and measurement checks pass. Git merges the branches without a text conflict, and the combined branch tests still pass. Measurement then fails the recorded condition. The runner also uses a separate integration probe that detects each problem; these examples show a gap in those branch tests, not a limit on what tests can express.

The recipes connect explicit inputs in the merged tree to the record: the search visibility switch, the storage policy and the promise in an email template. remeasure --run tests changed readings as a hypothesis without silently rewriting the canonical record. A plain check cannot see a tree change while its recorded input remains unchanged.

There are three useful checks:

  • Across branches during work: kpop consolidate --dry-run --from <branch-or-ref> reads another branch's committed record as proposed changes and tests it against the current record.
  • On the proposed merge result in CI: kpop check checks the combined record; kpop consolidate --dry-run also tests the hypotheses stored beside it.
  • Against the actual tree: kpop remeasure --run runs the deliberately configured recipes and tests their readings against the declared conditions.

A changed premise that needs review can be reported without failing CI; an affected hypothesis still needs review before it can be folded. Decisions are revised explicitly: a verdict laid over a standing judgment folds only when the record's own condition has broken it, or when a person names it (--take), and what it replaced stays beside the record.

These commands already run in kpopper's own CI workflow on pull requests and pushes to main, alongside the test suite. Record checks use the record's declared interpretation; the experimental Hub has a separate HTML verification step. See Add reasoning checks to CI for setup, including measurement recipes.

The coverage is what the record declares. These checks do not infer intent from arbitrary code or prove that all goals are mutually compatible. Keep relevant readings current and review measurement recipes as code. A contradiction expressed only in prose, or hidden behind unrelated IDs, can still require human review.

A third brain for work in progress

An agent's working instructions and the work it produces serve different readers. Plan section numbers can leak into code comments; a website can start describing the prompt that produced it. The useful boundary resembles the fourth wall: keep production instructions in the working context, and put what the audience needs in the finished work. A report may still need its assumptions, evidence and sources.

kpopper provides a separate place to preserve decisions and their reasons across sessions. The agent still has to respect that boundary; the record does not enforce it.

The excitement around building an organizational second brain is well deserved. A team's knowledge already lives across notes and agent memory, documents and research, conversations, plans and commitments, code and data. An agent can connect the relevant pieces into a shared, evolving picture of the work.

Andrej Karpathy's LLM Wiki describes how an agent can maintain that picture as a persistent wiki: synthesizing sources, surfacing contradictions and revisiting stale claims.

As that knowledge becomes a basis for action, its reasoning deserves an explicit record: why a conclusion was accepted, which sources and assumptions support it, and what would call for reconsideration.

kpopper gives that record structure. It connects conclusions to their grounds, preserves what they were reviewed against, and checks declared conditions as recorded facts change. We call this reasoning and review layer a third brain; the agent supplies the interpretation.

Role Question it helps answer
You What matters, and what should we do?
Your second brain: notes, documents and saved knowledge What have we learned and kept that can help?
kpopper, working with your agent What supports this decision, what has changed, and what needs review?

That distinction is useful when a perfectly retrievable note contains a decision whose premises have expired. Finding the note is one job; noticing that its recommendation needs another look is another.

There is a loose parallel with human memory: remembering can involve updating what was previously learned. In a laboratory study of episodic memory, reminders led participants to incorrectly include newly learned items when recalling an earlier list. Hupbach et al., 2007 provide one concrete example. This motivates an analogy, not a claim that kpopper models the brain or that neuroscience validates the product.

Operationally, the analogy is straightforward: retrieve the relevant context, compare it with new information, draw attention to a consequential mismatch, and review the conclusion. In kpopper those steps are explicit records and checks. The person or agent supplies the interpretation; the software follows the declared connections. You retain the decision.

The Lean proof assistant: from Fermat to agents

Lean programming language and proof assistant logo, with its trademark symbol.

In September 2026, Anthropic reported that Claude had formalized a proof of Fermat's Last Theorem in the Lean programming language, producing a complete computer-checked proof. The achievement was formalizing existing mathematics, building on Wiles's proof and community work. See Anthropic's account and the published proof.

Lean is part of kpopper's packaged reasoning runtime. Ordinary use needs no Lean installation. Its kernel checks formal proofs; compiled programs perform the supported calculations and checks. These serve related but distinct purposes:

  • Deterministic computation. The default reasoning core evaluates supported arithmetic, compound conditions and collection queries from explicit recorded inputs.
  • Specific formal guarantees. Named theorems cover defined properties of the evaluator and checked-session core. For example, successful evaluation of closed rational arithmetic agrees with its mathematical meaning, and an accepted session view preserves declared conflict signals. See the arithmetic proof scope and the checked-session guarantees.
  • An optional checked view. Experimental checked sessions expose a revision-bound view through CLI or MCP. A changed record rejects reads using the old revision. Enable this mode separately with kpop session enable; setup and compatibility instructions cover the packaged bundles.

These guarantees do not establish that a source is accurate, that the recorded premises logically imply an agent's verdict, or that an action is permitted. The adapters, renderer, compiler and runtime are not covered by an end-to-end formal proof. You do not need to write Lean to maintain a record. Lean's reference explains the distinction between proof checking and compiled execution.

Logo source and trademark information.

Experimental applications

kpopper's core keeps claims, their sources and dependencies, and the conditions that make decisions worth revisiting. Optional applications build on that core:

Application What it produces Status
kpopper Hub (hub) A browsable snapshot of the record, with layouts and an interactive graph. Experimental
Annotated Documents (annotated-doc) A standalone document with selected evidence and reviewable copy updates. Experimental

Release bundles include the compiled HTML applications. Request them explicitly; installation alone does not activate either application.

kpop experimental hub --open
kpop experimental annotated-doc guide

Check the changelog when using an older installation. Request these applications explicitly or give the agent a standing preference. Their interfaces and artifact formats may change. Ordinary installation, record checks and session hooks work without activating the HTML applications.

Experimental Annotated Documents application: the Autumn Garden Workshop report with an evidence card beside its registration-window passage. The author's interpretation is labelled Not checked and linked to the source notes.

See installation, boundaries and maturity, Annotated Documents, and kpopper Hub.

Share a focused excerpt

To share a small part of the record in a task, document or pull request, use kpop export. The excerpt separates readings captured at review from current recorded readings and marks values omitted from the selection. From a source checkout:

kpop export launch.announcement --record examples/launch-party/GROUNDING.yaml

Add --format markdown-mermaid to keep the text and append a diagram for destinations that support Mermaid. See focused exports for selection limits, original-field details and output options.

What is available, and what is next

This table describes the current repository. Check the changelog when updating an older installation; a merged feature may still be awaiting a release.

Use the native CLI, plugins and CI for the new default history-backed records. The reasoning runtime is packaged; existing legacy records require explicit adoption.

Legacy records with structured expressions and computed snapshots require at least 0.9.0 in every reader and writer that touches them; see reader compatibility. Upgrade every reader and writer together. On native Windows, legacy records support individual add/set writes and deliberate reviews; durable report batching requires POSIX file locking.

New history-backed records also require POSIX locking for creation and writes. On native Windows, create and author history on a POSIX host such as WSL; read-only assessment and the packaged Windows reasoning runtime remain available.

Layer Status Capability
Core Available YAML records, source references, judgment checks, dependency tracing, review snapshots, hypotheses and consolidation.
Core Available CLI and focused Markdown exports with optional Mermaid.
Application Experimental, optional kpopper Hub and Annotated Documents, with selected evidence and copy updates.
Core Available Checks on combined records and hypotheses in CI, including before-merge inspection of another branch's record.
Workflow Experimental; default for new Git projects Advanced mode: branch records and shared findings. Record merge conflicts and repeated CI remain an open design problem.
Integration Available within stated limits Background processing of explicit reports and selective delivery of important findings.
Integration Platform import route documented; runtime not yet validated ChatGPT Work installation and execution of this plugin.
Integration Experimental, opt-in Lean-checked session views, revision-bound reads and a project-bound MCP server.
Core Default for new records Deterministic core/v1 assessment and immutable history, with a packaged arithmetic runtime and automatic reader selection. Existing legacy records require explicit adoption.
Integration Available through the agent Guided starting choices: learn during ordinary work, map selected existing materials, or investigate a defined subject and period in depth.
Integration Available within host limits Optional first-use explanations, workspace guidance and the ability to skip or turn guidance off.

Mapping runs in the calling agent session, using its available tools and the sources you authorize. It does not install connectors or scan accounts by itself. See starting a knowledge record for the workflow and host requirements.

Popper: the philosopher behind falsifiability

No, he isn't really the father of K-pop. :) Karl Popper was a philosopher of science who argued that a scientific theory must be falsifiable: there must be a possible observation or experiment that would contradict it.

So let's put our own claim to the test: what evidence would make us give it up? If “father of K-pop” means Popper helped create the genre, we can investigate that: check his biography, look for songwriting and recording credits, and compare the claim with histories of Korean popular music.

But suppose we keep changing what “father” means whenever the evidence disagrees. Whatever we find, we invent another explanation that saves the claim. Now it can survive anything. We've made it impossible to test.

That is the challenge Popper posed to scientific theories: they must risk being wrong. A theory must rule something out—something an observation or experiment could reveal. If every possible result can be made to fit, it fails his criterion for empirical science.

For the historical comparison, the Library of Congress timeline traces Korean popular music and identifies the emergence of Seo Taiji and Boys as a turning point in the 1990s. Popper's biography tells the story of his work in philosophy. The banner's pop-star story is our invention.

Surviving tests does not prove a theory true, and an apparent counterexample still needs careful checking. The Stanford Encyclopedia of Philosophy explains this distinction between falsifiability and refutation in practice.

kpopper borrows that discipline to build a working picture of reality for your project. It connects claims and decisions to their evidence and assumptions, and preserves what would make us change that picture. When new evidence is recorded, its checks show which declared conditions fail and which decisions need another look. The person or agent interprets the evidence and revises the record. The picture stays open to criticism, refutation and revision as the work changes.

“Ready to launch” had a condition

In our fictional launch story, Karl Popper is releasing his debut K-pop single. An agent has prepared Friday's launch-party announcement. The venue has confirmed the booking. The decision: the announcement is ready, provided the booking stays confirmed.

The next day, the venue cancels. A later session picks up the launch plan. The announcement copy is unchanged; the reason it was ready to publish has disappeared.

kpopper preserves that connection:

known:
  venue.status: {v: confirmed, from: booking, as_of: "2026-09-09"}

judgments:
  launch.announcement:
    rests_on: [venue.status]
    verdict: "Friday's announcement is ready, provided the venue booking stays confirmed."
    because: "The announcement names the date and venue confirmed in the booking email."
    wrong_if: 'venue.status != "confirmed"'
    seen: {venue.status: confirmed}

sources:
  booking:
    name: "Venue confirmation"
    quoted: "Your booking for Friday is confirmed."
    read: "2026-09-09"

After the session records the cancellation, the next check reports:

launch.announcement: wrong_if holds (venue.status != "confirmed") - broken by its own condition

The next agent sees which decision needs review, which premise changed, and what the earlier decision was based on. It knows to revisit the announcement before reusing “ready to launch.”

Try the example · See a PR and CI case

What the check establishes

The launch example gives the announcement a concrete condition: the venue booking must stay confirmed. It is this declared condition that the check tests.

A green check means no failing condition was found by these checks; it does not establish that the recommendation is true, wise, complete or authorized.

This legacy example uses one comparison over declared inputs. The current core also supports compound conditions. Free-form reasoning still needs interpretation. If a condition cannot yet be checked, blocked_on records why. A decision that needs a person's judgment can instead carry reopened_by, describing the sign that would bring it back for review. Preferences and open questions need no invented scientific certainty. See the checking rules.

This is a practical use of falsification, not an automated implementation of the scientific method. Choosing good evidence and meaningful breaking conditions remains intellectual work.

kpopper. An ink illustration of Karl Popper holds a microphone and makes a finger heart, looking toward the wordmark and fictional quotation: It really whips the lemma's ass! Logic symbols rise from the blue word lemma. A separate italic attribution reads - Karl Popper, the father of K-pop. His speech bubble says OMG 이건 꼭 필요해! — roughly, OMG, I really need this!

Wait, what? kpopper is Karl Popper? Is he really the father of K-pop? Who is this guy?

Make it earn its place

Choose a starting point for your project: Coding agents · Claude Cowork / ChatGPT Work · Research

Try it on work you will revisit. After several sessions, ask whether returning takes less reconstruction, whether a changed premise surfaced a useful question, and whether the record costs less to maintain than it saves. Those are outcomes to measure in your work, not established productivity results.

Keep the record as small as the work allows. Its purpose is to help you move the project forward with reasons you can inspect and revise.

Try it on one project you will return to. Install kpopper, or run the merge examples before installing a plugin. If it helps, star the repository and tell us what changed in your work. Use the issue chooser to report a bug, suggest an improvement or ask a question. Documentation fixes, small reproducible examples and reports from different agent hosts are useful contributions; see Contributing for setup and checks.

kpopper is maintained by Ilan Bar Magen. The package is marked beta; the availability table describes current limits. Support and reviews depend on maintainer availability. Community participation follows the Code of Conduct; use the security policy for private vulnerability reports.

kpopper's own code is MIT-licensed. The packaged native runtime includes components with their own license notices and source and replacement instructions.

Command and storage reference · Contributing and validation · Changelog · Security · MIT license

About

kpopper is a checkable knowledge record for AI-assisted projects: decisions linked to their evidence, assumptions and falsifiers.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

12 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages