Summary
Please make the contract for finite, cold-start agentic replay explicit: execute each selected root once, preserve all owned descendants until the graph drains, and anchor dependency delays to captured transport events rather than credit bookkeeping.
This report comes from a downstream AgentX experiment. The measured failures below are from an unpublished local extension, not a reproduction against current upstream main. The upstream request is to carry these regression cases and clock-boundary requirements into supported finite replay. It should not change the existing steady-state or duration-limited modes implicitly.
Revisions and upstream applicability
Inspected upstream: 31a8089712bde563224b4c9a46d09f5b2ce96b3f, on September 22, 2026.
The downstream lineage starts from public SemiAnalysisAI/agentx-harness@60bce6b092641398d443627d31165219233ac441, followed by local changes. Local 3ad7ff9d was the starting checkpoint for this experiment; d8762826 added the finite typed-dependency adapter, 5a32f0bd still used credit-based timing, 6a547dc4 captured HTTP starts, and final 9a04123368ea2a7ec9364842f93a71602115d027 captured body EOF as well. These local hashes identify retained evidence; they are not published upstream revisions.
Current upstream has the relevant mechanisms, but not this complete contract:
ReplayBarrierCoordinator tracks completion barriers; it does not expose the local adapter's typed dispatch/completion delays or per-play time floors.
- Overlapping branches are initiated from credit issuance. A credit can precede HTTP start by request materialization, serialization and scheduling work.
- Return-driven dependency completion follows awaited bookkeeping; ordinary next-turn delays are also scheduled from that return path. The transport already captures request start and body EOF, but those timestamps do not travel through the dependency coordinator's credit-return path.
- Profiling send completion cancels pending scheduler/gate work. That cutoff behavior needs a separate decision for a finite whole-graph mode.
- Upstream already has an overlap-dispatch deduplication marker and parent lock. I am not claiming that current upstream generally duplicates descendants. Cold-start snapshot seeding, owner dispatch and return-time spawning need to remain jointly covered when introducing a finite adapter.
Related: #1231 covers a dropped post-child residual delay; #1226 proposes a graph runtime; #1424 discusses client start/TTFT boundaries. This issue concerns finite graph drain and the transport-event anchors, rather than asserting those reports are duplicates or that the unmerged graph runtime has these defects.
Small synthetic reproducer / acceptance fixture
Use two root lanes, sequential selection, one execution per root, start ratios 0/0, no cache warmup, no request/duration cutoff, and one worker. The entire invented graph below has nine requests and three child sessions. Every request has 128 integer prompt tokens; output lengths are 2 through 10 in the listed order. No production trace is needed.
Times are milliseconds. A request's deadline is the maximum, not the sum, of its root-relative floor and each predecessor's captured event plus that edge's delay.
| Request |
Session |
Root-relative floor |
Dependencies |
Dummy content duration |
| A0 |
A root |
0 |
none |
200 |
| A1 |
A root |
650 |
A0 EOF + 30; C1 EOF + 40 |
100 |
| A2 |
A root |
700 |
A1 EOF + 20 |
40 |
| C0 |
A child |
50 |
A0 HTTP start + 50 |
150 |
| C1 |
A child |
400 |
C0 EOF + 50 |
100 |
| BG |
A background child |
800 |
A1 EOF + 200 |
80 |
| B0 |
B root |
0 |
none |
250 |
| B1 |
B root |
670 |
B0 EOF + 20; D0 EOF + 60 |
50 |
| D0 |
B child |
20 |
B0 HTTP start + 40 |
400 |
The stored fixture expresses B's absolute floors as 80/750/100 and subtracts the root's 80 ms floor; the table uses the equivalent relative floors. Both roots are started independently. Run a second copy of the fixture with 124,480 tokens per prompt to exercise asymmetric preparation delays.
For the exact completion-boundary case, the dummy OpenAI SSE server emits its content, waits 30 ms, flushes usage and [DONE], waits another 100 ms, and only then closes the response body. Capture client start_perf_ns immediately before the HTTP operation and end_perf_ns after consuming the body through EOF. The two values must use the same host's monotonic clock.
The core oracle is:
# nodes: the table above; records: one successful HTTP record per node.
assert sorted(r.id for r in records) == sorted(n.id for n in nodes)
by_id = {r.id: r for r in records}
for node in nodes:
record = by_id[node.id]
assert 0 < record.start_perf_ns <= record.end_perf_ns
due = by_id[node.root_id].start_perf_ns + node.floor_ms * 1_000_000
for edge in node.dependencies:
parent = by_id[edge.parent_id]
anchor = (parent.start_perf_ns if edge.event == "dispatch"
else parent.end_perf_ns)
due = max(due, anchor + edge.delay_ms * 1_000_000)
assert record.start_perf_ns >= due # Zero early tolerance.
assert children_spawned == children_completed == 3
assert errors == cancellations == truncated_children == suppressed_joins == 0
Also check exact prompt digests and server-reported token usage. The background child must finish even when its root's last HTTP request has already completed. A separate two-node cold-start regression gives root and child equal source timestamps and checks that snapshot initialization plus owner dispatch plus parent return produces the child exactly once.
This is a regression fixture specification for the proposed contract, not a claim that upstream currently accepts the downstream typed-graph input format.
Complete synthetic input used by the downstream fixture
a:root:000/001/002 correspond to A0/A1/A2, a:child:000/001 to C0/C1, a:background:000 to BG, and the B-prefixed records follow the same naming convention. The adapter deterministically interns the invented block hashes in lexical request-ID order, repeats each compact token ID 64 times, and truncates to input_length. The large-prompt variant keeps this graph shape and uses additional invented hash blocks sufficient for 124,480 tokens. The format marker identifies that adapter's schema; it does not imply an upstream CLI feature.
{"schema":"dynamo.agentic_mooncake","version":2,"block_size":64,"hash_id_scope":"local","source":{"format":"synthetic-contract-test","digest":"fixture-v2"}}
{"request_id":"a:root:000","play_id":"a","session_id":"a:root","model":"fixture-model","input_length":128,"output_length":2,"hash_ids":[1,100],"not_before_ms":0,"dependencies":[]}
{"request_id":"a:root:001","play_id":"a","session_id":"a:root","model":"fixture-model","input_length":128,"output_length":3,"hash_ids":[1,101],"not_before_ms":650,"dependencies":[{"request_id":"a:root:000","trigger":"completion","delay_ms":30,"relation":"sequence"},{"request_id":"a:child:001","trigger":"completion","delay_ms":40,"relation":"join"}]}
{"request_id":"a:root:002","play_id":"a","session_id":"a:root","model":"fixture-model","input_length":128,"output_length":4,"hash_ids":[1,102],"not_before_ms":700,"dependencies":[{"request_id":"a:root:001","trigger":"completion","delay_ms":20,"relation":"sequence"}]}
{"request_id":"a:child:000","play_id":"a","session_id":"a:child","model":"fixture-model","input_length":128,"output_length":5,"hash_ids":[1,103],"not_before_ms":50,"dependencies":[{"request_id":"a:root:000","trigger":"dispatch","delay_ms":50,"relation":"spawn"}]}
{"request_id":"a:child:001","play_id":"a","session_id":"a:child","model":"fixture-model","input_length":128,"output_length":6,"hash_ids":[1,104],"not_before_ms":400,"dependencies":[{"request_id":"a:child:000","trigger":"completion","delay_ms":50,"relation":"sequence"}]}
{"request_id":"a:background:000","play_id":"a","session_id":"a:background","model":"fixture-model","input_length":128,"output_length":7,"hash_ids":[1,105],"not_before_ms":800,"dependencies":[{"request_id":"a:root:001","trigger":"completion","delay_ms":200,"relation":"spawn"}]}
{"request_id":"b:root:000","play_id":"b","session_id":"b:root","model":"fixture-model","input_length":128,"output_length":8,"hash_ids":[2,106],"not_before_ms":80,"dependencies":[]}
{"request_id":"b:root:001","play_id":"b","session_id":"b:root","model":"fixture-model","input_length":128,"output_length":9,"hash_ids":[2,107],"not_before_ms":750,"dependencies":[{"request_id":"b:root:000","trigger":"completion","delay_ms":20,"relation":"sequence"},{"request_id":"b:child:000","trigger":"completion","delay_ms":60,"relation":"replay_barrier"}]}
{"request_id":"b:child:000","play_id":"b","session_id":"b:child","model":"fixture-model","input_length":128,"output_length":10,"hash_ids":[2,108],"not_before_ms":100,"dependencies":[{"request_id":"b:root:000","trigger":"dispatch","delay_ms":40,"relation":"spawn"}]}
Observed versus expected
Expected: all nine requests execute once, every child drains, and no request starts before its latest controlling deadline. Delayed delivery of a transport notification may make dispatch late; it must not redefine the recorded event time or add another tool delay.
The preserved real multiprocess AIPerf-to-dummy-HTTP/SSE runs show:
| Local revision |
Fixture |
Outcome |
5a32f0bd |
128-token prompts, credit anchors |
FAIL: earliest controlling deadline residual −4.637 ms |
5a32f0bd |
124,480-token prompts, credit anchors |
FAIL: earliest residual −6.020 ms |
9a041233 |
128-token prompts, captured HTTP start/EOF |
PASS: all nine requests once; nonroot lateness 0.615–1.290 ms |
9a041233 |
124,480-token prompts, captured HTTP start/EOF |
PASS: all nine requests once; nonroot lateness 0.743–16.801 ms |
The failed runs already had complete descendant drain; they expose timing errors separately from lifecycle errors. Their completion oracle used last-SSE time only as a lower bound, so they do not establish exact EOF fidelity. The final fixtures exported exact EOF and measured it 100.32–106.05 ms after the last raw SSE event, deliberately preventing a last-content or [DONE] timestamp from passing as body EOF.
The cold-start and finite-drain regressions cover delayed/background descendants and duplicate-offer hazards in the local adapter. They do not establish an unmodified-upstream drop/duplicate rate.
Validated downstream approach
- Make root-admission completion distinct from graph completion. Retain delayed scheduler/gate work until all admitted trees drain; refresh sent counts when descendant credits are issued. Explicit cancellation and timeouts still cancel.
- For cold start, offer the root at turn zero and let ownership edges spawn descendants once. Preserve the existing parent lock/deduplication behavior.
- Carry the captured HTTP-start timestamp on the existing credit channel without awaiting notification I/O before launching HTTP. Carry captured body EOF with the terminal credit before result-forwarding awaits.
- Use absolute monotonic deadlines; timers only wake/recheck readiness. A late notification cannot release work before notification receipt, but should not shift the authored deadline. Preserve first-observation idempotence, reject late notifications for closed roots, and fail visibly on missing/invalid/error/cancelled terminal signals in strict finite mode.
- Export exact EOF separately. Existing last-content latency and last-raw-response fields keep their existing meanings. Client HTTP start is not server admission, and client EOF is not the engine's final token.
The final local revision passed 537 focused tests plus both multiprocess fixtures on a remote Linux CPU host. The fixtures shared one physical CPU deliberately to stress scheduling and preparation delay. These are functional results, not engine performance measurements. This report reused those results and performed current-upstream static inspection; it did not run a new upstream test suite or CI job. No GPU or customer trace was required.
I would keep this as an opt-in finite replay contract, or use the proposed graph runtime if it supplies the same guarantees, rather than silently changing AgentX's existing steady-state timing policy.
Summary
Please make the contract for finite, cold-start agentic replay explicit: execute each selected root once, preserve all owned descendants until the graph drains, and anchor dependency delays to captured transport events rather than credit bookkeeping.
This report comes from a downstream AgentX experiment. The measured failures below are from an unpublished local extension, not a reproduction against current upstream
main. The upstream request is to carry these regression cases and clock-boundary requirements into supported finite replay. It should not change the existing steady-state or duration-limited modes implicitly.Revisions and upstream applicability
Inspected upstream:
31a8089712bde563224b4c9a46d09f5b2ce96b3f, on September 22, 2026.The downstream lineage starts from public
SemiAnalysisAI/agentx-harness@60bce6b092641398d443627d31165219233ac441, followed by local changes. Local3ad7ff9dwas the starting checkpoint for this experiment;d8762826added the finite typed-dependency adapter,5a32f0bdstill used credit-based timing,6a547dc4captured HTTP starts, and final9a04123368ea2a7ec9364842f93a71602115d027captured body EOF as well. These local hashes identify retained evidence; they are not published upstream revisions.Current upstream has the relevant mechanisms, but not this complete contract:
ReplayBarrierCoordinatortracks completion barriers; it does not expose the local adapter's typed dispatch/completion delays or per-play time floors.Related: #1231 covers a dropped post-child residual delay; #1226 proposes a graph runtime; #1424 discusses client start/TTFT boundaries. This issue concerns finite graph drain and the transport-event anchors, rather than asserting those reports are duplicates or that the unmerged graph runtime has these defects.
Small synthetic reproducer / acceptance fixture
Use two root lanes, sequential selection, one execution per root, start ratios
0/0, no cache warmup, no request/duration cutoff, and one worker. The entire invented graph below has nine requests and three child sessions. Every request has 128 integer prompt tokens; output lengths are 2 through 10 in the listed order. No production trace is needed.Times are milliseconds. A request's deadline is the maximum, not the sum, of its root-relative floor and each predecessor's captured event plus that edge's delay.
The stored fixture expresses B's absolute floors as 80/750/100 and subtracts the root's 80 ms floor; the table uses the equivalent relative floors. Both roots are started independently. Run a second copy of the fixture with 124,480 tokens per prompt to exercise asymmetric preparation delays.
For the exact completion-boundary case, the dummy OpenAI SSE server emits its content, waits 30 ms, flushes usage and
[DONE], waits another 100 ms, and only then closes the response body. Capture clientstart_perf_nsimmediately before the HTTP operation andend_perf_nsafter consuming the body through EOF. The two values must use the same host's monotonic clock.The core oracle is:
Also check exact prompt digests and server-reported token usage. The background child must finish even when its root's last HTTP request has already completed. A separate two-node cold-start regression gives root and child equal source timestamps and checks that snapshot initialization plus owner dispatch plus parent return produces the child exactly once.
This is a regression fixture specification for the proposed contract, not a claim that upstream currently accepts the downstream typed-graph input format.
Complete synthetic input used by the downstream fixture
a:root:000/001/002correspond to A0/A1/A2,a:child:000/001to C0/C1,a:background:000to BG, and the B-prefixed records follow the same naming convention. The adapter deterministically interns the invented block hashes in lexical request-ID order, repeats each compact token ID 64 times, and truncates toinput_length. The large-prompt variant keeps this graph shape and uses additional invented hash blocks sufficient for 124,480 tokens. The format marker identifies that adapter's schema; it does not imply an upstream CLI feature.{"schema":"dynamo.agentic_mooncake","version":2,"block_size":64,"hash_id_scope":"local","source":{"format":"synthetic-contract-test","digest":"fixture-v2"}} {"request_id":"a:root:000","play_id":"a","session_id":"a:root","model":"fixture-model","input_length":128,"output_length":2,"hash_ids":[1,100],"not_before_ms":0,"dependencies":[]} {"request_id":"a:root:001","play_id":"a","session_id":"a:root","model":"fixture-model","input_length":128,"output_length":3,"hash_ids":[1,101],"not_before_ms":650,"dependencies":[{"request_id":"a:root:000","trigger":"completion","delay_ms":30,"relation":"sequence"},{"request_id":"a:child:001","trigger":"completion","delay_ms":40,"relation":"join"}]} {"request_id":"a:root:002","play_id":"a","session_id":"a:root","model":"fixture-model","input_length":128,"output_length":4,"hash_ids":[1,102],"not_before_ms":700,"dependencies":[{"request_id":"a:root:001","trigger":"completion","delay_ms":20,"relation":"sequence"}]} {"request_id":"a:child:000","play_id":"a","session_id":"a:child","model":"fixture-model","input_length":128,"output_length":5,"hash_ids":[1,103],"not_before_ms":50,"dependencies":[{"request_id":"a:root:000","trigger":"dispatch","delay_ms":50,"relation":"spawn"}]} {"request_id":"a:child:001","play_id":"a","session_id":"a:child","model":"fixture-model","input_length":128,"output_length":6,"hash_ids":[1,104],"not_before_ms":400,"dependencies":[{"request_id":"a:child:000","trigger":"completion","delay_ms":50,"relation":"sequence"}]} {"request_id":"a:background:000","play_id":"a","session_id":"a:background","model":"fixture-model","input_length":128,"output_length":7,"hash_ids":[1,105],"not_before_ms":800,"dependencies":[{"request_id":"a:root:001","trigger":"completion","delay_ms":200,"relation":"spawn"}]} {"request_id":"b:root:000","play_id":"b","session_id":"b:root","model":"fixture-model","input_length":128,"output_length":8,"hash_ids":[2,106],"not_before_ms":80,"dependencies":[]} {"request_id":"b:root:001","play_id":"b","session_id":"b:root","model":"fixture-model","input_length":128,"output_length":9,"hash_ids":[2,107],"not_before_ms":750,"dependencies":[{"request_id":"b:root:000","trigger":"completion","delay_ms":20,"relation":"sequence"},{"request_id":"b:child:000","trigger":"completion","delay_ms":60,"relation":"replay_barrier"}]} {"request_id":"b:child:000","play_id":"b","session_id":"b:child","model":"fixture-model","input_length":128,"output_length":10,"hash_ids":[2,108],"not_before_ms":100,"dependencies":[{"request_id":"b:root:000","trigger":"dispatch","delay_ms":40,"relation":"spawn"}]}Observed versus expected
Expected: all nine requests execute once, every child drains, and no request starts before its latest controlling deadline. Delayed delivery of a transport notification may make dispatch late; it must not redefine the recorded event time or add another tool delay.
The preserved real multiprocess AIPerf-to-dummy-HTTP/SSE runs show:
5a32f0bd5a32f0bd9a0412339a041233The failed runs already had complete descendant drain; they expose timing errors separately from lifecycle errors. Their completion oracle used last-SSE time only as a lower bound, so they do not establish exact EOF fidelity. The final fixtures exported exact EOF and measured it 100.32–106.05 ms after the last raw SSE event, deliberately preventing a last-content or
[DONE]timestamp from passing as body EOF.The cold-start and finite-drain regressions cover delayed/background descendants and duplicate-offer hazards in the local adapter. They do not establish an unmodified-upstream drop/duplicate rate.
Validated downstream approach
The final local revision passed 537 focused tests plus both multiprocess fixtures on a remote Linux CPU host. The fixtures shared one physical CPU deliberately to stress scheduling and preparation delay. These are functional results, not engine performance measurements. This report reused those results and performed current-upstream static inspection; it did not run a new upstream test suite or CI job. No GPU or customer trace was required.
I would keep this as an opt-in finite replay contract, or use the proposed graph runtime if it supplies the same guarantees, rather than silently changing AgentX's existing steady-state timing policy.