Summary
The #389 sync-health poller calls each agent's /api/git/status every 60 s. That endpoint runs git status --porcelain (takes .git/index.lock while refreshing the index) and then git fetch origin — against the agent's live workspace, outside _REPO_LOCK. Every agent that runs git inside its own turn is therefore racing a platform process it cannot see, roughly twice a minute, on every minute of every run. When the poller's git child is killed mid-window (execution end, container recreation, redeploy), a 0-byte .git/index.lock is orphaned and every subsequent git write in that workspace fails until a human or the agent clears it.
Root-caused first-hand today inside the receptionist agent container while investigating a recurring orphaned lock.
Evidence
python3 /app/agent-server.py (PID 52 in the container). Attribution to the sync-health poll is confirmed: get_git_sync_state returned last_check_at: 2026-09-12T17:15:55.422Z — the same second the observed lock window closed.
Lock file sampled at 20 Hz over 120 s in the receptionist workspace:
LOCK APPEARED 17:14:50.582Z / GONE 17:14:50.916Z (334 ms)
LOCK APPEARED 17:14:52.216Z / GONE 17:14:52.276Z ( 60 ms)
LOCK APPEARED 17:15:53.143Z / GONE 17:15:53.616Z (473 ms)
LOCK APPEARED 17:15:54.902Z / GONE 17:15:54.958Z ( 56 ms)
samples=2400 lock_present=16 (0.67% of wall clock, 2 windows/min)
Two overlapping git fetch origin from the same poller 0.9 s apart were also observed — the poll is not guarded against re-entry (and /api/git/status is additionally reachable from the UI and from the get_git_status MCP tool, so concurrent callers stack).
Code path
src/backend/services/sync_health_service.py — DEFAULT_POLL_INTERVAL = 60, _poll_cycle() over every git-enabled agent, reading /api/git/status.
docker/base-image/agent_server/routers/git.py:787 — git status --porcelain (index-lock-taking).
docker/base-image/agent_server/routers/git.py:~1325 — git fetch origin on the same request.
- Contrast: the auto-sync heartbeat in the same module is serialized via
_REPO_LOCK and routes through run_registered (the sweep-safe path). The read-only status endpoint does neither.
Impact
- Silently blocks agent commits. The
receptionist hit an orphaned lock 5 times in ~18 hours; each occurrence blocked its dashboard commits with no signal to the agent or the operator.
- This is fleet-wide by construction — the poller runs against every git-enabled agent. It very likely presents elsewhere as unexplained "git is stuck" / stalled-commit reports.
- The 60 s poll also makes the collision window recur on every minute of every long-running turn, so the odds scale with execution length.
Suggested fixes (in preference order)
git --no-optional-locks status --porcelain for the health poll. This is precisely what the flag exists for: a read-only status that does not take .git/index.lock. Applies to the /api/git/status read path (the auto-sync commit path legitimately needs the index).
- Guard the poller against re-entry — one in-flight
/api/git/status per agent; the overlapping-fetch observation says there is no such guard today.
- Startup reaper — on agent-server startup, clear a 0-byte
index.lock that no live process holds and that is older than a threshold. (Adjacent to the existing _reap_stale_git_litter, which does not cover this case for the status path.)
(1) alone removes the contention; (3) cleans up the installed base of already-orphaned locks.
Acceptance criteria
Related
abilityai/trinity#1561 (sync-health poller hammering dead agents) — same poller, different failure.
abilityai/trinity#1595 (auto-gc killed inside containers) — same "platform git child dies mid-operation" family.
abilityai/trinity#1505 (orphan sweep vs. the live agent-server subtree) — the sweep side of the kill that orphans the lock.
Filed by trinity-pm from first-hand evidence gathered by corbin inside the receptionist container (exec YfwjP9cLEgWfMlj8aXNhww), 2026-09-12.
Summary
The
#389sync-health poller calls each agent's/api/git/statusevery 60 s. That endpoint runsgit status --porcelain(takes.git/index.lockwhile refreshing the index) and thengit fetch origin— against the agent's live workspace, outside_REPO_LOCK. Every agent that runs git inside its own turn is therefore racing a platform process it cannot see, roughly twice a minute, on every minute of every run. When the poller's git child is killed mid-window (execution end, container recreation, redeploy), a 0-byte.git/index.lockis orphaned and every subsequent git write in that workspace fails until a human or the agent clears it.Root-caused first-hand today inside the
receptionistagent container while investigating a recurring orphaned lock.Evidence
python3 /app/agent-server.py(PID 52 in the container). Attribution to the sync-health poll is confirmed:get_git_sync_statereturnedlast_check_at: 2026-09-12T17:15:55.422Z— the same second the observed lock window closed.Lock file sampled at 20 Hz over 120 s in the receptionist workspace:
Two overlapping
git fetch originfrom the same poller 0.9 s apart were also observed — the poll is not guarded against re-entry (and/api/git/statusis additionally reachable from the UI and from theget_git_statusMCP tool, so concurrent callers stack).Code path
src/backend/services/sync_health_service.py—DEFAULT_POLL_INTERVAL = 60,_poll_cycle()over every git-enabled agent, reading/api/git/status.docker/base-image/agent_server/routers/git.py:787—git status --porcelain(index-lock-taking).docker/base-image/agent_server/routers/git.py:~1325—git fetch originon the same request._REPO_LOCKand routes throughrun_registered(the sweep-safe path). The read-only status endpoint does neither.Impact
receptionisthit an orphaned lock 5 times in ~18 hours; each occurrence blocked its dashboard commits with no signal to the agent or the operator.Suggested fixes (in preference order)
git --no-optional-locks status --porcelainfor the health poll. This is precisely what the flag exists for: a read-only status that does not take.git/index.lock. Applies to the/api/git/statusread path (the auto-sync commit path legitimately needs the index)./api/git/statusper agent; the overlapping-fetch observation says there is no such guard today.index.lockthat no live process holds and that is older than a threshold. (Adjacent to the existing_reap_stale_git_litter, which does not cover this case for the status path.)(1) alone removes the contention; (3) cleans up the installed base of already-orphaned locks.
Acceptance criteria
.git/index.lockin the agent workspace — re-run the 20 Hz sampler for 120 s and observelock_present=0./api/git/statuscalls for one agent do not produce overlappinggit fetch originchildren.index.lockrecovers without human intervention, and the recovery is observable (log line / sync-state), not silent — Tell the truth about state.Related
abilityai/trinity#1561(sync-health poller hammering dead agents) — same poller, different failure.abilityai/trinity#1595(auto-gc killed inside containers) — same "platform git child dies mid-operation" family.abilityai/trinity#1505(orphan sweep vs. the live agent-server subtree) — the sweep side of the kill that orphans the lock.Filed by
trinity-pmfrom first-hand evidence gathered bycorbininside the receptionist container (execYfwjP9cLEgWfMlj8aXNhww), 2026-09-12.