Skip to content

bug(telemetry): the sharing heartbeat's 24h timer restarts with the backend, so an install that restarts daily shares once and never again #2618

Description

@webmixgamer

Summary

The Tier-2 sharing heartbeat sleeps TELEMETRY_SHARING_INTERVAL_HOURS (24h) plus jitter from process start and then sends unconditionally. It never consults telemetry_sharing_last_shared_at, so a backend restart inside the window discards the elapsed time and restarts the countdown, and the missed send is never made up. An install that restarts more often than once a day delivers its consent-time backfill and then goes permanently silent, while Settings keeps reporting sharing as on and the last send as acknowledged.

Context

Found while verifying end to end that a consenting instance's shares reach the hosted receiver, after the benchmark read (abilityai/trinity-enterprise#190) landed.

The two layers disagree about what "a day" means. _resolve_window already reasons from last_shared_at and deliberately produces a cumulative, gap-free window ("a heartbeat covers everything since the last successful share"), so the payload layer is correct — a late send loses no data. The scheduler layer reasons from boot instead, so the send it would have described never happens.

Observed on one instance: consent granted, backfill delivered and acknowledged, no heartbeat since. Its backend restarts well inside every 24h window (a dev stack runs uvicorn with --reload, so an edit is enough), which is sufficient on its own to explain the silence. Any host that reboots nightly, or an operator who updates Trinity more often than daily, is in the same position.

Consequences:

  • Restart-heavy installs contribute one snapshot at consent time and then read as churned rather than active.
  • The fleet benchmark counts participants who shared inside a rolling window; a silent participant ages out, so the benchmark degrades toward not_enough_data and share-stale for everyone still sharing.
  • A failed send has the same shape: last_shared_at is written only on a 2xx, so a receiver outage costs a full interval rather than one wake.

Acceptance Criteria

  • The heartbeat decides whether to send by comparing telemetry_sharing_last_shared_at against the configured interval, not by counting from process start — a restart no longer resets the cadence.
  • The loop wakes more often than the send cadence, so an overdue install sends shortly after boot instead of a full interval later.
  • The "no boot burst" property is preserved and asserted by a test: nothing is sent during startup; the earliest possible send is one wake interval in.
  • An install with no successful share yet (empty last_shared_at, e.g. the consent-time backfill failed) is treated as due at the first wake, so the owed backfill retries rather than waiting a full interval.
  • A non-2xx or transport failure is retried at the next wake; last_shared_at continues to be written only on a genuine 2xx.
  • Single-flight across uvicorn workers is preserved: the tick marker's TTL keys off the send cadence, not the shortened wake interval, so two workers cannot both send inside one window.
  • Consent gate, hard-disable gate (TELEMETRY_SHARING_ENABLED / DO_NOT_TRACK) and the Redis fail-open behaviour are unchanged — this fix must not widen egress.
  • No new operator setting: TELEMETRY_SHARING_INTERVAL_HOURS keeps its 24h default and stays the only cadence knob.
  • Regression test covering: due (last share older than the interval) → sends; not due → skips; a simulated mid-window restart → still sends when due; consent off → sends nothing regardless of dueness.

Technical Notes

  • src/backend/services/telemetry_sharing_service.py
    • TelemetrySharingService._loop — the unconditional await asyncio.sleep(self.interval_seconds + jitter) followed by share_now(backfill=False); this is the defect.
    • _claim_tick — the cross-worker tick marker, TTL max(interval_seconds // 2, 60); its TTL must follow the send cadence if the wake interval shrinks.
    • share_now — writes telemetry_sharing_last_shared_at only inside the 2xx branch.
    • _resolve_window / _days_since — already correct; the fix should make the scheduler agree with them.
  • src/backend/config.py — TELEMETRY_SHARING_INTERVAL_HOURS, default 24.
  • src/backend/main.py — the staggered +9s service start.
  • Tests: tests/unit/test_ent12_telemetry_sharing.py, tests/unit/test_ent437_telemetry_consent.py.
  • Docs: docs/memory/feature-flows/telemetry-sharing.md; docs/user-docs/operations/telemetry.md if the cadence wording changes.
  • Related: bug(telemetry): the send log does not record where a share went, so a test-receiver 200 reads as a production acknowledgement #2571 (same send-log surface, different defect).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions