Skip to content

Repository files navigation

TDAW-BFL Lab

A reproducible research lab for studying history-aware adaptive weighting in Byzantine-robust federated learning.

The project trains a small MNIST classifier across multiple logical clients, injects controlled model-poisoning attacks, and compares standard FedAvg with Temporal-Difference Adaptive Weighting (TDAW). The custom detection, trust, clipping, and aggregation path is implemented in Rust; PyTorch and Flower provide training and orchestration; a Rust control API and Next.js dashboard make experiments inspectable.

Important

The primary research modes are not privacy-preserving: they expose individual client model updates to the server and do not provide secure aggregation or production security. Flower SecAgg+ is implemented as a separate aggregate-only comparison and cannot run server-side TDAW scoring.

Tip

New to the repository? Start with Setup and operation. If you are configuring Flower, keep the complete Flower deployment guide open as the operator reference. If you just ran mise run purge-experiments, use the post-purge restart; do not rotate Flower credentials.

Why this project exists

Current-round anomaly detectors can react quickly to obviously malicious model updates, but they ignore how a client's behavior changes over time. TDAW tests whether a small amount of client history can improve resistance to intermittent or on-off attacks without causing excessive rejection of honest clients.

For each submitted model delta, the server:

  1. extracts an interpretable fingerprint from magnitude, direction, sign, temporal, and optional validation evidence;
  2. normalizes client evidence with robust round statistics based on medians and median absolute deviations;
  3. updates bounded client trust, penalizing deterioration faster than it rewards recovery;
  4. rejects hard failures, assigns capped normalized weights, and clips accepted update norms;
  5. aggregates the weighted model deltas and records explainable per-client decisions.

The study is deliberately narrow: it prioritizes a clear ablation, deterministic execution, and interpretable evidence over large models, broad benchmarks, or claims of universal Byzantine robustness.

System overview

flowchart LR
  C["MNIST clients"] -->|"model deltas"| P["Python orchestration\nlocal runner or Flower"]
  P --> R["Rust TDAW core\nfingerprint · trust · weights · clipping"]
  R -->|"aggregate delta + decisions"| P
  P --> A["Versioned run artifacts\nJSONL · config · summaries"]
  A --> API["Rust control API"]
  API <--> DB[("PostgreSQL metadata")]
  UI["Next.js dashboard"] <--> API
Loading

Rust is the authoritative implementation of the custom algorithm. Python handles tensor conversion, PyTorch training, deterministic data partitioning, attack injection, and Flower integration without duplicating TDAW policy.

Experiment modes

Mode Runtime Server visibility Intended use
local-research Deterministic single Python process Individual client deltas and derived evidence Scientific reference and final experiment matrix
flower-visible Long-lived Flower SuperLink and explicit SuperNodes Individual client deltas and derived evidence Deployment-equivalence and orchestration study
flower-secaggplus Flower SecAggPlusWorkflow and secaggplus_mod Aggregate model update, participant counts, utility, and protocol cost only Privacy/observability trade-off comparison

The control API supervises one local runner or submits a run to an already-running Flower federation. The authenticated Docker fleet has 10 explicitly named SuperNodes; visible-update experiments choose 2 through 10 logical clients, while SecAgg+ experiments choose 3 through 10, using the matching prefix of that long-lived fleet. It never receives the Docker Engine socket. Only the control API receives PostgreSQL credentials; training nodes receive neither database credentials nor shared artifact-write access.

What the repository contains

Path Purpose
crates/tdaw-core Pure Rust fingerprinting, robust statistics, temporal trust, weighting, clipping, and aggregation
crates/tdaw-py Thin PyO3/maturin bridge between NumPy and the Rust core
crates/control-api Axum control API, experiment lifecycle, PostgreSQL metadata, and safe artifact projections
apps/federation PyTorch/MNIST simulator, attacks, deterministic partitions, Flower apps, and artifact recording
apps/web Next.js experiment runboard, round/client inspection, charts, and two-run comparison
configs Committed smoke, deployment, ablation, and final-matrix configurations
schemas Versioned experiment, event, and OpenAPI contracts
deploy PostgreSQL and long-lived Flower Docker Compose profiles
docs Algorithm, architecture, threat model, experiment design, reports, and decisions

Main experimental result

The frozen study ran one IID, 10-client, 20-round on-off sign-flip workload across three seeds. It compared FedAvg, current-round fingerprint weighting, temporal EMA weighting, and full temporal-difference hysteresis. All 12 runs completed and passed artifact verification.

Condition Final accuracy Active-round malicious weight mass
FedAvg 79.52% 20.00%
Fingerprint only 88.54% 0.18%
Temporal EMA 88.53% 0.34%
Full hysteresis 88.52% 0.22%

In this specific workload, every TDAW variant substantially reduced malicious contribution weight relative to FedAvg. Temporal memory did not outperform current-round fingerprint weighting: EMA reacted more slowly at attack onset and carried lower trust into benign rounds, while hysteresis recovered some of that gap without improving final utility. These are descriptive results from three seeds, not a general robustness or significance claim.

Final accuracy and active malicious weight mass across the frozen study

See the generated result tables and workshop report for confidence intervals, round trajectories, runtime measurements, limitations, and related work.

Setup and operation

This section covers the normal local workflow and the default authenticated Flower subprocess profile. The complete Flower deployment guide is the source of truth for credential boundaries, Linux permissions, deployment evidence, the insecure plumbing-only profile, and optional external process isolation.

Choose the setup you need

Goal Required services Start here
Verify the algorithm locally No persistent services Bootstrap and run the smoke experiment
Use the dashboard with local-research PostgreSQL, control API, web app Run the control plane and dashboard
Run flower-visible or flower-secaggplus PostgreSQL, control API, web app, authenticated Flower fleet Provision Flower for the first time
Resume after purge-experiments Same as before the purge; identities are preserved Restart after purging experiments

Prerequisites

  • mise for the pinned Rust, Python, Node.js, uv, and pnpm toolchain;
  • Docker Engine, Docker Desktop, or OrbStack with Docker Compose v2 for PostgreSQL and Flower deployment modes;
  • openssl and ssh-keygen for authenticated Flower credential generation;
  • jq for the documented JSON readiness checks;
  • network access on first setup to install dependencies, pull images, and obtain MNIST.

The optional Flower process isolation profile requires Docker Compose 2.24.4 or newer. On Linux, the Flower bind mounts must be readable by container UID 49999, and .data/runs must be writable by that UID; use the host's normal ownership or ACL policy rather than weakening or committing credentials.

Bootstrap a fresh clone

Install the toolchain and verify the local research path:

mise trust
mise install
mise run bootstrap
mise run doctor
mise run run-smoke

bootstrap installs the locked Python and web dependencies and builds the Rust/Python extension. doctor checks tool versions, imports, contracts, and repository structure. run-smoke downloads MNIST on first use and completes a deterministic five-round local experiment, which is also the recommended data prefetch before starting Flower's read-only MNIST mounts.

Generated data is ignored by Git:

Location Contents
.data/mnist Downloaded MNIST cache
.data/runs Per-experiment configs, events, summaries, and logs
.data/flower Flower CLI config, registrations, identity maps, evidence, and run registry
deploy/flower/secrets Generated local TLS, authentication, and optional AppIO credentials
Docker volume tdaw-control_postgres-data Control-plane experiment metadata
Docker volume tdaw-flower-visible_superlink-state Persistent Flower registrations and SuperLink state

Run the control plane and dashboard

Start these in separate terminals:

# Terminal 1: starts PostgreSQL, runs migrations, and serves 127.0.0.1:8787
mise run api
# Terminal 2: serves the dashboard on 127.0.0.1:3000
mise run web

mise run api already invokes mise run db-up; there is no need to start PostgreSQL separately. mise run dev prints the expected terminal layout without starting long-lived processes.

Verify the control plane:

curl -fsS http://127.0.0.1:8787/healthz | jq
curl -fsS http://127.0.0.1:8787/api/v1/deployment-readiness | jq

Open http://localhost:3000. The dashboard creates experiments from committed presets, checks deployment readiness before starting Flower runs, follows lifecycle state through SSE-backed SWR refreshes, summarizes current-round node milestones and accumulated timing, surfaces bounded backend failure diagnostics, visualizes round-level utility and robustness, drills into per-client evidence, and compares two compatible runs.

Default loopback ports are:

Service Address Override
Dashboard http://127.0.0.1:3000 Next.js CLI/environment
Control API http://127.0.0.1:8787 TDAW_API_ADDR
PostgreSQL 127.0.0.1:5432 TDAW_POSTGRES_PORT or TDAW_DATABASE_URL
Flower SuperLink Fleet API 127.0.0.1:9093 Compose configuration

The local PostgreSQL defaults are user tdaw, database tdaw, and password tdaw-local-only. Set TDAW_POSTGRES_PASSWORD and TDAW_POSTGRES_PORT before the first database start, or set TDAW_DATABASE_URL explicitly when using another database or a password that needs URL escaping.

Important runtime variables are:

Variable Purpose Default or requirement
TDAW_DATA_DIR Shared run, MNIST, and Flower control root mise sets <repo>/.data
TDAW_API_ADDR Control API listen address 127.0.0.1:8787
TDAW_POSTGRES_PASSWORD Local Compose PostgreSQL password tdaw-local-only
TDAW_POSTGRES_PORT Local Compose PostgreSQL host port 5432
TDAW_DATABASE_URL Complete external/overridden API database URL Derived from the two PostgreSQL variables
FLWR_HOME Host Flower CLI config root Export <repo>/.data/flower/flwr-home
TDAW_GIT_COMMIT OCI/runtime source revision Required 40-character lowercase Git commit for Flower Compose
TDAW_GIT_DIRTY OCI/runtime working-tree provenance Required literal true or false for Flower Compose
TDAW_FLOWER_ISOLATION_MODE API-selected Flower isolation profile subprocess; set process only with matching process artifacts
TDAW_FLOWER_DEPLOYMENT_EVIDENCE Explicit deployment-evidence file override Derived from the selected isolation profile

How the Flower fleet works

Flower is long-lived deployment infrastructure, not a container-per-experiment launcher. The default authenticated profile contains one SuperLink and 10 named SuperNodes. Each SuperNode has a stable P-384 key and topology identity:

  • flower-visible chooses a canonical prefix of 2 through 10 nodes;
  • flower-secaggplus chooses a canonical prefix of 3 through 10 nodes;
  • changing an experiment's client count does not create, destroy, or re-register containers;
  • stopping an experiment cancels its Flower run but leaves the shared federation running;
  • subprocess isolation is the default; the separately measured process profile has 22 services and is documented in the Flower operator guide.

The control API never receives the Docker Engine socket. It validates the registered identity map and deployment evidence, obtains the live node listing through Flower's CLI boundary, and only then accepts a Flower start request.

Provision Flower for the first time

Do this after bootstrap and the local smoke experiment. Provisioning is a one-time deployment operation, not something to repeat after every run or restart.

Prepare the shared paths and source provenance used by Compose:

export TDAW_DATA_DIR="$PWD/.data"
mkdir -p "$TDAW_DATA_DIR/runs" \
  "$TDAW_DATA_DIR/flower/run-registry" \
  "$TDAW_DATA_DIR/flower/control"

export TDAW_GIT_COMMIT="$(git rev-parse HEAD)"
if test -z "$(git status --porcelain)"; then
  export TDAW_GIT_DIRTY=false
else
  export TDAW_GIT_DIRTY=true
fi

Generate 10 SuperNode keypairs, TLS/AppIO material, HMAC secrets, and an isolated Flower CLI config:

mise run flower-provision
export FLWR_HOME="$TDAW_DATA_DIR/flower/flwr-home"

The task prints the exact next commands and public-key registration commands. Keep FLWR_HOME exported for every host-side flwr command; this selects the repository-local tdaw-local connection instead of changing ~/.flwr/config.toml.

Validate and build the default authenticated profile. Deployment evidence requires observed build, startup, and readiness durations; do not enter invented values or fold up --build into startup timing. For zsh, the following records the image-build wall time:

zmodload zsh/datetime
started_at=$EPOCHREALTIME

docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  config --quiet

docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  build

export FLOWER_BUILD_MS=$(( (EPOCHREALTIME - started_at) * 1000.0 ))

Start only SuperLink so the public keys can be registered:

docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  up --detach --wait --wait-timeout 300 superlink

Create .data/flower/registrations, then run all 10 flwr supernode register commands printed by flower-provision, from client-00 through client-09. A single command has this shape:

mkdir -p .data/flower/registrations
uv run --project apps/federation --locked \
  flwr supernode register \
  deploy/flower/secrets/supernode-00.pub \
  tdaw-local --format json \
  > .data/flower/registrations/client-00.json

Do not reuse that response for another node. Each file must report "success": true and a unique node-id. After all 10 exist, materialize the trusted topology-to-node mapping:

mise run flower-materialize-map

For an accurate full-fleet startup measurement, stop the registration-only SuperLink without deleting its volume, then start all 10 SuperNodes without --build:

docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  down

started_at=$EPOCHREALTIME
docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  up --detach --wait --wait-timeout 300
export FLOWER_STARTUP_MS=$(( (EPOCHREALTIME - started_at) * 1000.0 ))

Never add --volumes to this routine stop: that destroys SuperLink registration state and invalidates the transport map.

Measure and validate the complete readiness join:

started_at=$EPOCHREALTIME

uv run --project apps/federation --locked \
  flwr supernode list tdaw-local --format json \
  > .data/flower/online-nodes.json

docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  config --format json \
  > .data/flower/compose.auth.resolved.json

uv run --project apps/federation --locked python -m \
  tdaw_fl.flower.deployment_readiness \
  --topology deploy/flower/topology.json \
  --transport-map .data/flower/transport-node-map.json \
  --registration-responses .data/flower/registrations \
  --public-keys deploy/flower/secrets \
  --online-nodes .data/flower/online-nodes.json \
  --compose .data/flower/compose.auth.resolved.json

export FLOWER_READINESS_MS=$(( (EPOCHREALTIME - started_at) * 1000.0 ))

Success prints:

10 authenticated manifest nodes are online and Compose policy is valid

Capture the sanitized image, policy, topology, resource, and timing evidence required by the API:

mise run flower-capture-deployment -- \
  --build-ms "$FLOWER_BUILD_MS" \
  --startup-ms "$FLOWER_STARTUP_MS" \
  --readiness-ms "$FLOWER_READINESS_MS"

This writes .data/flower/deployment-evidence.subprocess.json. The capture command records the durations supplied by the caller; it does not measure them itself. Re-capture after rebuilding images, changing topology/policy, changing the selected isolation profile, or changing source provenance embedded in the images.

Start mise run api and mise run web in separate terminals, then verify the entire fleet contract:

curl -fsS http://127.0.0.1:8787/api/v1/deployment-readiness |
  jq -e '
    .status == "ready"
    and .modes["local-research"].ready == true
    and .modes["flower-visible"].ready == true
    and .modes["flower-secaggplus"].ready == true
    and .flower.isolation_mode == "subprocess"
    and .flower.provisioned_node_count == 10
    and .flower.online_node_count == 10
    and .flower.runnable_client_capacity == 10
    and (.flower.issues | length) == 0
  '

Restart an existing Flower deployment

Routine host restarts and docker compose down preserve credentials, registrations, maps, evidence, and the SuperLink identity volume. Re-export current provenance and restart the services; do not provision again:

export TDAW_DATA_DIR="$PWD/.data"
export FLWR_HOME="$TDAW_DATA_DIR/flower/flwr-home"
export TDAW_GIT_COMMIT="$(git rev-parse HEAD)"
if test -z "$(git status --porcelain)"; then
  export TDAW_GIT_DIRTY=false
else
  export TDAW_GIT_DIRTY=true
fi

docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  up --detach --wait --wait-timeout 300

Then start mise run api and mise run web. Add --build and re-capture deployment evidence only when the application image or its recorded provenance actually changed.

Restart after purging experiments

mise run purge-experiments intentionally stops Flower, deletes the PostgreSQL volume, removes run/sweep/final-matrix artifacts and per-run Flower control records, and recreates empty run-registry directories. It preserves:

  • MNIST and dependency caches;
  • generated Flower credentials and FLWR_HOME;
  • all 10 registration responses and transport maps;
  • the tdaw-flower-visible_superlink-state identity volume;
  • captured deployment evidence and built images.

Before running the purge, stop host mise run api, mise run web, and experiment processes yourself; the task controls Compose services and generated files, not arbitrary host processes.

After the purge, follow only this recovery sequence:

  1. export the provenance variables shown in Restart an existing Flower deployment;
  2. restart authenticated Flower with the two-file docker compose ... up --detach --wait command shown there;
  3. run mise run api—it recreates PostgreSQL and applies migrations automatically;
  4. run mise run web;
  5. check /healthz and /api/v1/deployment-readiness.

Do not run flower-provision --force, re-register keys, re-materialize the map, or recapture unchanged images after an experiment-only purge.

Stop services without deleting state

Stop mise run web and mise run api with Ctrl-C, then stop the persistent services:

mise run db-down

docker compose \
  -f deploy/flower/compose.yaml \
  -f deploy/flower/compose.auth.yaml \
  down

These commands preserve both named volumes and all ignored identity artifacts. The next normal startup uses Restart an existing Flower deployment. Do not use down --volumes unless you are deliberately entering the coordinated identity-reset procedure.

Credential rotation and full reset

mise run flower-provision -- --force --confirm-identity-state-reset is only for an intentional identity rotation, such as migrating a legacy five-node deployment. It is not a repair or restart command. The safety guard requires the old registration responses and transport maps to be archived and removed after the federation is stopped and the tdaw-flower-visible_superlink-state volume is removed. All 10 new keys must then be registered and materialized again.

mise run purge is broader: after confirmation it deletes the PostgreSQL and Flower volumes plus all Git-ignored project state, including MNIST, credentials, dependencies, caches, builds, and experiment artifacts. After a full purge, repeat bootstrap, the local smoke/data prefetch, and the complete first-time Flower provisioning flow.

For the exact coordinated rotation procedure and optional process-isolation artifacts, use the authenticated provisioning guide.

Common setup failures

Symptom Likely cause Action
flower_superlink_unavailable Flower is stopped, FLWR_HOME is wrong, or the Fleet API cannot be reached Export the repository-local FLWR_HOME, start the authenticated Compose profile, and run flwr supernode list tdaw-local --format json
flower_fleet_invalid Missing/stale map or evidence, identity mismatch, offline nodes, or image/topology drift Run the committed Python readiness validator; only rotate identity if the artifacts and persistent volume are intentionally being replaced
credential rotation requires ... archived and removed --force was used while old identity artifacts still exist Stop: use the normal restart path unless you truly intend to destroy and replace the deployment identity
Compose requests TDAW_GIT_COMMIT or TDAW_GIT_DIRTY Image/runtime provenance was not exported Export the exact 40-character commit and a truthful true/false dirty flag before Compose commands
Compose reports a missing bind-mount source MNIST, run, or registry directories were not prepared Run bootstrap/smoke and create the three shared directories shown above; Compose deliberately does not create them
Containers are healthy but the dashboard still says unavailable The API is stopped/stale or its Flower CLI/data-root configuration differs from Compose Restart mise run api, then inspect /api/v1/deployment-readiness; backend diagnostics are shown in the experiment detail/error UI
Port 3000, 8787, 5432, or 9093 is in use Another local process or stack owns the loopback port Stop the conflicting process or configure the supported override before startup

The Flower deployment guide contains the deeper trust-boundary rationale and exact profile-specific commands. The root README intentionally documents the supported default path so a first-time setup does not depend on discovering that file by accident.

Reproduce the frozen study

The matrix runner checkpoints every completed or failed cell. The analyzer rejects incomplete matrices and verifies the recorded configs, artifacts, event streams, partition manifests, and paired seeds before producing tables and plots.

mise run run-final-matrix
mise run analyze-final-matrix -- <matrix-directory>

The committed analysis is in docs/results/final-temporal. Earlier malicious-fraction and non-IID coverage remains available in the P1 experiment report, but it is not part of the final temporal ablation.

Development and validation

Use mise as the entry point for repository tasks:

mise run fmt
mise run lint
mise run test
mise run check

mise run check performs repository validation, environment diagnostics, formatting checks, Rust/Python/web linting, unit and integration tests, and production builds. PostgreSQL-backed lifecycle tests require the local Compose service.

Useful focused commands include:

mise run test-core
mise run test-python
mise run web-lint
mise run web-typecheck

The two confirmed cleanup tasks have deliberately different recovery paths:

Task Scope Next step
mise run purge-experiments Experiment metadata/artifacts only; preserves Flower identity, MNIST, dependencies, caches, and images Use Restart after purging experiments
mise run purge All generated and Git-ignored project state, including both Docker volumes, credentials, MNIST, dependencies, caches, and builds Repeat Bootstrap a fresh clone, then provision Flower again if needed

Both preserve tracked changes, untracked source files, global tool installations, and unrelated Docker state. Stop host API, dashboard, and experiment processes before either task.

Documentation

Scope and limitations

  • The default model is a small 784 → 128 → 10 MNIST MLP, not a large or production workload.
  • The main experiments use synchronous rounds and a trusted centralized research server.
  • Attack labels are simulator ground truth used for evaluation; they are never detector input.
  • A clean server calibration set is an explicit assumption when candidate validation is enabled.
  • Full model deltas are not stored in JSONL, PostgreSQL, or the browser, but visible research modes still expose them to trusted server code during aggregation.
  • SecAgg+ protects individual updates from server application code at the cost of per-client TDAW observability.

Use model update or model delta for the client payload. With local training across multiple mini-batches or epochs, it is not a literal per-batch gradient.

License

First-party source code and documentation are available under the MIT License. Third-party dependencies and vendored components retain their own terms; see LICENSE-NOTICE.md.

About

Reproducible TDAW research lab for history-aware Byzantine-robust federated MNIST training, with a Rust core, Flower modes, and a Next.js dashboard.

Topics

Resources

Security policy

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages