Minimal streaming server — takes RTMP/SRT inputs and restreams them to multiple RTMP/SRT outputs. Built on official SRS for ingest, srt-bonding-relay for bonded SRT ingress, and FFmpeg for outputs. Node.js + TypeScript backend.
Designed to handle tens of simultaneous pipelines (inputs) and hundreds of output forwards running continuously across long events. See Capacity & Limits for the tested envelope.
OBS / ffmpeg ──RTMP────────► SRS (21935) ──FFmpeg──► YouTube / Facebook / ...
──SRT─────────► SRS (10080) ──FFmpeg──► rtmp:// or srt://
──SRT bonding─► srt-bonding-relay (10081) ──► SRS
| Component | Description |
|---|---|
| SRS | Ingest broker — accepts RTMP and SRT streams |
| srt-bonding-relay | Standalone bonded SRT relay (GitHub release in prod, sibling repo in dev) |
| Node.js app | REST API + dashboard on port 8080 |
| FFmpeg | One process per output, spawned and managed by the app |
| SQLite | Persistent state for pipelines, outputs, stream keys, settings |
The server is built and operated for the envelope below. It has not been tested or designed for anything beyond it — treat these as the supported ceiling, not a target to exceed :)
| Limit | Supported ceiling |
|---|---|
| Inputs (pipelines) | up to 50 |
| Outputs (forwards) | up to 500 total |
| Outputs using custom (transcoding) encoding | only a few at a time — see below |
| Parallel dashboard clients | up to ~10 |
Past these figures, expect host CPU/RAM/network and the number of concurrent FFmpeg processes (one per output) to become the limiting factors well before the dashboard or API does.
Keep almost all outputs in copy mode. A copy output is a passthrough — it
remuxes the input and forwards it with negligible CPU cost, so hundreds can run on
modest hardware. A custom encoding (720p, 1080p, vertical_rotate, …) makes
FFmpeg transcode the video, which is CPU-intensive: each such output consumes
roughly a full core's worth of work. Only a handful of outputs should use custom
encoding at any one time; everything else should be copy. Putting many outputs
into custom-encoding mode will saturate the CPU long before the 500-output ceiling
and starve the copy outputs and the dashboard alike.
Parallel dashboard clients. Health is computed once every 5s and shared by all clients, so extra browser tabs do not multiply SRS/SRT-FFprobe work. API responses are gzip-compressed, and config changes from another session are picked up automatically via the health poll's config revision. Around 10 simultaneous dashboard clients is fine; higher counts are not tuned or tested.
This app now runs natively on Linux. The production setup uses three systemd services:
| Service | Purpose |
|---|---|
srs.service |
Native SRS binary, started as /usr/local/bin/srs -c /etc/restream-srs/srs.conf |
srt-bonding-relay.service |
Shared SRT bonding relay, started on UDP port 10081 |
restream-srs.service |
Node.js dashboard/API, started from /opt/restream-srs/dist/index.js |
The installer downloads the official SRS release binary and a pinned srt-bonding-relay binary from the standalone live-miracles/srt-bonding-relay GitHub releases.
Startup is ordered so the Node.js control plane is reachable before SRS accepts
publishes. srs.service waits for the unauthenticated readiness endpoint
http://127.0.0.1:8080/api/ready before starting; this prevents SRS publish
hooks from racing the dashboard/API process during boot.
Production install:
sudo apt-get update -q && sudo apt-get install -y -q git # in case git is not installed
sudo git clone https://github.com/live-miracles/restream-srs /opt/restream-srs
sudo bash /opt/restream-srs/scripts/server-install.shSet server timezone (optional, for readable log timestamps):
SRS and journald log timestamps in the system's local timezone, which defaults to UTC on most fresh servers. Set it to your local timezone so log times make sense at a glance, e.g. for IST (Asia/Kolkata):
sudo timedatectl set-timezone Asia/KolkataList available timezone names with timedatectl list-timezones. No reboot is needed, but already-running services may have the old timezone cached in memory, so restart them to apply it to new log lines immediately:
sudo systemctl restart srs.service restream-srs.service srt-bonding-relay.serviceUpdate an installed server:
sudo bash /opt/restream-srs/scripts/server-install.shStop services:
sudo bash /opt/restream-srs/scripts/server-down.shOpen the dashboard: http://SERVER_IP:8080 — default password is admin.
Default ports from srs.conf and srt-bonding-relay.json:
| Port | Protocol | Purpose |
|---|---|---|
| 21935 | TCP | RTMP input (non-default port to avoid 1935 scanner noise) |
| 10080 | UDP | SRT input (passphrase required) |
| 10081 | UDP | SRT bonding input (passphrase required) |
| 8080 | TCP | Dashboard + API |
Do not expose 1985 (SRS HTTP API), 8080 (if the dashboard is served through a tunnel), or 8081 (relay status) — the app talks to those over loopback.
The dashboard is meant to be reached through a Cloudflare Tunnel plus
Cloudflare Access (see Security), so port 8080 can stay closed in
the firewall. The installer does not set this up; cloudflared runs as its own
systemd service, outside the three services above. Ingest (RTMP/SRT) and
outputs never go through the tunnel, so restarting cloudflared only
interrupts dashboard/API access.
- Create the tunnel in Cloudflare Zero Trust → Networks → Tunnels →
Create a tunnel (type Cloudflared). Copy the tunnel token from the
install command it shows (the long string after
--token). - Add a public hostname to the tunnel, e.g.
restream.example.com, with serviceHTTP→localhost:8080. - Protect it with Cloudflare Access: Zero Trust → Access → Applications → add a self-hosted application for that hostname with an allow policy for your team. Without this the hostname is open to the internet and only the dashboard password stands in the way.
- Install
cloudflaredfrom the Cloudflare package repository (or the.debfrom its GitHub releases). - Store the token in a root-only environment file rather than on the
command line. Do not use
cloudflared service install <token>: it puts the token inExecStart, where every local user can read it viaps//proc/<pid>/cmdlineand the world-readable unit file.sudo install -d -m 0750 /etc/cloudflared sudo sh -c 'umask 077; printf "TUNNEL_TOKEN=%s\n" "<paste-token-here>" > /etc/cloudflared/tunnel.env'cloudflared tunnel runreads theTUNNEL_TOKENenvironment variable itself, so no--tokenflag is needed. - Create the unit
/etc/systemd/system/cloudflared.service:[Unit] Description=Cloudflare Tunnel client After=network-online.target Wants=network-online.target [Service] TimeoutStartSec=15 Type=notify EnvironmentFile=/etc/cloudflared/tunnel.env ExecStart=/usr/bin/cloudflared --no-autoupdate tunnel run Restart=on-failure RestartSec=5s [Install] WantedBy=multi-user.target
- Start it and check:
sudo systemctl daemon-reload sudo systemctl enable --now cloudflared journalctl -u cloudflared -n 20 --no-pager # look for "Registered tunnel connection" ps -eo args | grep '[c]loudflared' # the token must not appear here
To rotate the token, generate a new one in the Zero Trust dashboard, update
/etc/cloudflared/tunnel.env, and sudo systemctl restart cloudflared. Treat
the token as compromised (and rotate it) if it was ever printed in a log or
terminal, or was previously passed on the command line.
The repository copies both runtime config files during install:
/etc/restream-srs/srs.conf/etc/restream-srs/srt-bonding-relay.json
SRS only reads its config at startup, and the SRT bonding relay only reads its JSON config file at startup.
The repo configs ship a public default SRT passphrase; on first install the installer generates a random per-server secret and writes it to both files. On reinstall it leaves each file's passphrase exactly as deployed — it doesn't compare, regenerate, or sync the two. SRS rejects SRT connections without the passphrase at the handshake, for publish and play alike; the dashboard's SRT publish URLs include it.
To change or disable it, edit both files by hand and restart SRS and the
relay: passphrase ; in srs.conf, "passphrase": "" in
srt-bonding-relay.json. An empty value in both disables the passphrase
check entirely — SRS and the relay treat empty the same as unset.
Ingest is further locked down by SRS HTTP hooks handled by the app:
on_publish rejects unknown stream keys, and on_play rejects any play not
from loopback (only the app's own FFmpeg ever pulls streams), so the public
ports are ingest-only. See
Known issues for why rejected attempts are logged
but not converted into firewall bans.
The dashboard is protected by a password (default is admin). Change it in Settings → Change Password after logging in.
To reset a forgotten password:
sudo bash /opt/restream-srs/scripts/server-reset-password.shThis resets the password to admin and restarts the service.
Dashboard/API auth is intentionally simple. It's a single shared password
(see Authentication) behind a session cookie — no
per-user accounts, no MFA, no CSRF token beyond SameSite=Strict, and the
app itself never terminates TLS (it's plain app.listen on 8080; the
cookie isn't marked Secure). That's a deliberate tradeoff, not an
oversight: this app is designed to be reached through a zero-trust tunnel
(e.g. Cloudflare Tunnel + Cloudflare Access) rather than exposed to the raw
internet. The tunnel is expected to provide TLS, identity-based access
control, and the real perimeter; the built-in password is just a lightweight
second factor for whoever's already authenticated through that layer, plus a
login-rate-limit (see src/api/auth.ts) against brute force if it's ever
reached directly. If you expose port 8080 straight to the internet instead
of through a tunnel, put a real reverse proxy (TLS + rate limiting) in front
of it rather than relying on the dashboard password alone.
RTMP/SRT ingest ports (21935, 10080, 10081) are different — they have to stay open to the raw internet so OBS/hardware encoders can reach them, so they're hardened at the SRS/relay layer instead:
- SRT (
10080,10081) requires a passphrase, checked at the handshake before a connection is accepted at all — see SRS and SRT relay config. - Every publish attempt (RTMP or SRT) is checked against a per-pipeline
stream key via SRS's
on_publishhook; unknown keys are rejected before any restream starts. on_playrejects any playback request not from loopback, so the public ports are ingest-only — nothing can pull a stream back out through SRS directly.- RTMP listens on a non-default port (
21935, not1935) to cut down on generic scanner noise. - Rejected publish/play attempts are logged, but this app does not install fail2ban; see Known issues.
Never expose 1985 (SRS HTTP API) or 8081 (relay status) — the app only
talks to those over loopback.
For how to publish to a pipeline (ffmpeg test commands for RTMP/SRT) and how to configure translation outputs — which mix a separate translator pipeline's audio into an output — see the in-app User Manual (the open-book icon in the dashboard navbar). Translation output mixing, with per-track selection, speech-detection ducking, and configurable delay/restore timing, is available as a built-in audio encoding option on any output.
All routes below sit behind the session-cookie auth middleware except
/api/ready, /api/auth/login, and the SRS publish hook. An output pushes to a
single destination: url and audioEncoding, plus videoEncoding. The input is
pulled back over whatever protocol it was published with (SRT input → SRT pull,
RTMP input → RTMP pull), so there is no pull-method setting.
| Method | Path | Description |
|---|---|---|
| GET | /api/config |
Pipelines, outputs, encodings, stream keys, server name |
| GET | /api/health |
Live input/output status snapshot (refreshed every 5s) |
| GET | /api/version |
App version / build info |
| GET | /api/metrics/system |
Host CPU, RAM, disk and network stats |
| GET | /api/srs-logs |
Recent SRS up/down events and log tails for SRS, the dashboard, and the relay |
| GET | /api/ready |
Unauthenticated readiness check used by systemd before starting SRS |
| POST | /api/pipelines |
Create pipeline (auto-names and assigns stream key) |
| GET | /api/pipelines/:id |
Pipeline details including relay health and bonded-input activity |
| POST | /api/pipelines/:id |
Rename pipeline { name }, optionally reassign key { name, streamKeyId } |
| DELETE | /api/pipelines/:id |
Delete pipeline (stream key is freed, not deleted) |
| GET | /api/pipelines/:id/logs |
Pipeline online/offline event log |
| POST | /api/pipelines/:id/preview/start |
Start an HLS preview { audioTrack? } |
| POST | /api/pipelines/:id/preview/keepalive |
Refresh the preview TTL — previews with no keepalive for 90s are stopped automatically (covers closed browser tabs) |
| POST | /api/pipelines/:id/preview/stop |
Stop the HLS preview |
| POST | /api/pipelines/:id/outputs |
Create output { name, videoEncoding, url, audioEncoding } |
| POST | /api/pipelines/:id/outputs/bulk |
Bulk create outputs { outputs: [{ name, videoEncoding, url, audioEncoding }] } — validates all before creating any |
| DELETE | /api/pipelines/:id/outputs |
Clear all outputs for the pipeline — returns 409 if any output is still running |
| POST | /api/pipelines/:id/outputs/:outId |
Update output (same body as create) |
| DELETE | /api/pipelines/:id/outputs/:outId |
Delete output (stops it first if running) |
| POST | /api/pipelines/:id/outputs/start-all |
Start all outputs (staggered at 200 ms intervals, returns immediately) |
| POST | /api/pipelines/:id/outputs/stop-all |
Stop all outputs |
| POST | /api/pipelines/:id/outputs/:outId/start |
Start output |
| POST | /api/pipelines/:id/outputs/:outId/stop |
Stop output |
| POST | /api/settings |
Update settings { name, publicHost } |
| POST | /api/settings/regenerate-stream-keys |
Regenerate all stream keys |
| POST | /api/auth/login |
Login { password } — sets session cookie |
| POST | /api/auth/logout |
Logout — clears session cookie |
| POST | /api/auth/change-password |
Change password { currentPassword, newPassword } |
| POST | /api/srs/on_publish |
Unauthenticated SRS publish hook (called by SRS, not the dashboard) |
Prerequisites: Node.js 22+ plus curl, unzip, tar, and xz-utils.
1. Install dependencies and local media binaries:
npm install
npm run dev-install # downloads SRS and pinned FFmpeg into ./objsRerunning npm run dev-install also refreshes the sibling ../srt-bonding-relay
repo when it has no local changes and rebuilds ./objs/srt-bonding-relay.
To use a local SRS binary:
SRS_LOCAL_BIN=./build/srs npm run dev-installThe standalone relay now lives in the sibling ../srt-bonding-relay repo during development and in live-miracles/srt-bonding-relay GitHub Releases for production installs.
2. Start SRS (terminal 1):
npm run srs # runs ./objs/srs -c srs.conf in the foreground3. Start the SRT bonding relay (terminal 2):
npm run relay # runs the ./objs/srt-bonding-relay binary built by dev-install
npm run relay:update # pulls the latest ../srt-bonding-relay source and rebuilds ./objs/srt-bonding-relayThe relay also exposes a local HTTP status endpoint on the default
status_port (127.0.0.1:8081) in development. The dashboard backend polls
that endpoint to show the top-level relay health and the per-pipeline
bonded-input status.
4. Start the app (terminal 3):
npm run dev # tsx watch + tsc watch + tailwind watchThe app reads runtime settings from restream.json in the app root.
restream.json:
| Field | Default | Description |
|---|---|---|
port |
8080 |
App HTTP port |
database_path |
./db.sqlite |
SQLite database path |
srs_config_path |
./srs.conf |
SRS config path |
ffmpeg_path |
ffmpeg |
FFmpeg binary for outputs and previews |
ffprobe_path |
ffprobe |
FFprobe binary for input media probing and validation |
output_watchdog.warmup_ms |
90000 |
Warmup before stall checks, shared by the output progress watchdog and the translation-mixer watchdogs below |
output_watchdog.stall_ms |
45000 |
Output progress stall window before restarting FFmpeg |
output_watchdog.translator_meter_stale_ms |
10000 |
How long a translation output's translator audio meter can go quiet (including never producing a sample) before the mixer is restarted |
output_watchdog.interval_ms |
5000 |
Output watchdog polling interval |
output_watchdog.socket_warmup_ms |
15000 |
Socket watchdog warmup before socket-state checks |
output_watchdog.socket_grace_ms |
30000 |
Socket warning grace window before restarting FFmpeg |
output_watchdog.memory_limit_mb |
200 |
FFmpeg RSS limit (MB) before an output is killed and restarted; applies to copy and any videoEncoding not listed below |
output_watchdog.memory_limit_mb_by_encoding |
{vertical_rotate: 450, "720p": 650, "1080p": 950} |
Per-videoEncoding RSS limit override (MB); libx264 transcode profiles run at a much higher legitimate baseline than copy |
Relative file paths are resolved from the app root. Command names like ffmpeg
and ffprobe are left as command names.
SRS values are inferred from srs.conf:
- RTMP pull/publish port from top-level
listen - SRT pull/publish port from
srt_server { listen ... } - SRS HTTP API port from
http_api { listen ... } - SRS log tail path from
srs_log_file, if present — otherwise the dashboard falls back tojournalctl -u srs.service(production'ssrs.conflogs to console/journald instead of a file; only local dev's still setssrs_log_file)
The dashboard's own log and the relay's log are always read via
journalctl -u restream-srs.service / -u srt-bonding-relay.service — both only
ever log to the journal. All three require restream-srs.service's
SupplementaryGroups=systemd-journal (set by server-install.sh); without it,
the Server Logs view shows no output for whichever service's journal it can't read.
The control plane also writes structured diagnostics to
/var/lib/restream-srs/diagnostics/ in production. The rotated JSONL files keep
output starts/exits/restarts, SRS and relay transitions, input and bonded-leg
transitions, and FFmpeg timestamp/media-clock warnings for seven days (the last
ffmpeg stderr tail at exit is included on the exit event; raw stderr isn't
streamed to diagnostics to keep incident volume from crowding out the window).
Once per minute it also stores compact health snapshots containing SRT counters,
leg rates/health, relay forwarding state, and FFmpeg output progress so an
incident can be investigated after the live one-hour chart window has passed.
The snapshots also record resource usage: per-output FFmpeg RSS/CPU and memory
limit (translation mixers included), CPU and RSS for the Node control plane
(plus its V8 heap), SRS and the SRT relay (the same samples the dashboard
shows, up to ~10 s old), and host total/available memory and swap. Separately,
host-memory-low (available memory under 15% of total) and
host-memory-recovered (back above 20%) events are logged on transition, the
former with the five largest processes by RSS.
Files are
rotated daily or at 100 MB, whichever comes first, with a 5 GB total diagnostics
cap that removes the oldest rotated files first. The installer configures
persistent journald with a seven-day retention limit, 1 GB system journal cap,
256 MB runtime cap, and one-day journal file rotation.
Relay ports and status polling are read from srt-bonding-relay.json, located beside srs_config_path.
Installer/development overrides:
| Variable | Used by | Description |
|---|---|---|
REPO_URL |
server-install.sh |
Repository URL cloned into /opt/restream-srs on first install |
SRT_RELEASE_TAG |
server-install.sh |
Relay GitHub release tag to install |
SRT_URL |
server-install.sh |
Custom relay release archive URL |
SRT_SHA256 |
server-install.sh |
Expected SHA256 for the relay archive; blank skips verification |
WIPE_DB |
server-install.sh |
Set y/n to force or skip wiping the existing SQLite DB |
SRS_LOCAL_BIN |
dev-server-install.sh |
Local executable SRS binary to copy into ./objs/srs |
On GCP Ubuntu hosts, google-osconfig-agent.service may periodically run
apt-get update / package inventory checks and briefly consume a large fraction
of one vCPU. On a 4-vCPU VM this can show up as a ~25% CPU spike. This is host
maintenance noise, not a restream pipeline failure.
Temporary live-event mitigation, cleared automatically on reboot:
sudo systemctl stop google-osconfig-agent.service
sudo systemctl mask --runtime google-osconfig-agent.serviceIf Ubuntu Pro APT/ESM jobs are also noisy:
sudo systemctl mask --runtime apt-news.service esm-cache.serviceThis temporarily affects GCP OS patch/inventory reporting or Ubuntu Pro update messaging; re-enable after the event or reboot.
SRS's native srt_to_rtmp feature (which remuxes an SRT publish into the RTMP
layer so it can be played back over RTMP/HLS) emits audio with bursty,
discontinuous timestamps — roughly 60 ms gaps between 21 ms packets. The
resulting RTMP/HLS plays with constantly breaking/crackling audio, and because
RTMP/FLV carries only one audio stream it also collapses a multi-track SRT source
to a single track. No combination of SRS settings (hls_dts_directly, etc.) made
the audio clean.
How it's avoided: srt_to_rtmp is turned off in srs.conf. SRT inputs
stay in the native SRT/MPEG-TS domain and are pulled back over SRT (raw MPEG-TS),
which preserves every audio track and keeps timestamps intact; the HLS preview is
generated by the app's own ffmpeg rather than SRS's native HLS. RTMP inputs are
unaffected — they never went through srt_to_rtmp — and are pulled over RTMP.
The SRS log fills with srt serve error code=6001(SrtIo)(SRT read or write failed) : srt play recv thread : ... socket status=SRTS_BROKEN at a steady background rate.
This is expected noise, not an incident: it fires on every teardown of an SRT
play (pull) connection — including the ffprobe media check this app runs when a
publisher connects to an SRT input, and its retries every
FFPROBE_FAILED_REFRESH_MS while that probe keeps failing — not just on
genuine connection failures.
Confirmed in SRS 6.0 source (src/app/srs_app_srt_conn.cpp,
src/protocol/srs_protocol_srt.cpp): SrsMpegtsSrtConn::cycle() logs at
srs_error() for any non-success return from the play loop, with no branch for
"peer closed cleanly" vs. "connection broke" — and the SRT recv thread treats any
non-timeout read() result (including a normal peer disconnect) as a hard error
that propagates up to that log call. check_error() just stringifies whatever
srt_getsockstate() reports at that instant, and libsrt sets a torn-down read-side
socket to SRTS_BROKEN regardless of how gracefully the peer closed.
No client-side fix exists. SRT socket options like linger only affect
how the closing side waits to flush unsent data — since ffprobe is a
receive-only client, they don't change what SRS observes or logs. This was
verified by adding &linger=1 to the ffprobe pull URL and confirming SRS still
logged SRTS_BROKEN on every probe close. Fixing this would require patching
and rebuilding SRS itself.
Mitigation: filter SRTS_BROKEN / srt play recv thread lines out of log
monitoring/alerting rather than treating them as incidents.
When a translation output's translator input drops, the source's ducked
volume does not recover for tens of seconds — sometimes a minute or more.
Reconnecting the translator is fast by comparison (a few seconds). This
asymmetry is expected: connecting is an edge-triggered, unambiguous event (a
new publish is immediately visible), but disconnecting is inferred from
absence, which SRS can't trust right away — the app only reacts once its
health poll observes SRS's own stream list as gone, and SRS itself only
concludes an SRT peer is dead after its srt_server.peer_idle_timeout
elapses with nothing received, confirmed via SRTS_BROKEN in the SRS log
(see the entry above). In the rarer case where the translation mixer's own
FFmpeg process also stalls (rather than just losing the translator leg
cleanly), recovery instead waits on the coarser output-progress-stalled
watchdog (stallMs, default 45s) to restart it.
Why this is not tuned lower. peer_idle_timeout is a single setting on
SRS's shared SRT listener — it applies to every SRT input, not just
translators, including the main broadcast source. SRT exists specifically to
tolerate lossy/unstable networks; shortening this timeout to speed up
translator recovery would risk the main source being torn down on a brief,
real network blip during a live event. A translator SRT drop is rare, and
the output degrades gracefully in the meantime (it keeps publishing source
audio throughout), so that tradeoff is not worth the risk to the primary
broadcast path.
Mitigation: none needed — the output stays up on source audio for the entire gap. If translator disconnects become frequent enough to matter, investigate the translator's own network path rather than this timeout.
The app runs these recovery loops:
| Watchdog | Scope | Restart condition | Notes |
|---|---|---|---|
| Health poll / input recovery | SRS reachability, live pipeline inputs, desired running outputs | When SRS and the pipeline input become ready, outputs whose desired state is running are started or restarted with staggered timing |
Computed once every 5s and shared by dashboard clients. Inputs become live from SRS publisher presence; ffprobe fills in media/track metadata and retries every FFPROBE_FAILED_REFRESH_MS until it succeeds, then stops for that publisher. |
| Output progress watchdog | Every running FFmpeg output process | After warmup, if the input is ready but FFmpeg total_size / out_time_ms stop advancing for the configured stall window |
Protocol-agnostic backstop; covers SRT outputs and local RTMP relays |
| Remote RTMP socket watchdog | Running outputs with a remote RTMP/RTMPS destination | After socket warmup and grace, if the destination socket is missing or remains in a closing state such as CLOSE-WAIT |
Uses one ss -H -tanp snapshot per watchdog interval; a local RTMP/RTMPS destination is ignored because local input/output sockets are ambiguous |
| Output memory watchdog | Every running FFmpeg output process, including translation mixers (checked every 5 s; a mixer over its limit is restarted with a translation-mixer-restart event, reason memory limit exceeded) |
After warmup, if process RSS crosses memory_limit_mb (or its per-encoding override) |
Reads /proc/<pid>/status; unconditional — a leaking process can still show advancing total_size/healthy sockets, so this doesn't wait on the other two. Once RSS crosses 70% of the limit the output surfaces a yellow "High memory usage" warning in the dashboard, before the watchdog actually restarts it at 100%. The limit is doubled for outputs on a 4K (≥3840px on either dimension) input — the 2x multiplier is a placeholder, not a measured baseline; see #11 |
| Translator audio meter watchdog | Running translation-mixer outputs | After the same startup warmup grace as the output progress watchdog, if the translator audio meter has gone quiet for translator_meter_stale_ms (including a translator that never delivered a single sample) |
Watches independently of the SRS live-state poll, so a translator that connects but sends a dead/silent track is still caught. Valid translator silence is not treated as a disconnect — FFmpeg continues emitting silent meter samples in that case — so this only fires when the meter itself stops updating. Restarts the mixer so it attaches to the current translator session. |
All output watchdogs above use the same restart path: they write a detailed
last_error, kill the stuck FFmpeg process, and let the normal retry loop start a
fresh process while the output's desired state remains running. The socket
watchdog is advisory: if ss fails or times out, it does not restart anything,
and a socket warning does not prevent the other watchdogs from acting.
Input media validation is separate from the output watchdogs. The health service stays on the regular 5s poll, but ffprobe backs off once a publisher's probe succeeds:
| Input check | Scope | Cadence | Notes |
|---|---|---|---|
| Media validation | Connected RTMP/SRT inputs without a successful probe yet for this publisher | Immediate probe (staggered across pipelines), then retried every FFPROBE_FAILED_REFRESH_MS (30s) while it keeps failing |
Stops re-probing once a probe succeeds for the current publisher; a new probe cycle starts on the next publisher change. Input liveness and output recovery do not wait for a valid ffprobe result — ffprobe fills in media/track details when available. |
The relay exists to work around two constraints:
-
ffmpeg cannot accept bonded SRT group connections. ffmpeg's SRT handler does not set
SRTO_GROUPCONNECT=1on its listener socket, so bonded connection attempts from encoders like AJA Bridge Live are rejected at the handshake level. -
SRS has no native SRT bonding support. SRS accepts normal SRT publishers, but it does not accept bonded/redundant SRT groups directly.
How it is fixed: srt-bonding-relay starts as its own systemd service, listens on 10081 with SRTO_GROUPCONNECT=1, accepts each bonded source session, reads its incoming streamid, and opens a normal SRT publisher connection to SRS using the same streamid. It only connects to SRS after an encoder connects, so SRS's idle-publisher timeout is not triggered by an empty boot-time publisher.
The input is pulled back over its own protocol, so an output on an SRT input
always pulls over SRT. When the destination rejects such a stream (e.g. a wrong
YouTube stream key) or drops the RTMP connection mid-publish, ffmpeg can get
stuck instead of exiting. An output on an RTMP input (pulled over RTMP)
usually exits ffmpeg immediately with a clear error (Error opening output files: Input/output error).
The difference is timing: with RTMP pull, input stream info is available right
away, so ffmpeg opens the destination immediately and the rejection surfaces at
write_header time, exiting non-zero. With SRT pull, ffmpeg must first probe the
MPEG-TS input; by the time it connects, the destination accepts the handshake
then drops the connection mid-publish, and ffmpeg deadlocks — SRT's large input
buffers keep the input thread fed, so the broken-pipe error on the output write
may not propagate promptly. The process can stay alive while buffering input and
consuming RAM even though the external output is no longer uploading.
How it is handled: the output service tracks ffmpeg's -progress pipe:1
fields (total_size, out_time_ms, and bitrate) plus the stderr tail. It also
takes a single local TCP socket snapshot (ss -H -tanp) every watchdog interval
and maps RTMP/RTMPS destination sockets back to each ffmpeg pid. After a warmup
period, an output is marked yellow if its destination socket is missing or in a
closing state such as CLOSE-WAIT; if that persists past the socket grace
window, the watchdog records last_error, kills that ffmpeg process, and lets
the normal retry loop start a fresh output. The same retry path is used when the
input is live but ffmpeg output bytes/time stop advancing for a sustained window.
The dashboard shows the warning/restart reason and timestamp; the error details
include the pid, stall duration, last ffmpeg progress values, and stderr tail.
Limitation: the watchdog monitors aggregate ffmpeg output progress and
pid-level TCP socket state for the process. Local RTMP/RTMPS outputs are excluded
from the TCP socket check because ffmpeg's local input pull and local output push
both connect to local SRS and are ambiguous in ss; those relays are still
covered by the output-progress watchdog.
On a glitchy SRT input (real packet loss/corruption upstream, not a passphrase or
config issue — see the incident write-up below), FFmpeg's AAC decoder can
misparse a corrupted frame as a bogus multichannel layout (logged as [SWR] Full-on remixing from 22.2 has not yet been implemented!). Reinitializing the
resampler for that bogus layout leaks memory instead of failing cleanly, and RSS
can grow unbounded — observed reaching 1.5-1.6GB before the kernel OOM-killer
stepped in.
That kernel-level kill is the actual danger, not the leak itself: FFmpeg output
processes share a cgroup with the restream-srs control-plane service, so
systemd treats any OOM kill inside that cgroup as the whole service failing and
restarts it — which cascades (via Requires=) into restarting SRS and the SRT
bonding relay too, forcing every stream on the box to reconnect simultaneously.
That mass reconnect can reintroduce the same corrupt-resync conditions on another
output, repeating the cycle. Full timeline and log evidence from the 2026-07-10
incident: fail-reports/2026-07-10-pipeline1-output-oom-cascade.md.
How it is handled: two layers. buildFfmpegArgs passes -fflags +discardcorrupt -err_detect crccheck+bitstream so packets already flagged
broken by the demuxer are dropped before reaching the decoder, reducing (but not
eliminating) the chance of the misparse happening at all. As a hard backstop,
the output memory watchdog (see the watchdogs table above) kills and restarts an
output once its RSS crosses memory_limit_mb (default 500MB — roughly 6-8x a
healthy output's normal 65-90MB baseline, and far below the 1.5GB+ range where
the kernel OOM-killer struck) — well before the kernel ever needs to get
involved, so the cascading restart doesn't happen.
This repo intentionally does not install or manage fail2ban. Earlier versions used it as a narrow RTMP/app-hook protection, but it added root-owned config, sudo helper scripts, dashboard whitelist/status UI, and extra test surface while leaving the most important SRT gap uncovered.
Neither SRT listener logs a bad-passphrase rejection in a form fail2ban could act on:
- SRS's own SRT server (port
10080) rejects handshakes that fail itspassphrasecheck — either the peer declares no encryption at all, or declares it butSRTO_ENFORCEDENCRYPTIONfails the key exchange. Either way, the rejection happens inside the SRT handshake, beforesrt_accept()returns, so SRS's own log (ERROR:UNSECURE/ "Password required or unexpected") never records the peer's IP. srt-bonding-relay's listener (port10081) used to log a matchable[srt-relay] rejected connection (bad passphrase) from <HOST>line from its ownsrt_accept()KMSTATE check, and a dedicatedsrt-bonding-relayfail2ban jail watched it. As ofsrt-bonding-relayv2.1.0 that line is no longer logged, so a jail would have nothing left to match.
The app still logs rejected on_publish and on_play attempts, which would be
enough for a fail2ban jail covering repeated bad RTMP stream-key attempts from
the same IP. That protection is intentionally not included because it is only a
partial defense: it does not cover SRT bad passphrases, does little against
distributed attackers, and is not a replacement for a proper edge/firewall
deployment for high-risk events.
This is not believed to be a practical brute-force risk: the configured passphrase is a long random string, and each guess requires a live SRT handshake round trip, so online guessing is computationally infeasible regardless of banning. The real effect is possible log/CPU noise from a misconfigured device or scanner retrying indefinitely. For small events, keep the deployment simple and rely on strong stream keys, a strong SRT passphrase, and the non-default RTMP port. For larger or higher-risk events, keep the origin private and put a media-aware ingest edge or provider firewall/DDoS layer in front of it.