Appearance
Multi-node archive
Coming soon
This feature depends on the management service (platform/backend), which is not publicly available yet — it's coming soon. Standalone single-node recording, restreaming and export work today without it.
When a camera moved between recorder nodes during the export window you've requested, its archive segments are physically split across multiple servers. Multi-node archive is the control layer that collects fragments from the right nodes, stitches them into one MP4, and serves it to the user.
This feature requires both a shared database and a management service (platform/backend). Standalone recorders keep handling single-node exports on their own — nothing here changes that.
When the feature is active
Multi-node archive runs when both conditions hold:
- All recorders point at the same PostgreSQL instance (see Configuration — Database).
- The management service
platform/backendis deployed next to them, and the UI/SDK calls its address for export, not a recorder's.
In this mode the management service initiates the export and the recorders act as workers: one handles its own fragment, another does the final concat. The end user sees a normal job with a single job_id and the standard submit → poll → download cycle.
Flow
┌────────┐ POST /api/archive/{cam}/export
│ user │ ──────────────────────────────────► ┌────────────┐
└────────┘ │ platform/ │
│ backend │
└─────┬──────┘
│
┌──────────────────────────────────────────────────┤
│ 1. SELECT recordings WHERE camera AND window
│ 2. Group by host node into chronological "runs"
│ 3. INSERT archive_jobs + N×archive_job_runs
│ 4. asyncio background orchestrator task
│
▼ phase 1: sub-export on each owning recorder
┌─────────┐ ┌─────────┐ ┌─────────┐
│ node S1 │ POST /export?from&to (run 0) │ node S2 │ │ node S3 │
│ │ ─────────────────────────────►│ │ │ │
│ ffmpeg │ │ ffmpeg │ │ ffmpeg │
│ -c copy │ │ -c copy │ │ -c copy │
└────┬────┘ └────┬────┘ └────┬────┘
│ part_0.mp4 ready │ │
│ │ │
phase 2: backend picks an "assembler" — the node owning the biggest share
│ │ │
┌────▼─────────────────────────────────────────────────────────────┐
│ assembler (e.g. S2) │
│ POST /api/archive/{cam}/assemble │
│ body: [(addr_S1, job_0), (addr_S2, job_1), (addr_S3, job_2)] │
│ ──► downloads the 3 parts in parallel │
│ ──► ffmpeg concat parts.txt → final.mp4 │
│ ──► (timelapse only) ffmpeg setpts=PTS/N → final-tl.mp4 │
└─────────────────────────────────────────────────────────────────┬──┘
│
┌─────────────────────────────────────────────────────────────┘
│
▼ phase 3: user downloads
┌────────┐ GET /api/archive/jobs/{id}/download
│ user │ ◄────────────────────────────────── backend proxies stream
└────────┘ from the assemblerWhat a "run" is
The management service groups segments from recordings into contiguous per-host blocks — runs. Each run becomes one POST to one recorder.
Example. A camera lived on S1 for 12h, migrated to S2 for 2h, then returned to S1. The user asks for the whole window — three runs come out:
| # | Node | Window | Action |
|---|---|---|---|
| 0 | S1 | T..T+12h | POST /export on S1 |
| 1 | S2 | T+12h..T+14h | POST /export on S2 |
| 2 | S1 | T+14h..end | POST /export on S1 again as a separate job |
Even though runs 0 and 2 share a host, there's a foreign-host gap between them, so they go as two separate sub-export requests with two different windows. If we sent one wide request to S1 covering both, the resulting MP4 would have an internal jump where S2's slice should be.
Who the "assembler" is
Once all parts are ready, one recorder downloads them and runs the concat. The management service picks it by:
"the node owning the longest portion of the requested window"
Properties:
- No-migration case (one node holds everything): the assembler is that same node, downloads one file from
127.0.0.1, concat is a no-op. - With migrations, this choice minimises bytes-on-wire — the assembler only pulls foreign parts.
- The management service is never the assembler. It only orchestrates and proxies the final download.
Management service endpoints
These endpoints live on platform/backend. They look like recorder endpoints but actually orchestrate multiple nodes.
Submit an export
POST /api/archive/{camera_id}/export
Content-Type: application/json
{
"from_ts": 1780054469,
"to_ts": 1780058069
}Response: 202 Accepted
json
{
"job_id": "9e23...",
"camera_id": "camera_001",
"kind": "export",
"from_ts": 1780054469,
"to_ts": 1780058069,
"state": "pending",
"created_at": 1780100000,
"runs": [
{"ordinal": 0, "node_name": "S1", "from_ts": 1780054469, "to_ts": 1780056269, "recorder_state": "pending"},
{"ordinal": 1, "node_name": "S2", "from_ts": 1780056269, "to_ts": 1780058069, "recorder_state": "pending"}
]
}The runs array surfaces the plan up-front — you see which nodes will be involved and which window each gets.
Submit a time-lapse
POST /api/archive/{camera_id}/timelapse
Content-Type: application/json
{
"from_ts": 1780054469,
"to_ts": 1780100000,
"speed": 30
}speed is an integer in [2, 120]. Sub-exports on the recorders stay plain stream-copy exports; the setpts=PTS/speed pass is applied by the assembler on the concatenated file.
Status
GET /api/archive/jobs/{job_id}Possible state values:
| Value | Meaning |
|---|---|
pending | job just created, recorders not yet contacted |
exporting | runs are working on their owning recorders |
assembling | all runs done; assembler is stitching |
ready | final MP4 ready to download |
failed | job failed; check reason |
runs[].recorder_state mirrors the recorder's local job state (pending / running / done / failed). runs[].attempts shows how many times the management service had to resubmit that sub-export (see resilience below).
After failed, reason carries a human-readable explanation, e.g. run 2 on S3 unrecoverable after 3 attempts.
Download
GET /api/archive/jobs/{job_id}/downloadStreams the assembled MP4 through the management service. The service does not buffer to its own disk — bytes flow straight from the assembler recorder to the client. This keeps the backend lightweight even with a small VPS.
Possible responses:
200+video/mp4— the file stream.409 Conflict— job isn'treadyyet.410 Gone— job was ready but the assembler has since pruned its output (1 h TTL). Re-submit.502 Bad Gateway— assembler temporarily unreachable. Try again in a minute.
List jobs
GET /api/archive/jobs?camera_id=camera_001&limit=50&offset=0History per camera (or across all cameras). Returns the compact ArchiveJobListItemDto form (no runs array).
Cancel / cleanup
DELETE /api/archive/jobs/{job_id}Removes the database row and best-effort deletes intermediate files from every recorder involved (sub-exports + assembler's output). If a recorder is unreachable, its own 1 h TTL sweeper catches the leftover.
Resilience
Multi-node exports take minutes; any component can restart in the middle of one.
Case 1: a recorder dies mid sub-export
The plan and the window are persisted, so on the next poll the management service sees 404 from the recovered recorder, resubmits that sub-export with the same window, and stores the new recorder job_id. attempts increments; the limit is 3 per run before the job fails.
The recording files on disk weren't lost (they live on persistent storage, not in the recorder's in-memory job map), so the resubmit produces the same output.
Case 2: the assembler dies mid-concat
The management service sees 404 from the assembler, picks another recorder (excluding the dead one), and sends the same /assemble task with the same parts list. Sub-export outputs still exist on their recorders (within the 1 h TTL window), so the new assembler picks them up. Limit: 2 assembler attempts per job.
Case 3: the management service (backend) restarts
This is the trickiest one and we handle it explicitly.
What's preserved: archive_jobs and archive_job_runs in PostgreSQL. On restart the backend knows the plan of every job and which recorder_job_id each recorder was given.
At startup the backend scans all jobs in non-terminal states and spawns a fresh orchestrator task for each. The orchestrator is idempotent:
- If a run already has a
recorder_job_id, we don't resubmit — we just poll. - If a recorder returns
404(it also restarted), the normal run retry path kicks in. - If the job was in
assemblingand the assembler is still alive, the orchestrator joins its poll loop directly. - If the assembler is gone, a fresh one is picked (with the previous one excluded).
End result: the user doesn't lose the job over a backend restart — a brief pause, then polling resumes. In the worst case some in-flight work fails cleanly and the user re-submits.
Case 4: a needed node isn't in the servers registry
If a node_name from archive_job_runs doesn't have a matching servers.name row (decommissioned node, typo), the job fails immediately with reason node 'X' not in servers registry. Fix the registry or prune stale recordings.
Performance & parallelism
- The management service does no heavy work — no ffmpeg, no buffering. Just PG queries, recorder HTTP calls, and a streaming proxy. A small VPS is enough.
- Heavy work (decode/encode, ffmpeg concat) is distributed across recorders — more nodes mean more concurrent jobs in flight.
- Per-node limits remain identical to single-node export: 4 parallel exports and 2 parallel time-lapse jobs per recorder; extras queue up in
recorder_state: pending. - The assembler downloads parts in batches of 3 — CPU/network protection against chatty-migration scenarios with many runs.
Known limitations in this version
- No authentication. The management service calls recorders and recorders call each other without an
Authorizationheader. This is intended for closed internal networks. Auth is the next iteration. - Every recorder must be in
serverswith the correctname(the same value the recorder writes asnode_name). Without this row the management service can't dispatch sub-exports there. - Small windows that fit on a single node still go through the assembling phase. That's an extra concat pass over a single input file — milliseconds on a 4 h export, but it shows in the log.
- Resume after backend restart picks up polling but doesn't recover in-flight ffmpeg progress on recorders. If a sub-export was almost finished we either keep it (if the recorder survived) or resubmit from scratch (if it didn't).
Troubleshooting
| Symptom | Look at |
|---|---|
state=failed, reason mentions not in servers registry | A servers.name row is missing for one owner node. Compare archive_job_runs.node_name with servers.name. |
state=failed, reason mentions unrecoverable after N attempts | Either the recorder is down or its segment filter doesn't return the data. Check GET /api/archive/{cam}/segments?from=&to= on that recorder. |
state=failed, reason mentions assembler ... after N tries | All assembler attempts failed. Maybe only one node in the window and it's bad. Check the assembler's logs — ffmpeg stderr will be there. |
download → 410 | More than an hour has passed; the recorder pruned the MP4. Re-submit. |
download → 502 | Assembler unreachable right now. Transient; retry. |
runs[].attempts > 0 | At least one recorder restart happened during this job. Informational, not an error. |
Compatibility with single-node mode
Direct recorder calls POST /api/archive/{cam}/export (see Management API) keep behaving exactly as before — that's the single-node export using the local filesystem. Multi-node behaviour shows up only when the export is initiated through the management service.
So if you migrate from a single-node to a multi-node deployment, the UI doesn't have to change: both endpoints return the same job_id/state/download_url shape. Just point the UI at the management service's address instead of a specific recorder.