Recordings & Replay
How Synento's layout-agnostic recording model works and how to replay sessions with configurable layouts.
How Synento's layout-agnostic recording model works and how to replay sessions with configurable layouts.
The Stream-Native Recording Model
In Synento, recording is not a separate step. Every session automatically creates a recording the moment it starts:
Capture (per participant) → S2 append log (separated tracks + event timeline)
↓
┌─────────────────────┴─────────────────────┐
↓ ↓
Composable replay (client) Export (server, MP4)
layout descriptor + director track same layout vocabularyThere is no encoding delay. As soon as the last participant leaves a session, the recording is instantly replayable at full fidelity. Presentation is never baked in — the stored data is a timeline of media tracks and typed events; the player (or exporter) composes the final view at playback time.
What Gets Recorded
Each session under a/{projectId}/rooms/{sessionId}/ persists several stream types:
1. Separated media tracks (media/{userId}/{kind})
Each participant writes up to three independent track streams:
| Stream | Kind | Contents |
|---|---|---|
media/{userId}/cam | Camera video | WebCodecs EVC1 (H.264/VP9/VP8) or MJPEG image bytes |
media/{userId}/mic | Microphone audio | Raw audio samples (s16le, 16 kHz mono) |
media/{userId}/screen | Screen share | Same video formats as cam |
Track separation lets playback select only what a layout needs (e.g. active-speaker uses audio independently of video; screen-share focus reads the screen track alone).
Video payloads are self-describing:
- WebCodecs chunk (preferred):
EVC1magic + 18-byte header (codec id, keyframe flag, dimensions, timestamp µs) + elementary stream (annexb H.264, or VP9/VP8). - MJPEG image (fallback): raw JPEG/WebP bytes.
The server relay stores and forwards payloads opaquely — it never re-encodes.
2. Event timeline (events stream)
A versioned, append-only JSON timeline. Every event shares one envelope:
{ "v": 1, "type": "participant.joined", "ts": 1699876000123, "actor": "alice", "data": {} }Core event types:
| Type | Purpose |
|---|---|
participant.joined / participant.left | Lifecycle |
track.added / track.removed | Track lifecycle (kind: cam, mic, screen) |
screenshare.started / screenshare.stopped | Screen-share lifecycle |
speaking.started / speaking.stopped | Voice activity (from client-side VAD at capture) |
director.cut | Computed layout decision (focus participant, reason) |
caption.cue | Timed caption overlay |
branding.applied | Logo/watermark/accent colour |
highlight.marked | Highlight clip span |
participant.redacted | GDPR/removal annotation |
The meta stream still receives legacy join/leave/control records for compatibility, but events is the canonical timeline for reconstructing session state.
3. Director track (director stream)
When a session ends, the server computes active-speaker cuts from the speaking timeline and persists them as director.cut events. This makes active-speaker playback deterministic and cheap — clients read pre-computed focus switches instead of re-deriving them from raw audio on every replay.
Director cuts honour a minimum hold time (default 1.5 s) to avoid rapid switching.
4. Chat (chat stream)
Text messages between participants as JSON (base64-encoded):
{"text": "Hello everyone!", "userId": "bob", "timestamp_ms": 1234567}5. Captions (captions/{lang} stream, optional)
Per-language caption cues, using the same caption.cue event envelope. The session manifest advertises a captions stream when present.
Session Manifest
GET /v1/sessions/:sessionId/manifest returns the reconstruction contract — enough to bootstrap playback without reading every stream:
{
"data": {
"version": 1,
"session_id": "sess-abc123",
"epoch": 1699876000000,
"start_ts": 1699876000000,
"end_ts": 1699876300000,
"participants": [
{
"user_id": "alice",
"tracks": [
{ "kind": "cam", "stream_id": "a/.../media/alice/cam", "start_ts": 1000, "end_ts": 5000 },
{ "kind": "mic", "stream_id": "a/.../media/alice/mic", "start_ts": 1100, "end_ts": 4980 }
]
}
],
"streams": {
"events": "a/.../events",
"chat": "a/.../chat",
"director": "a/.../director",
"captions": "a/.../captions/en"
},
"capabilities": [
"separated-tracks",
"event-timeline",
"director-track",
"configurable-layouts",
"captions"
]
},
"error": null
}Use the manifest to fetch only the tracks a layout needs (bandwidth win) and to discover optional features (director track, captions).
Public sessions: GET /v1/public/sessions/:sessionId/manifest.
Session Lifecycle & Recording States
| State | Description | Timestamp Field |
|---|---|---|
| Created | Session created, no participants yet | created_at set |
| Started | First participant joined | started_at set |
| Streaming | Active session with participants connected | — |
| Ended / Recording Ready | Last participant left, recording complete; director track computed | ended_at set |
A session's "recording" is the persisted streams between start and end. To check if a recording is ready: look for ended_at !== null.
Accessing Recordings (Replay)
Step 1: Get the manifest (recommended)
curl https://api.synento.com/v1/sessions/{sessionId}/manifest \
-H "Authorization: Bearer sk_live_xxxxx"Returns participants, separated track index, time bounds, and capabilities.
Alternatively, use the lighter playback timeline for stream IDs and bounds:
curl https://api.synento.com/v1/sessions/{sessionId}/playback/timeline \
-H "Authorization: Bearer sk_live_xxxxx"Step 2: Read playback windows (iterative)
The playback API reads all relevant streams (events, director, chat, and every participant track) in time-based windows:
curl "https://api.synento.com/v1/sessions/{sessionId}/playback?from_ts=0&window_size=30000" \
-H "Authorization: Bearer sk_live_xxxxx"Continue while next_base_seq is non-null. When it is null, playback is complete.
Each event includes stream_id (full S2 path), timestamp, and body_base64. Media track events carry raw payload bytes (no outer wire header).
Step 3: Compose the view (client side)
Rather than rendering one participant at a time, the composition layer decodes the tracks a layout needs and draws them into a single canvas:
- Reduce session state — process
eventsanddirectorevents up to the playhead (participants, active speaker, screen sharers). - Resolve layout — map a layout descriptor + session state to normalized regions (0..1 coordinates).
- Decode tracks — WebCodecs
VideoDecoderforEVC1chunks,createImageBitmapfor MJPEG; one offscreen surface per visible track. - Draw — composite regions onto the output canvas; optional caption and branding overlays.
Available layout types:
| Layout | Behaviour |
|---|---|
active-speaker | Focus the current speaker (from director/speaking events); thumbnails for others |
presenter | Focus a designated participant |
grid | Equal tiles for all active cameras |
screen-share | Full-frame screen share with camera thumbnails; falls back to active-speaker |
selected | Grid of chosen participants only |
custom | Explicit normalized regions |
The dashboard replay player (web/app) exposes a layout selector (Speaker / Grid / Presenter / Screen). The player SDK provides the same engine for embeds.
Replay Controls
- Seek — jump to any timestamp (
from_tsparameter) - Speed control — play at 0.5x, 1x (default), or 2x
- Layout switching — change composition instantly without re-fetching (client-side)
- Pause / resume — stop consuming events without losing position
Player SDK (@synento/player)
Two limits to know before you choose this path.
In-browser replay is video-only. The player composes camera and screen tracks; it does not decode the microphone tracks. Audio exists in the MP4 export, which is muxed server-side. If your users need to hear a recording, export it and serve the file — see Exporting to MP4 — rather than reaching for the player.
PlaybackEngineauthenticates with an API key, so the snippet below is for trusted, server-side or first-party contexts only. Never ship ask_live_key to a browser. For a customer-facing embed today, proxy the playback endpoints through your own backend (which holds the key) exactly as the Synento dashboard does, or make the session public and read the unauthenticated/v1/public/sessions/:id/...endpoints.
The player SDK wraps windowed playback and provides a full composition stack:
import {
PlaybackEngine,
Compositor,
resolveLayout,
reduceSessionState,
parseSessionEvents,
availableTracks,
type LayoutDescriptor,
} from "@synento/player";
const engine = new PlaybackEngine(sessionId, apiBase, apiKey);
await engine.fetchTimeline();
const canvas = document.querySelector("canvas")!;
const compositor = new Compositor({ canvas, width: 1280, height: 720 });
const descriptor: LayoutDescriptor = {
type: "active-speaker",
orientation: "landscape",
captions: true,
};
engine.isPlaying = true;
await engine.play(() => {
// Pass accumulated window events + current playhead to the compositor.
compositor.render(descriptor, windowEvents, engine.currentTimeMs);
});
compositor.dispose();Lower-level helpers (resolveLayout, reduceSessionState, availableTracks, resolveActiveCaption, resolveBranding) are pure and reusable — the server export layout engine shares the same vocabulary.
See @synento/player for the full API.
Exporting to MP4
Export is async and parameterized by layout. The web dashboard can enqueue and download exports while logged in (session cookie via the API proxy). For scripts and integrations, pass a project API key on each request:
# Enqueue export (returns job id)
curl "https://api.synento.com/v1/sessions/{sessionId}/export?layout=presenter&orientation=vertical&focus=alice" \
-H "Authorization: Bearer sk_live_xxxxx"Response (202):
{ "data": { "job_id": "..." }, "error": null }Query parameters:
| Parameter | Values | Default |
|---|---|---|
layout | grid, presenter | grid |
orientation | landscape (1280×720), vertical (720×1280) | landscape |
focus | Participant user id (for presenter layout) | first participant |
Poll status:
curl https://api.synento.com/v1/exports/{jobId} \
-H "Authorization: Bearer sk_live_xxxxx"When status is completed, download via download_url:
curl https://api.synento.com/v1/exports/{jobId}/download \
-H "Authorization: Bearer sk_live_xxxxx" \
--output session.mp4Export behaviour:
- Reads separated
camandmictracks per participant - Codec-aware: H.264 sessions mux annexb directly; MJPEG sessions use
image2pipewith frame-gap repetition - Composes video via FFmpeg using the same layout vocabulary as the client compositor
- Export artifacts are stored by Synento and served from the download endpoint
Participant redaction
Remove a participant's media while keeping the rest of the session:
curl -X DELETE https://api.synento.com/v1/sessions/{sessionId}/participants/{userId} \
-H "Authorization: Bearer sk_live_xxxxx"This deletes all of the participant's track streams (cam, mic, screen) and appends a participant.redacted event to the timeline. Playback and export skip redacted tracks and can render a placeholder.
Architecture diagram
flowchart TB
subgraph rec [Recording layer — layout agnostic]
cam[media/userId/cam]
mic[media/userId/mic]
scr[media/userId/screen]
ev[events timeline]
dir[director track]
end
rec --> comp[Composition layer]
comp --> lay[Layout descriptor]
comp --> cli[Client compositor — @synento/player]
comp --> srv[Server exporter — FFmpeg]
lay --> cli
lay --> srv
cli --> view[Interactive playback]
srv --> art[MP4 / vertical / clips]