# Recordings & Replay

> How Synento's layout-agnostic recording model works and how to replay sessions with configurable layouts.

How Synento's layout-agnostic recording model works and how to replay sessions with configurable layouts.

## The Stream-Native Recording Model

In Synento, **recording is not a separate step**. Every session automatically creates a recording the moment it starts:

```
Capture (per participant) → S2 append log (separated tracks + event timeline)
                                    ↓
              ┌─────────────────────┴─────────────────────┐
              ↓                                           ↓
     Composable replay (client)                    Export (server, MP4)
     layout descriptor + director track            same layout vocabulary
```

There is no encoding delay. As soon as the last participant leaves a session, the recording is instantly replayable at full fidelity. **Presentation is never baked in** — the stored data is a timeline of media tracks and typed events; the player (or exporter) composes the final view at playback time.

## What Gets Recorded

Each session under `a/{projectId}/rooms/{sessionId}/` persists several stream types:

### 1. Separated media tracks (`media/{userId}/{kind}`)

Each participant writes up to three independent track streams:

| Stream | Kind | Contents |
|--------|------|----------|
| `media/{userId}/cam` | Camera video | WebCodecs `EVC1` (H.264/VP9/VP8) or MJPEG image bytes |
| `media/{userId}/mic` | Microphone audio | Raw audio samples (s16le, 16 kHz mono) |
| `media/{userId}/screen` | Screen share | Same video formats as `cam` |

Track separation lets playback select only what a layout needs (e.g. active-speaker uses audio independently of video; screen-share focus reads the screen track alone).

Video payloads are **self-describing**:

- **WebCodecs chunk** (preferred): `EVC1` magic + 18-byte header (codec id, keyframe flag, dimensions, timestamp µs) + elementary stream (annexb H.264, or VP9/VP8).
- **MJPEG image** (fallback): raw JPEG/WebP bytes.

The server relay stores and forwards payloads opaquely — it never re-encodes.

### 2. Event timeline (`events` stream)

A versioned, append-only JSON timeline. Every event shares one envelope:

```json
{ "v": 1, "type": "participant.joined", "ts": 1699876000123, "actor": "alice", "data": {} }
```

Core event types:

| Type | Purpose |
|------|---------|
| `participant.joined` / `participant.left` | Lifecycle |
| `track.added` / `track.removed` | Track lifecycle (kind: `cam`, `mic`, `screen`) |
| `screenshare.started` / `screenshare.stopped` | Screen-share lifecycle |
| `speaking.started` / `speaking.stopped` | Voice activity (from client-side VAD at capture) |
| `director.cut` | Computed layout decision (focus participant, reason) |
| `caption.cue` | Timed caption overlay |
| `branding.applied` | Logo/watermark/accent colour |
| `highlight.marked` | Highlight clip span |
| `participant.redacted` | GDPR/removal annotation |

The `meta` stream still receives legacy join/leave/control records for compatibility, but **`events` is the canonical timeline** for reconstructing session state.

### 3. Director track (`director` stream)

When a session ends, the server computes active-speaker cuts from the speaking timeline and persists them as `director.cut` events. This makes active-speaker playback **deterministic and cheap** — clients read pre-computed focus switches instead of re-deriving them from raw audio on every replay.

Director cuts honour a minimum hold time (default 1.5 s) to avoid rapid switching.

### 4. Chat (`chat` stream)

Text messages between participants as JSON (base64-encoded):

```json
{"text": "Hello everyone!", "userId": "bob", "timestamp_ms": 1234567}
```

### 5. Captions (`captions/{lang}` stream, optional)

Per-language caption cues, using the same `caption.cue` event envelope. The session manifest advertises a captions stream when present.

## Session Manifest

`GET /v1/sessions/:sessionId/manifest` returns the **reconstruction contract** — enough to bootstrap playback without reading every stream:

```json
{
  "data": {
    "version": 1,
    "session_id": "sess-abc123",
    "epoch": 1699876000000,
    "start_ts": 1699876000000,
    "end_ts": 1699876300000,
    "participants": [
      {
        "user_id": "alice",
        "tracks": [
          { "kind": "cam", "stream_id": "a/.../media/alice/cam", "start_ts": 1000, "end_ts": 5000 },
          { "kind": "mic", "stream_id": "a/.../media/alice/mic", "start_ts": 1100, "end_ts": 4980 }
        ]
      }
    ],
    "streams": {
      "events": "a/.../events",
      "chat": "a/.../chat",
      "director": "a/.../director",
      "captions": "a/.../captions/en"
    },
    "capabilities": [
      "separated-tracks",
      "event-timeline",
      "director-track",
      "configurable-layouts",
      "captions"
    ]
  },
  "error": null
}
```

Use the manifest to fetch only the tracks a layout needs (bandwidth win) and to discover optional features (director track, captions).

Public sessions: `GET /v1/public/sessions/:sessionId/manifest`.

## Session Lifecycle & Recording States

| State | Description | Timestamp Field |
|-------|-------------|-----------------|
| **Created** | Session created, no participants yet | `created_at` set |
| **Started** | First participant joined | `started_at` set |
| **Streaming** | Active session with participants connected | — |
| **Ended / Recording Ready** | Last participant left, recording complete; director track computed | `ended_at` set |

A session's "recording" is the persisted streams between start and end. To check if a recording is ready: look for `ended_at !== null`.

## Accessing Recordings (Replay)

### Step 1: Get the manifest (recommended)

```bash
curl https://api.synento.com/v1/sessions/{sessionId}/manifest \
  -H "Authorization: Bearer sk_live_xxxxx"
```

Returns participants, separated track index, time bounds, and capabilities.

Alternatively, use the lighter playback timeline for stream IDs and bounds:

```bash
curl https://api.synento.com/v1/sessions/{sessionId}/playback/timeline \
  -H "Authorization: Bearer sk_live_xxxxx"
```

### Step 2: Read playback windows (iterative)

The playback API reads all relevant streams (events, director, chat, and every participant track) in time-based windows:

```bash
curl "https://api.synento.com/v1/sessions/{sessionId}/playback?from_ts=0&window_size=30000" \
  -H "Authorization: Bearer sk_live_xxxxx"
```

Continue while `next_base_seq` is non-null. When it is `null`, playback is complete.

Each event includes `stream_id` (full S2 path), `timestamp`, and `body_base64`. Media track events carry raw payload bytes (no outer wire header).

### Step 3: Compose the view (client side)

Rather than rendering one participant at a time, the **composition layer** decodes the tracks a layout needs and draws them into a single canvas:

1. **Reduce session state** — process `events` and `director` events up to the playhead (participants, active speaker, screen sharers).
2. **Resolve layout** — map a layout descriptor + session state to normalized regions (0..1 coordinates).
3. **Decode tracks** — WebCodecs `VideoDecoder` for `EVC1` chunks, `createImageBitmap` for MJPEG; one offscreen surface per visible track.
4. **Draw** — composite regions onto the output canvas; optional caption and branding overlays.

Available layout types:

| Layout | Behaviour |
|--------|-----------|
| `active-speaker` | Focus the current speaker (from director/speaking events); thumbnails for others |
| `presenter` | Focus a designated participant |
| `grid` | Equal tiles for all active cameras |
| `screen-share` | Full-frame screen share with camera thumbnails; falls back to active-speaker |
| `selected` | Grid of chosen participants only |
| `custom` | Explicit normalized regions |

The dashboard replay player (`web/app`) exposes a layout selector (Speaker / Grid / Presenter / Screen). The player SDK provides the same engine for embeds.

## Replay Controls

- **Seek** — jump to any timestamp (`from_ts` parameter)
- **Speed control** — play at 0.5x, 1x (default), or 2x
- **Layout switching** — change composition instantly without re-fetching (client-side)
- **Pause / resume** — stop consuming events without losing position

## Player SDK (`@synento/player`)

> **Two limits to know before you choose this path.**
>
> **In-browser replay is video-only.** The player composes camera and screen
> tracks; it does not decode the microphone tracks. Audio exists in the
> **MP4 export**, which is muxed server-side. If your users need to *hear* a
> recording, export it and serve the file — see
> [Exporting to MP4](#exporting-to-mp4) — rather than reaching for the player.
>
> **`PlaybackEngine` authenticates with an API key**, so the snippet below is
> for **trusted, server-side or first-party contexts only**. Never ship a
> `sk_live_` key to a browser. For a customer-facing embed today, proxy the
> playback endpoints through your own backend (which holds the key) exactly as
> the Synento dashboard does, or make the session public and read the
> unauthenticated `/v1/public/sessions/:id/...` endpoints.

The player SDK wraps windowed playback and provides a full composition stack:

```typescript
import {
  PlaybackEngine,
  Compositor,
  resolveLayout,
  reduceSessionState,
  parseSessionEvents,
  availableTracks,
  type LayoutDescriptor,
} from "@synento/player";

const engine = new PlaybackEngine(sessionId, apiBase, apiKey);
await engine.fetchTimeline();

const canvas = document.querySelector("canvas")!;
const compositor = new Compositor({ canvas, width: 1280, height: 720 });

const descriptor: LayoutDescriptor = {
  type: "active-speaker",
  orientation: "landscape",
  captions: true,
};

engine.isPlaying = true;
await engine.play(() => {
  // Pass accumulated window events + current playhead to the compositor.
  compositor.render(descriptor, windowEvents, engine.currentTimeMs);
});

compositor.dispose();
```

Lower-level helpers (`resolveLayout`, `reduceSessionState`, `availableTracks`, `resolveActiveCaption`, `resolveBranding`) are pure and reusable — the server export layout engine shares the same vocabulary.

See [`@synento/player`](https://www.npmjs.com/package/@synento/player) for the full API.

## Exporting to MP4

Export is **async** and **parameterized by layout**. The web dashboard can enqueue and download exports while logged in (session cookie via the API proxy). For scripts and integrations, pass a project API key on each request:

```bash
# Enqueue export (returns job id)
curl "https://api.synento.com/v1/sessions/{sessionId}/export?layout=presenter&orientation=vertical&focus=alice" \
  -H "Authorization: Bearer sk_live_xxxxx"
```

Response (202):

```json
{ "data": { "job_id": "..." }, "error": null }
```

Query parameters:

| Parameter | Values | Default |
|-----------|--------|---------|
| `layout` | `grid`, `presenter` | `grid` |
| `orientation` | `landscape` (1280×720), `vertical` (720×1280) | `landscape` |
| `focus` | Participant user id (for presenter layout) | first participant |

Poll status:

```bash
curl https://api.synento.com/v1/exports/{jobId} \
  -H "Authorization: Bearer sk_live_xxxxx"
```

When `status` is `completed`, download via `download_url`:

```bash
curl https://api.synento.com/v1/exports/{jobId}/download \
  -H "Authorization: Bearer sk_live_xxxxx" \
  --output session.mp4
```

Export behaviour:

- Reads separated `cam` and `mic` tracks per participant
- **Codec-aware**: H.264 sessions mux annexb directly; MJPEG sessions use `image2pipe` with frame-gap repetition
- Composes video via FFmpeg using the same layout vocabulary as the client compositor
- Export artifacts are stored by Synento and served from the download endpoint

## Participant redaction

Remove a participant's media while keeping the rest of the session:

```bash
curl -X DELETE https://api.synento.com/v1/sessions/{sessionId}/participants/{userId} \
  -H "Authorization: Bearer sk_live_xxxxx"
```

This deletes all of the participant's track streams (`cam`, `mic`, `screen`) and appends a `participant.redacted` event to the timeline. Playback and export skip redacted tracks and can render a placeholder.

## Architecture diagram

```mermaid
flowchart TB
  subgraph rec [Recording layer — layout agnostic]
    cam[media/userId/cam]
    mic[media/userId/mic]
    scr[media/userId/screen]
    ev[events timeline]
    dir[director track]
  end
  rec --> comp[Composition layer]
  comp --> lay[Layout descriptor]
  comp --> cli[Client compositor — @synento/player]
  comp --> srv[Server exporter — FFmpeg]
  lay --> cli
  lay --> srv
  cli --> view[Interactive playback]
  srv --> art[MP4 / vertical / clips]
```
