Synento docs

Core Concepts

Understanding the building blocks of Synento.

View as Markdown

Understanding the building blocks of Synento.

Architecture Overview

Synento follows an API-first, agent-native architecture: every feature is exposed through REST endpoints, the dashboard is simply a client of those same APIs, and nothing requires a human in that dashboard to set up — a coding agent can create a project, publish media, and read it back through the same calls.

┌──────────────┐    REST     ┌─────────────┐   WebSocket    ┌──────────────┐
│  CLI / SDK   │ ◄────────► │  Synento    │ ◄─────────────►│  WebSocket   │
│  Dashboard   │            │  Server     │                │    Hub       │
│  External    │            │  (Fastify)  │                │   (Hub.ts)   │
│  Integrations│            └─────────────┘                └──────────────┘
└──────────────┘                                                      │
                                                                    Recording
                                                                     (StreamStore)

Hierarchy: Organisation → Project → Session

Organisation

The top-level container representing a team, company, or account. All projects belong to an organisation.

  • Created automatically when you register via the dashboard
  • Billing applies at the organisation level (free tier: 3,000 participant-minutes and 50 GiB, no card required)

Project

A scoped workspace under an organisation. Each project has:

  • Environment (e.g., development, staging, production) — optional, for your own organization
  • API Key (secret key: sk_live_xxxxx) — used to authenticate all API requests for this project

Projects are the unit of billing. Usage is tracked per project and aggregated to the organisation level for billing purposes.

Session

The core unit of live streaming and recording. Each session:

  • Has a unique ID (UUID format)
  • Can have multiple participants connected simultaneously via WebSocket
  • Is automatically recorded as it happens — no separate recording setup needed (stream-native)
  • Is instantly replayable once it ends, using windowed playback API
type Session = {
  id: string;            // Unique session ID (UUID)
  projectId: string;     // Parent project
  name: string;          // Display name (optional)
  createdBy: string;     // Creator identifier
  isPublic?: boolean;    // Whether anyone can access this session's replay
  startedAt: string | null; // When first participant joined
  endedAt: string | null;   // When last participant left
  createdAt: string;     // Creation timestamp
};

The Stream-Native Recording Model

This is what makes Synento fundamentally different from traditional video platforms.

Key Principle: Recording = Persisted Stream

There is no separate recording pipeline. When participants stream via WebSocket, every message is immediately persisted to the StreamStore (a multi-stream append log). This same stored data powers both:

  1. Replay — iterating the stream in chronological (or variable-speed) order
  2. Export — reading the full stream and encoding it to MP4

Stream Types

Synento persists a layout-agnostic recording: separated media tracks plus a typed event timeline. Presentation (grid, active-speaker, etc.) is resolved at playback time.

StreamPurposeExample
media/{userId}/camCamera video per participantWebCodecs EVC1 or MJPEG bytes
media/{userId}/micMicrophone audio per participantRaw s16le audio
media/{userId}/screenScreen share per participantSame video formats as cam
eventsStructured session timeline{ "v": 1, "type": "speaking.started", "actor": "alice", … }
directorComputed layout decisions{ "type": "director.cut", "data": { "focus": "bob" } }
chatText messages{ "text": "Hello!", "userId": "u1" }
captions/{lang}Caption cues (optional){ "type": "caption.cue", "data": { "text": "…", "end": … } }
metaLegacy join/leave/controlJSON (superseded by events for new integrations)

Composition Layer

The recording stores what happened; playback decides how to show it:

Tracks + events  →  Session state  →  Layout descriptor  →  Composed view
                         ↑                    ↑
                   director track      grid / active-speaker / presenter / …
  • Client: @synento/player Compositor + LayoutEngine decode visible tracks and draw regions.
  • Server: FFmpeg export uses the same layout vocabulary (grid, presenter, orientation).
  • Manifest: GET /v1/sessions/:id/manifest indexes tracks, bounds, and capabilities for fast bootstrap.

What This Means for Developers

  • No encoding delay — replay is available immediately after a session ends
  • Instant layout switching — change grid ↔ active-speaker without re-fetching
  • Selective track fetch — manifest-driven bandwidth savings (fetch only visible tracks)
  • Windowed reading — large sessions don't need to be loaded entirely
  • Variable speed replay — replay at 0.5x, 2x, or any speed using time-based buffering

Access Control Model

Sessions support three levels of access:

LevelWho Can AccessHow
Private (default)Only holders of a valid connection tokenAPI key + session ID
PublicAnyone with the session URLNo auth required
Link-onlyOnly holders of an expiring replay linkToken in URL (TTL 1 min – 30 days)

See Access Control Guide for detailed API usage.

Connection Tokens & WebSocket Authentication

Every participant needs a connection token to join or replay a session:

  1. Request via POST /v1/sessions/{id}/connection with API key auth
  2. Receive an access_token (JWT) that expires in a configurable time window
  3. Connect to WebSocket: ws://host/ws?session={sessionId}&access_token={token}&user_id=alice

Tokens include claims for:

  • user_id — participant identifier
  • session_id — the session to join
  • Mode-specific parameters: replay, from_ts, speed for replay mode

Usage Tracking & Billing

Usage is tracked per project and aggregated at the organisation level:

  • Streaming minutes: calculated from started_at to ended_at per session
  • Storage bytes: size of exported MP4 files

Billing tiers — the authoritative source is GET /v1/billing/plans:

TierParticipant-minutesStorageRetentionCost
Free3,000 included50 GiBup to 30 days$0, no card
Pro4,000 included50 GiBup to 90 days$19/mo, then $0.004/participant-minute and $0.06/GiB-month

Two things worth being precise about, because both have caused real errors:

  • Usage is metered in participant-minutes, not room-minutes: a four-person call for ten minutes is 40, not 10.
  • Storage is GiB (2³⁰ bytes), the unit S2 bills in, and the API fields say so (includedStorageGib, storageUsedGib). Only storageBytesUsed is bytes. Scaling a GiB figure by 1e9 renders a 50 GiB allowance as "46.57 GB".

Replay API & Windowed Playback

The replay system uses window-based iteration to handle sessions of any size:

  1. GET /v1/sessions/:id/manifest — reconstruction contract (participants, tracks, bounds, capabilities)
  2. GET /v1/sessions/:id/playback/timeline — compact metadata (participants, duration, stream ids)
  3. GET /v1/sessions/:id/playback?from_ts=0&window_size=30000 — read 30-second windows of all tracks + events
  4. Repeat until next_base_seq is null

Each window returns events from separated media tracks, the events timeline, director track, and chat. The Player SDK composes them into configurable layouts (grid, active-speaker, presenter, screen-share).

See Recordings & Replay Guide for layout descriptors, export parameters, and participant redaction.

On this page