/work Systems

Frolic Multiplayer Framework

Custom and open-source multiplayer frameworks powering real-time gameplay across the Frolic gaming platform.

Role
Senior Engineer
Company
Frolic
Period
2021 — 2023
Status
Live
Node.js Unity (C#) WebSocket Redis Distributed Systems

Scope

Shared multiplayer substrate for multiple game genres and teams

Team context

Led the framework contract while game teams owned game logic and frontend integration.

Owned

  • Framework architecture
  • State-sync and reconnect contracts
  • Transport and room/session layers

Constraints

  • Different game types need different latency tolerances
  • Mobile disconnects must recover without corrupting state
  • Game teams should not need distributed-systems expertise

Decisions

  • Authoritative server with delta replication
  • Standard WebSocket primitives with framework logic on top
  • Schema-declared state instead of free-form messages

Outcome

  • Framework reused across shipped game titles and genres
  • New game teams consumed the substrate without backend rewrites

Concurrency and latency numbers are operationally sensitive; outcomes are described qualitatively.


Case study narrative

Problem & Constraints

Frolic is a multi-game platform — each game is a different shape (turn-based, real-time, head-to-head, multi-player rooms), but they all need to share the same multiplayer substrate. Building a one-off backend per game means rewriting auth, matchmaking, room state, anti-cheat, and reconnection for every title.

The constraints:

  • Latency varies by game type. Real-time twitch games need different tolerances than turn-based card games. The framework had to expose the right knobs without forcing game teams to learn distributed-systems theory.
  • State has to survive disconnects. Mobile networks drop, players background apps, devices die. Reconnect-and-resume must be cheap and correct.
  • Operational cost matters. A multiplayer framework that requires a dedicated SRE per game is not a framework — it’s a maintenance trap.
  • Game teams shouldn’t need to think about backends. The contract had to be small enough that a game developer can model their gameplay against it without learning the internals.

System architecture

Shared multiplayer substrate

  1. Game client Sends intent and renders replicated state Untrusted edge
  2. Connection layer Session identity, ordering, heartbeat, reconnect Transport contract
  3. Room/session Lifecycle, membership, ready state, matchmaking handoff Framework contract
  4. State engine Authoritative transitions, deltas, snapshots, replay Consistency boundary
  5. Game logic Pure domain rules supplied by each game team Extension point
  6. Persistence/ops Durable results, diagnostics, and session history Operational boundary
Illustrative logical architecture. It shows contracts and state ownership, not the private production topology.

Approach & Architecture

The framework split responsibilities cleanly:

  • Connection + transport layer — handled WebSocket lifecycle, reconnect, message ordering, backpressure. Game code never touched sockets directly.
  • Room + session layer — matchmaking, room creation, player join/leave, ready states, lifecycle hooks the game subscribed to.
  • State sync layer — authoritative server state with delta replication. Game code declared its state schema; the framework handled sync, conflict resolution, and snapshot/replay.
  • Game logic layer — the only place game teams wrote real code. Pure functions that consumed state and player events and produced state transitions.

Where it made sense, we leaned on existing open-source primitives. Where the open-source choice would have meant fighting the framework or paying a runtime cost we couldn’t afford, we built the piece ourselves. The split was deliberate: don’t reinvent transports, don’t accept somebody else’s gameplay model.

Illustrative state contract

The framework contract makes player intent explicit and keeps authoritative state on the server. Game code does not send arbitrary socket messages.

type PlayerIntent = {
  roomId: string;
  playerId: string;
  sequence: number;
  command: "MOVE" | "PLAY_CARD" | "READY";
  payload: unknown;
};

type StateTransition<S> = {
  previousVersion: number;
  nextVersion: number;
  state: S;
  emittedEvents: DomainEvent[];
};

// Illustrative only — each game supplies its own reducer.
function applyIntent<S>(state: S, intent: PlayerIntent): StateTransition<S> {
  assertMember(intent.playerId, intent.roomId);
  assertNextSequence(intent.sequence);
  return gameReducer(state, intent);
}

Sequence numbers make duplicate delivery detectable. State versions make stale writes visible. The reducer boundary gives game teams a testable model while the framework owns transport, replication, and recovery.

Request sequence

Disconnect and resume

  1. 01
    Clientloses transport

    Keeps the last applied state version and stops optimistic advancement.

  2. 02
    Connection layerexpires heartbeat

    Marks the member disconnected without immediately destroying the room.

  3. 03
    Clientreconnects with session token

    Presents room identity and last confirmed state version.

  4. 04
    State enginecomputes recovery

    Sends missing deltas when safe or a complete snapshot when versions diverge.

  5. 05
    Room lifecyclerestores membership

    Rejoins the player only after state acknowledgement.

The reconnect path is designed as normal lifecycle behavior because mobile suspension and network changes are expected, not exceptional.

My Role

I owned the framework architecture and the contract between framework and game teams. That meant designing the abstractions, building the core layers (transport, room/session, state sync), and supporting the first wave of games that consumed it.

What the team owned: individual game logic, art and frontend integration, infrastructure operation, and the matchmaking heuristics tuned per game.

Tradeoffs & Decisions

Authoritative server vs lockstep. Considered both. Lockstep is cheaper at runtime but punishing to develop against on mobile networks where latency is variable. We picked authoritative server with delta replication — more bandwidth but dramatically simpler reasoning for game teams.

Custom transport vs off-the-shelf. We did not build our own transport — we used standard WebSocket primitives and put framework logic on top. Building a transport from scratch is one of the most common ways multiplayer frameworks burn a year of engineering time without shipping a game.

Schema-declared state vs free-form messages. Schema-declared was harder to migrate but every other property of the framework — replay, debug, conflict resolution, snapshot — fell out for free. Worth it.

Tradeoff matrix

Authoritative server versus lockstep

OptionStrengthsCostsDecision
Authoritative serverChosen Clear ownership, easier anti-cheat, simple reconnect reasoningHigher server compute and bandwidthSelected for the shared framework
Deterministic lockstep Compact network traffic and distributed simulationStrict determinism, latency sensitivity, difficult mobile recoveryRejected as the universal default
Peer authority Low infrastructure cost for small roomsTrust, host migration, and fairness complexityUnsuitable for the platform contract
The choice optimized for game-team ergonomics and unreliable mobile networks rather than minimum server cost.

Observability and migration discipline

Framework-level telemetry needs to explain a player session, not merely a process. Useful identifiers include room, player, connection generation, state version, command sequence, and game build. With those fields, a support report such as “my turn disappeared after reconnect” can be reconstructed as an ordered state story.

State-schema changes are the dangerous edge. A safe evolution path includes:

  1. Readers that tolerate the previous schema version.
  2. Explicit migration functions with fixture-based replay tests.
  3. Snapshot versioning so a rollback does not strand active rooms.
  4. A compatibility window across server and client releases.
  5. Metrics for rooms still running older schemas before removal.

Outcome

The framework supported multiple shipped game titles across genres without per-game backend rewrites. Specific concurrency and latency numbers are operational data and not publicly disclosable, but the framework met its design targets and continued to absorb new game types after the initial cohort.

Bounded proof

Platform leverage

Reuse
Multiple genres

One substrate supported different gameplay models without one backend per title.

Recovery
Reconnect-first

Mobile network interruption was designed into the normal session lifecycle.

Team contract
Logic only

Game teams focused on domain rules while the framework owned distributed concerns.

Public-safe evidence focuses on reuse and ownership rather than private concurrency figures.

What I’d Revisit

  • Clearer migration story for state schema. State migrations were possible but not pleasant. With more time I would have built the migration tooling alongside the schema declaration, not after.
  • Earlier observability investment. We added debug tooling reactively. Reactive debug tooling is a tax you pay forever. The framework should ship with player-session replay from day one.
  • More opinionated matchmaking primitives. We stayed neutral on matchmaking and let each game build its own. In retrospect a strong default with escape hatches would have shortened time-to-launch for new titles.


Want to discuss this work in more detail? Get in touch.

Back to all work