/work Systems
Frolic Multiplayer Framework
Custom and open-source multiplayer frameworks powering real-time gameplay across the Frolic gaming platform.
Scope
Shared multiplayer substrate for multiple game genres and teams
Team context
Led the framework contract while game teams owned game logic and frontend integration.
Owned
- Framework architecture
- State-sync and reconnect contracts
- Transport and room/session layers
Constraints
- Different game types need different latency tolerances
- Mobile disconnects must recover without corrupting state
- Game teams should not need distributed-systems expertise
Decisions
- Authoritative server with delta replication
- Standard WebSocket primitives with framework logic on top
- Schema-declared state instead of free-form messages
Outcome
- Framework reused across shipped game titles and genres
- New game teams consumed the substrate without backend rewrites
Concurrency and latency numbers are operationally sensitive; outcomes are described qualitatively.
Case study narrative
Problem & Constraints
Frolic is a multi-game platform — each game is a different shape (turn-based, real-time, head-to-head, multi-player rooms), but they all need to share the same multiplayer substrate. Building a one-off backend per game means rewriting auth, matchmaking, room state, anti-cheat, and reconnection for every title.
The constraints:
- Latency varies by game type. Real-time twitch games need different tolerances than turn-based card games. The framework had to expose the right knobs without forcing game teams to learn distributed-systems theory.
- State has to survive disconnects. Mobile networks drop, players background apps, devices die. Reconnect-and-resume must be cheap and correct.
- Operational cost matters. A multiplayer framework that requires a dedicated SRE per game is not a framework — it’s a maintenance trap.
- Game teams shouldn’t need to think about backends. The contract had to be small enough that a game developer can model their gameplay against it without learning the internals.
System architecture
Shared multiplayer substrate
- Game client Sends intent and renders replicated state Untrusted edge
- Connection layer Session identity, ordering, heartbeat, reconnect Transport contract
- Room/session Lifecycle, membership, ready state, matchmaking handoff Framework contract
- State engine Authoritative transitions, deltas, snapshots, replay Consistency boundary
- Game logic Pure domain rules supplied by each game team Extension point
- Persistence/ops Durable results, diagnostics, and session history Operational boundary
Approach & Architecture
The framework split responsibilities cleanly:
- Connection + transport layer — handled WebSocket lifecycle, reconnect, message ordering, backpressure. Game code never touched sockets directly.
- Room + session layer — matchmaking, room creation, player join/leave, ready states, lifecycle hooks the game subscribed to.
- State sync layer — authoritative server state with delta replication. Game code declared its state schema; the framework handled sync, conflict resolution, and snapshot/replay.
- Game logic layer — the only place game teams wrote real code. Pure functions that consumed state and player events and produced state transitions.
Where it made sense, we leaned on existing open-source primitives. Where the open-source choice would have meant fighting the framework or paying a runtime cost we couldn’t afford, we built the piece ourselves. The split was deliberate: don’t reinvent transports, don’t accept somebody else’s gameplay model.
Illustrative state contract
The framework contract makes player intent explicit and keeps authoritative state on the server. Game code does not send arbitrary socket messages.
type PlayerIntent = {
roomId: string;
playerId: string;
sequence: number;
command: "MOVE" | "PLAY_CARD" | "READY";
payload: unknown;
};
type StateTransition<S> = {
previousVersion: number;
nextVersion: number;
state: S;
emittedEvents: DomainEvent[];
};
// Illustrative only — each game supplies its own reducer.
function applyIntent<S>(state: S, intent: PlayerIntent): StateTransition<S> {
assertMember(intent.playerId, intent.roomId);
assertNextSequence(intent.sequence);
return gameReducer(state, intent);
}
Sequence numbers make duplicate delivery detectable. State versions make stale writes visible. The reducer boundary gives game teams a testable model while the framework owns transport, replication, and recovery.
Request sequence
Disconnect and resume
- 01 Clientloses transport
Keeps the last applied state version and stops optimistic advancement.
- 02 Connection layerexpires heartbeat
Marks the member disconnected without immediately destroying the room.
- 03 Clientreconnects with session token
Presents room identity and last confirmed state version.
- 04 State enginecomputes recovery
Sends missing deltas when safe or a complete snapshot when versions diverge.
- 05 Room lifecyclerestores membership
Rejoins the player only after state acknowledgement.
My Role
I owned the framework architecture and the contract between framework and game teams. That meant designing the abstractions, building the core layers (transport, room/session, state sync), and supporting the first wave of games that consumed it.
What the team owned: individual game logic, art and frontend integration, infrastructure operation, and the matchmaking heuristics tuned per game.
Tradeoffs & Decisions
Authoritative server vs lockstep. Considered both. Lockstep is cheaper at runtime but punishing to develop against on mobile networks where latency is variable. We picked authoritative server with delta replication — more bandwidth but dramatically simpler reasoning for game teams.
Custom transport vs off-the-shelf. We did not build our own transport — we used standard WebSocket primitives and put framework logic on top. Building a transport from scratch is one of the most common ways multiplayer frameworks burn a year of engineering time without shipping a game.
Schema-declared state vs free-form messages. Schema-declared was harder to migrate but every other property of the framework — replay, debug, conflict resolution, snapshot — fell out for free. Worth it.
Tradeoff matrix
Authoritative server versus lockstep
| Option | Strengths | Costs | Decision |
|---|---|---|---|
| Authoritative serverChosen | Clear ownership, easier anti-cheat, simple reconnect reasoning | Higher server compute and bandwidth | Selected for the shared framework |
| Deterministic lockstep | Compact network traffic and distributed simulation | Strict determinism, latency sensitivity, difficult mobile recovery | Rejected as the universal default |
| Peer authority | Low infrastructure cost for small rooms | Trust, host migration, and fairness complexity | Unsuitable for the platform contract |
Observability and migration discipline
Framework-level telemetry needs to explain a player session, not merely a process. Useful identifiers include room, player, connection generation, state version, command sequence, and game build. With those fields, a support report such as “my turn disappeared after reconnect” can be reconstructed as an ordered state story.
State-schema changes are the dangerous edge. A safe evolution path includes:
- Readers that tolerate the previous schema version.
- Explicit migration functions with fixture-based replay tests.
- Snapshot versioning so a rollback does not strand active rooms.
- A compatibility window across server and client releases.
- Metrics for rooms still running older schemas before removal.
Outcome
The framework supported multiple shipped game titles across genres without per-game backend rewrites. Specific concurrency and latency numbers are operational data and not publicly disclosable, but the framework met its design targets and continued to absorb new game types after the initial cohort.
Bounded proof
Platform leverage
- Reuse
- Multiple genres
- Recovery
- Reconnect-first
- Team contract
- Logic only
One substrate supported different gameplay models without one backend per title.
Mobile network interruption was designed into the normal session lifecycle.
Game teams focused on domain rules while the framework owned distributed concerns.
What I’d Revisit
- Clearer migration story for state schema. State migrations were possible but not pleasant. With more time I would have built the migration tooling alongside the schema declaration, not after.
- Earlier observability investment. We added debug tooling reactively. Reactive debug tooling is a tax you pay forever. The framework should ship with player-session replay from day one.
- More opinionated matchmaking primitives. We stayed neutral on matchmaking and let each game build its own. In retrospect a strong default with escape hatches would have shortened time-to-launch for new titles.
Want to discuss this work in more detail? Get in touch.
Back to all work