Gessa Docs
Product · Explanation

Concept

Explanation: Graph Room Sequencer

How Gessa runs one live collaboration session per document with a single-owner lease, an epoch fence as the real safety boundary, an owner-authoritative presence roster, one multiplexed durable transport, single-hop owner proxying, and headless AI reflection.
engine v1.0.234since v1Copy for LLM
engine v1.0.232

Last verified 2026-09-03 against engine v1.0.232.

When several people, and the AI, edit one Gessa document at once, the edits have to converge on one authoritative model without a partitioned server ever committing a stale write. The Graph Room Sequencer (GRS) is the runtime that owns that live collaboration session per document. It elects one owner per document, keeps an owner-held presence roster, carries every channel for the open document over one multiplexed WebSocket, replays durable channels on reconnect, meters abusive connections, routes a write to the owner when it arrives on the wrong pod, drains gracefully, and reports itself to CloudWatch. This page explains the model: who owns a document, why the lease is not the safety boundary, how presence and the durable channels behave, how a misrouted write reaches the owner, and what is deliberately not built. Every concrete field, constant, and capacity number lives in the Graph Room Sequencer reference, which this page links down to and never restates.

The framing from Backend Authority carries through: the server owns the document and the client is a non-authoritative view. GRS is the discipline that keeps one server the writer while many clients watch and propose. It is the collaboration-session sibling of the Netcode Model: netcode runs the playable room, GRS runs the authoring room.

One owner per document, and why the lease is not the safety boundary

Each document, keyed by its gameId, has exactly one authoritative owner on one pod. The owner holds a lease that records the pod, a monotonic epoch, a routable WebSocket endpoint, and its lease and expiry times. The epoch is the load-bearing value: it is monotonic per document and it increments only on a fresh takeover, staying put on a renew and on an idempotent re-acquire by the process that already holds it.

A lease can be lost to a network partition. A zombie leader that still believes it holds the lease must not be able to commit. GRS makes that impossible by treating the lease as a liveness and routing hint and putting safety at the durable write: every durable graph write stamps the owner epoch, and the durable write rejects a write whose epoch is behind the stored fence. So the worst a partitioned leader can do is fail its own commit. The fence decision and the durable owner-epoch record live in ProjectGraphService, which GRS composes rather than owns.

Because a store can be rebuilt after data loss, a fresh takeover does not simply add one to the previous epoch; it floors the new epoch to wall-clock time so it can never restart below a fence a durable database already advanced past:

epoch=max(prevEpoch+1,⌊nowMs⌋)

The in-memory store is the reference implementation of this contract; RedisLeaseStore implements the identical contract atomically with server-side Lua registered as named commands and a persistent counter for the epoch, so the same guarantees hold across pods.

The maintenance loop and failover

Every room runs a single self-rescheduling timer. One tick renews or re-acquires the lease, and then, only if this pod is the owner, reaps expired presence peers. The cadence is not taken at face value; it is clamped so at least one renewal always lands inside a lease lifetime:

maintainIntervalMs=min(max(1000,requested),max(1000,⌊ttlMs/2⌋))

The 1000 ms floor is real: for a lease TTL of 2000 ms or more the cadence is at most half the TTL, but for a shorter TTL the floor dominates and the "at most half" property does not hold. Production runs a much longer lease, so the property holds there; the reference states the exact formula and names the test that pins it.

Failover is implicit rather than orchestrated. A dead owner stops renewing, its lease lapses, and the next surviving pod's tick acquires the document with a strictly higher epoch. The old owner, if it ever comes back, is already fenced out at the durable write.

The presence roster

The roster of who is in a document is owner-held, with a per-peer TTL. Correctness splits along one line. Join, update, heartbeat, and leave all mutate the one shared roster and are pod-agnostic, so a joining client reads the complete roster in one read with no delta reconstruction to get wrong. TTL-driven reaping, by contrast, is owner-only: a non-owner's reap tick is a no-op, so an expiry-driven departure has exactly one source and cannot be double-emitted from two pods. A reap-versus-rejoin race is handled explicitly: if a peer that was just reaped appears again in the post-sweep roster, its stale departure is suppressed rather than flapped. The Redis roster mirrors this with a hash keyed by document plus a TTL-scored set for expiry.

The same heartbeat carries one more piece of liveness. A session's exclusive resource locks are renewed off its presence pings, and a presence lock update re-syncs them, so a collaborator that goes silent loses its held locks by TTL rather than holding them forever. The lock primitive itself lives in the project-graph service, not GRS; GRS owns only the coupling of lock renewal to the heartbeat. The reference names the exact functions.

One multiplexed transport

Every channel for an open document rides one WebSocket on a single session route, and every frame is the same small shape: a channel, an optional sequence number, and a payload. A single normative policy, channelDurability, classifies each channel once and defines "logged", "replayable", and "durable" as one predicate, so no code can treat a channel as replayable but not logged. The durable channels are the project-graph op log and the chat, review, and layout timelines; the ephemeral channels are presence and cursor; control carries heartbeats, in-stream auth revalidation, and the owner-moving handoff.

Two rules make the durable channels trustworthy. First, a durable frame must carry a sequence number and an ephemeral or control frame must not, and the multiplexer tracks a high-watermark per durable channel. Second, outbound delivery is two-phase: the transport asks whether a frame may be admitted, sends it, and only then marks it delivered, so a serialization or socket failure can never silently advance a cursor past a frame the client did not receive.

On reconnect, a client resumes rather than reloads when it safely can. planChannelResume computes the contiguous, deduplicated window after the client's last acknowledged sequence up to the current head. If that window cannot be covered contiguously, or the client is somehow ahead of the server, or a retained sequence is not a whole number, it signals a gap and the client cold-hydrates from a fresh snapshot instead of stitching a broken log.

Everything the server sends is fail-closed. Outbound frames pass through schema validation and the asset client-boundary monitor before they leave, and the op-result and caught-up frames are treated as critical control: if backpressure would drop one, the session forces a resync rather than dropping it silently.

The write path: verify, then apply or proxy

A write does not trust the local view of ownership. When an op frame arrives, the pod calls a verified-ownership check that returns an epoch only if the cached lease atomically renews in the shared store at that moment, closing the window where a lease has just moved. If the renew fails, the pod clears its cached lease and treats itself as a non-owner.

A non-owner never applies the op locally, which is what prevents split-brain. Instead, if the lease names a routable owner endpoint, the pod forwards the exact op one hop to the owner over the owner's normal session socket, guarded by a hop header so a forwarded write is never forwarded again. A pod:-scheme endpoint is not WebSocket-routable and the proxy refuses it rather than fabricating a write. If there is no routable owner, the client is simply told it is not the owner and asked to retry.

When the owner applies the op and the durable write reports an epoch-fence rejection carrying the stored epoch, the owner repairs its epoch to one above the stored fence and retries the op once. The durable op payloads are project-graph mutations that ProjectGraphService applies against the authored graph under its own conflict and revision policy; the canonical authoring verbs behind those mutations are in the Action Catalog.

ReferenceAction CatalogResolved signature, schema and example

Fairness

A per-connection token bucket keeps one busy connection from starving the others. The op-write rate is set lower for an AI actor than a human, so an eager agent can never crowd out a person's local echo; the cursor channel gets a generous ephemeral budget; and chat and review share one budget. Presence and control frames are never metered because they are heartbeats and auth. One durable channel, layout, is also not metered; the reference lists it plainly among the unmetered channels so the gap is documented rather than implied.

Headless AI presence

An AI agent commits authoring transactions on the server with no WebSocket session of its own, yet a human collaborator should still see the AI as a peer in the document. GRS reflects a server-side AI commit into the same roster and fan-out a human uses. The reflection is filtered to AI activity, it coalesces onto a deterministic presence id per document and actor so successive AI transactions appear as one stable peer rather than a fresh session per commit, it is injected only when the room already holds a live human session, and it is aged out by the same owner reap loop. That deterministic id is an identifier shaped like an RFC-4122 version-5 UUID, derived by hashing a fixed seed, not a namespaced UUID; the reference states exactly what it is.

There is a matching retract path in the code that would remove the synthetic peer immediately, but it has no production caller today: the live retirement path for a quiet AI peer is TTL reaping. This page does not present immediate retract as a live behavior.

Draining

When a pod is asked to hand a document off cleanly, the owner checkpoints the current head, broadcasts an owner-moving control frame so clients expect the move, stops accepting local writes, and releases every lease so a surviving pod can take over immediately with a higher epoch. Draining refuses new attaches rather than creating an ownerless room.

What is not built

Several boundaries are deliberate, and reading them as gaps rather than bugs matters:

How this connects to the rest of the engine

GRS is the authoring-session counterpart of the playable-room runtime. The playable side, its clock, authority lease, replication, and durability, is the Netcode Model; the two share the pattern of a single-writer lease whose safety is a durable fence rather than the lease itself. For the concrete GRS surfaces, invariants, numbers, and proof surfaces, read the Graph Room Sequencer reference.

Status: draft explanation page for engine v1. The GRS collaboration-session model is live (registered on the request path in createApp); its documented behavior tracks the conformance tests named on the reference page, and the items listed under "What is not built" are deliberate boundaries, not defects.

Was this helpful?Report an issueContact support

On this page