---
title: "Explanation: Overload and Degradation Model"
description: "How Gessa's server-authoritative runtime measures a room, connection, stream, and shard for pressure and sheds work in a fixed order before it fails, and which parts of that ladder are wired and which are declared."
engineVersion: v1.0.234
date: 2026-09-28
license: "(c) Gessa, proprietary. Cite with attribution to https://gessa.ai/docs/. Terms: https://gessa.ai/terms/."
canonical: https://gessa.ai/docs/explanation/backpressure-degradation/
---

# Explanation: Overload and Degradation Model

{% version engineVersion="v1.0.232" coordinate="docs-public.v0" /%}

_Last verified 2026-09-05 against engine v1.0.232._

This page awaits re-extraction under decision D26 of the docs program; treat every number on it as provisional until its dossier receipt is stamped.

A multiplayer room can be asked to do more than its process can afford: a burst of state, a slow client that will not drain, more rooms than one instance should host. Gessa's runtime answers that with **graceful degradation**, a fixed order in which it sheds cheap-to-lose work first and refuses or stops last, rather than a single cliff where everything fails at once. This page explains the model: what is measured, what is shed and in what order, and (just as important) which rungs are wired into a live actuator today and which are declared but inert. It frames the **why**. Every constant, threshold, and capacity number lives in the [Backpressure and Degradation reference](../reference/backpressure-degradation.md), which this page links down to and never restates.

The posture is the one from [Backend Authority](backend-authority.md) and the [Netcode Model](netcode-model.md): the server owns the room, and every shed decision is the server's. A client is told what happened through a typed code; it never negotiates its own service level.

## Pressure is measured, not guessed

Per room, `summarizeRuntimeBackpressure` classifies replication health from two independent signals. The first is **acknowledgement lag**: for each connected-or-degraded connection, how far the server's latest sequence has run ahead of what the client last acknowledged.

```math
ackLag(c) = \max(0,\ latestServerSeq - c.lastAckedServerSeq)
```

The summary reads the maximum and the p95 of that lag across the room and trips the highest of three named thresholds, or escalates from the count of connections already flagged degraded. The second signal is **gateway pressure**: when the routes layer rejects an oversized envelope or a command over budget, it records that rejection on the connection, and the summary reconstructs it. The two are combined by rank, worst wins:

```math
state = \max\nolimits_{rank}\big(\ ackLagState,\ \ degradedCountState,\ \ gatewayPressureState\ \big),\quad none < watch < throttling < shedding
```

This classification is **observational**. It labels a room so operators and clients can see the strain; it does not itself shed anything. The shedding is done by the mechanisms below, and the label is what a client reads to know a room is under load.

## Two ladders, not one cliff

Overload is handled by two ordered ladders. The **per-room** ladder, `RUNTIME_LOAD_DEGRADATION_ORDER`, is a fixed sequence: downsample or drop state streams, then drop low-priority replication, then reject room commands, then evict slow consumers, then mark the room unhealthy. It is a **declaration**, not a dispatcher: the array performs no action of its own, and each rung is realized by a separate module (the stream governor, the slow-consumer machine, the routes command-shed, the tick fuse) described below.

The **per-shard** ladder, `RUNTIME_SHARD_LOAD_DEGRADATION_ORDER`, is meant to compose above the per-room one when rooms run on worker-thread shards: thin snapshots, reduce snapshot rate, pause non-essential lanes, refuse joins, migrate or close. Its selector `evaluateShardDegradation` is a pure function of a shard's signals that returns the single highest active rung, walked from calm upward:

```math
tickRatio = \frac{tickP99Ms}{\max(1,\ frameBudgetMs)}
```

```math
rung = \begin{cases}
5 & tickRatio \ge 2 \ \lor\ ring \ge 1\\
4 & tickRatio \ge 1.5 \ \lor\ ring \ge 0.95 \ \lor\ roomCount \ge roomCapacity\\
3 & tickRatio \ge 1.25 \ \lor\ ring \ge 0.85\\
2 & tickRatio \ge 1 \ \lor\ ring \ge 0.7\\
1 & tickRatio \ge 0.8 \ \lor\ ring \ge 0.5\\
0 & \text{otherwise}
\end{cases}
```

The two top rungs set `refusesJoins`, and the thresholds are budget-relative so the ladder holds at any tick rate. The values in that piecewise definition are the literals inside `evaluateShardDegradation` itself, which is their source of truth; the ladder's conformance test proves the rung order and the typed code each rung emits, not that any interior threshold is drift-protected.

{% warning severity="caution" title="The per-shard ladder is not wired into a live actuator" %}
`evaluateShardDegradation` has no caller anywhere in the server. The shard host gathers a shard's load signals through `requestGovernor`, but nothing evaluates them into a live shed action, and the production tick-authority path is unsharded (see the capacity floor below). Treat the shard rungs (thin snapshots, reduce snapshot rate, pause non-essential lanes) as declared-but-inert today. The typed codes exist, and the logic is unit-tested, but the wiring that would make a running shard shed on its own signals was not found.
{% /warning %}

## The state-stream governor

Author-driven state streams are the loudest ingress source, so they pass their own governor, `RuntimeStateStreamBudget`, before they are admitted. A sample runs a chain of sliding one-second windows: per-stream rate, then per-stream bytes, then per-player aggregate bytes, then per-room aggregate bytes. The **first** scope it would exceed rejects it with a typed reason and a retry hint, and the hint is specific: a rate rejection tells the client to drop its next sample, a byte or aggregate rejection tells it to downsample. Crucially, an accepted sample whose utilization has reached the degrade ratio still carries a downsample hint, so a well-behaved client backs off **before** it starts losing samples. Ahead of all of that, a per-player distinct-stream cap (`enforcePlayerStreamStateLimit`) bounds how many streams one player may open, so the byte and rate windows are never asked to police an unbounded fan-out. This realizes the first per-room rung.

## The slow-consumer machine

Replication is unreliable by design, but a client that stops draining its socket still costs the server memory. Per outbound socket, `slowConsumerStage` maps the socket's buffered bytes through a fixed ladder to a watch, throttle, or evict stage, and the stage is latched so it can only rise. Under **throttle**, only a low-priority (replication) frame is dropped, with reason `backpressure_throttle`; higher-priority frames still go out. Under **evict**, the subscriber is detached with WebSocket close code 1008. These two stages are exactly the realizations of the per-room ladder's drop-low-priority-replication and evict-slow-consumers rungs.

## The capacity floor

Before a room ever joins the tick loop, `evaluateRoomPlacement` checks the placement against a budget and returns admit or a typed refusal naming the first breached scope, and `assertRoomPlacementWithinBudget` throws that refusal (`runtime_shard_over_budget`, category capacity) so an over-budget room never starts ticking. Instance scopes are always checked; shard scopes are checked only when the placement is sharded. The **live** tick-authority path places with the unsharded flag, so in production only the per-instance scopes are ever consulted, and the per-shard budget and the shard host that would use it run only under the test harness.

The capacity numbers themselves are honest about their status. The placement caps are marked provisional in the source, conservative placeholders re-measured per release rather than tuned service levels, so at their generous defaults admission does not refuse in normal production. Separately, a durability-bounded per-loop knee records the room and player counts a single loop was validated to sustain; it is a **load-test guardrail** (the load harness throws when a plan exceeds it), not a production admission gate, and its room knee sits below the enforced per-instance cap.

## What the runtime promises here, and what it does not

Read plainly, the model as implemented is: a room's strain is measured and classified; state streams and slow consumers are governed live and shed in the declared order; room placement is refused live before an over-budget room can start; and one capacity tier is benchmark-validated. What is **not** yet wired: the per-shard ladder does not actuate (no actuator, unsharded live path), the placement caps are provisional rather than measured, and the classification summary drives no shed on its own path. The reference marks each of these as a limit rather than hiding it.

## Where the values and proofs live

This page states the shape of the model. For every constant and threshold with the test or check that pins it, the typed error codes, and the proof surfaces, read the [Backpressure and Degradation reference](../reference/backpressure-degradation.md). For the wider runtime this sits inside (the room clock, authority, replication lanes, durability, and rejoin), read the [Netcode Model](netcode-model.md) and the [Netcode Contract reference](../reference/netcode-contract.md). The player-safe copy for each overload code is in the generated runtime error registry.

{% generated-reference file="docs/spec/generated/error-catalog.md" label="Runtime error registry" /%}
