#4320 config push
1 / 1

Reviewer walkthrough · NVIDIA/OpenShell#4320

Deliver gateway configuration over the supervisor session

Supervisors poll the gateway every 10 seconds today. This PR lets the gateway push configuration over the ConnectSupervisor stream instead. Supervisors apply each update and acknowledge it. You turn it on with config_delivery_mode = "push". Poll stays the default.

For users

Nothing changes unless an operator opts in. With push on, a policy or settings change reaches the sandbox without waiting for the next poll.

For reviewers

Start with the session protocol and the per-session delivery task. Most of the correctness argument lives there. The rest is plumbing, moved code, and generated bindings.

Getting around

Use ← → or the buttons at the bottom. Slides tagged interactive have controls.

    High level

    How a config change reaches a running sandbox

    Same boxes, two triggers. In poll mode the supervisor asks on a timer, and the gateway builds a snapshot per request. In push mode the write itself starts the build, and the gateway sends the result down the session stream. Both modes apply through the same supervisor code.

    Before poll, default

    After push, opt-in

    Highlighted boxes differ between the two modes. The provider environment follows the same path. In poll mode the supervisor fetches it only when provider_env_revision changes.

    High level

    Terms used in the code

    These names show up in code, metrics, and logs.

    TermMeaning
    Componentsandbox_config (policy, settings, middleware, admission) or provider_environment (env values, credentials, managed files). The gateway always sends a full snapshot of one component, never a diff.
    BootstrapBoth components, sent inside SessionAccepted. The supervisor applies it before the workload starts.
    Dirty markWhen a write commits, the gateway marks every affected session dirty. The mark names the components to rebuild and a lane. It carries no data. More writes merge into the same mark. The code calls the write a publication and the per-session state marks.
    LanePriority of the pending work. sandbox means a change to this one sandbox. fanout means a change that hits many sessions, plus the periodic reconcile.
    Build permitLimits concurrent builds. Two FIFO pools: shared, and a reserve that only the sandbox lane can use.
    Delivery taskOne tokio task per push session. It waits for a dirty mark, takes a permit, builds, and hands the snapshot to the session.
    AckConfigUpdateResult from the supervisor. Until it arrives, the session holds any newer snapshot of that component.
    ReconcileEvery 30 s each session rebuilds its sandbox configuration anyway, to repair anything a missed mark left stale.

    Mechanism · config_delivery/session.rs, supervisor_session.rs interactive

    Session protocol, from hello to live updates

    A push session never starts from polled state. The gateway builds both components, the supervisor applies them before the workload runs, and the gateway persists admission before launch. After that, updates flow one component at a time. Click a message or step through.

    Step 1

      Mechanism · config_delivery/mod.rs, task.rs

      Mark dirty, wait for a permit, build, send

      Marking a session dirty records intent only. The task reads state after it holds a permit, so each build sees every write that marked it. That is why merging marks is safe.

      Write commitspolicy, setting,provider, attach publish_*()pick scope, seq += 1notify peers (HA) mark dirtycomponents ∪= newlane = max(lane) acquire permitfanout: sharedsandbox: either pool take marks, buildone input load,timeout per component owner checkdeliver_configslot, then stream request path never blocks, O(sessions) delivery task, one per session Marks that land during a build wait for the next round. Failed components come back after a 1 s to 30 s backoff.

      Who marks what

      WriteScopeLane
      Sandbox policy or settings, provider attach or detachone sandboxsandbox
      Provider record change, credential refreshsandboxes attaching itfanout
      Provider profile changeworkspace or allfanout
      Global policy or settings, interceptor profile revisionall connectedfanout
      30 s reconcileitselffanout

      Every write marks both components. Only reconcile narrows to sandbox_config.

      Lanes and queueing

      • Reconcile uses the fanout lane, the same as a fleet-wide change.
      • A sandbox edit raises the session to the sandbox lane, even while it is already waiting. From then on it can take a reserve permit.
      • Permits are FIFO. Tokio's Semaphore is fair, so 100 waiting sessions get permits in the order they asked.
      • More fanout marks don't wake a waiting task, so it keeps its place in line.
      • A lane raise restarts the wait. The session moves to the back of the shared queue but also joins the reserve queue, which is usually short.

      task.rs · Permits::acquire

      Permit sizing

      Builds = 2 × DB pool size, at least 4. A quarter of them go to the reserve.

      StorePoolSharedReserve
      Postgres10155
      SQLite file582
      SQLite memory131

      Why the DB pool? Each build is a burst of small store queries, so the pool is the shared bottleneck. Two builds per connection keep it busy without piling waiters onto the pool's acquire timeout. Both pool sizes are hard-coded, so nothing here is configurable today. Provider builds also wait on credential drivers like Vault, and this sizing ignores that.

      Mechanism · task.rs interactive

      Watch marks merge and lanes compete

      Eight push sessions on a gateway with the smallest sizing, 3 shared permits and 1 reserve. Time runs slow so you can follow it. Try Global setting, then right away Edit policy on sb-3. sb-3 skips the fanout line on the reserve permit.

      t = 0.0 s
      Waiting for a permit, FIFO
      dirtybuildingsent, awaiting ackheld for ackackedfailed, backing off
      Write
      Conditions
      Clock
      0writes
      0merged marks
      0builds
      0sent

      Simplified model. One tick is 0.5 s. Input load takes 1 tick, a sandbox-config build 1, a provider build 2, and the supervisor acks after 2. Reconcile runs every 30 s, starting 30 to 60 s in. HA, owner checks, and the bootstrap are left out.

      Mechanism · task::run, RECONCILE_INTERVAL interactive

      Reconcile catches what dirty marks miss

      Dirty marks live in memory, and peer notifications can fail. So every 30 s each session rebuilds its sandbox configuration, plus any component the supervisor last failed to apply. The first pass lands at a random point 30 to 60 s after the session registers.

      Green is the bootstrap at connect, blue is a reconcile build. The bars count builds per 5 s across 40 sessions. Without the jitter, a gateway restart would line every session up on the same tick for good.

      Why only the sandbox configuration?

      Provider builds resolve credentials and may call Vault. Reconcile skips them unless something changed.

      1. Reconcile rebuilds sandbox_config.
      2. The supervisor sees a new provider_env_revision in it. It stages the snapshot and reports AWAITING_COMPONENT.
      3. The gateway rebuilds the other half, provider_environment.
      4. Once the pair matches, the supervisor activates both and sends the result it still owes.

      Polling supervisors follow the same rule. They fetch the provider environment only when its revision changes. If both halves are staged, neither triggers a rebuild of the other, so they can't loop. The next write or reconcile sorts them out.

      Mechanism · timeouts and retries

      Every wait has a bound

      Delivery is best effort and runs beside the session. A failed build never tears a session down. It retries, and reconcile covers whatever retries miss.

      Timeouts

      Sandbox-config build5 s
      Provider-env build20 s
      Bootstrap build, startup prepare45 s
      Waiting for an invalid image policy to be fixed5 min
      Holding for an ack1 min
      Peer notify5 s

      The shared input load counts against each component's own deadline.

      Build failure backoff

      Per session, 1, 2, 4, 8, 16, then 30 s. A new write or a session close cuts the wait short. A success resets it.

      A bootstrap build retries up to 3 times if its two halves disagree on provider revision, attachment epoch, or policy hash.

      Supervisor-side failures

      • The gateway remembers a rejected update per component, and reconcile rebuilds it until one succeeds.
      • A bad bootstrap result closes the session. The supervisor reconnects with ReconnectBackoff.
      • Failure messages carry no payloads, credentials, or provider values.
      • If middleware is unreachable at startup, the supervisor starts degraded, the same as in poll mode.
      • A failed session setup demotes the sandbox. The shutdown handoff from #4321 still applies.

      Mechanism · PeerNotifyConfigUpdate

      Several replicas: tell the owner, send no secrets

      A write can land on any replica. Only the replica that owns a sandbox's session can deliver to it. So the writing replica sends the owner a notification with no config in it, and the owner rebuilds from the shared store.

      Replica A, takes the write publish_sandbox(sb-7) no local session, notify peer Shared DBsandbox, policy, providers,owner index (replica, session_id) Replica B, owns sb-7 handle_peer_notify mark, build, deliver supervisor sb-7 PeerNotifyConfigUpdate{sandbox, session_id, components} reads state writes The notification names a scope and components. No snapshot data crosses replicas.

      Merged and bounded

      Notifications to the same target merge while one is pending or running. At most 8 run at once, each with a 5 s timeout.

      Tied to one session

      A sandbox notification carries the session_id the sender saw. If the session has moved, the receiver answers stale_owner. A build that finishes after the owner moved notifies the new owner.

      Skipped on one replica

      is_single_replica() skips the owner lookup and peer notifications entirely. Workspace, provider, and global changes go to every peer.

      Also helps poll mode

      Builds got cheaper in both modes

      Push builds when a write commits, not when a supervisor asks, so a fleet-wide change means a burst of builds. These changes cut the cost of each one, for poll-only gateways too.

      One input load per build

      The provider catalog, provider records, global settings, and latest sandbox policy load once into SandboxConfigInputs and feed both components. The bootstrap, combined live builds, and the readiness snapshot all use it.

      grpc/policy/config_inputs.rs

      Cached provider profiles

      Every build used to call the interceptor provider-profile sources. Now the gateway serves them from a snapshot it refreshes every 10 s. The interval is fixed. When the revision changes, all connected sessions get marked dirty.

      Reads fail closed before the first fetch and after 30 s of failed refreshes.

      provider_profile_sources.rs

      Provider environment builds

      They read refresh state once instead of twice and load providers concurrently. Only this component resolves credentials.

      grpc/provider.rs, provider_refresh.rs

      Rollout interactive

      Opt-in, safe across versions, easy to undo

      Push needs a gateway in push mode and a supervisor that says it supports it. Every other pairing polls exactly as today. Pick one.

      Rolling back

      Set config_delivery_mode = "poll" and restart the gateway. Supervisors reconnect, get config_push_enabled = false, and go back to polling. There is no migration. Poll mode reads no state that this PR adds.

      Mixed replicas

      An older replica doesn't implement PeerNotifyConfigUpdate, and senders count that as unsupported_peer. Its sessions poll anyway, so nothing is lost. Push sessions on newer replicas still reconcile every 30 s.

      Phase 0, this PR

      Ship it off

      Default is poll. New supervisors advertise support. CI runs both e2e:rust and e2e:rust:push.

      Phase 1

      Opt in

      Turn push on for dev and internal gateways. Watch queue wait, build failures, apply results, and bootstraps. Load test push against poll, then HA with 3 replicas.

      Phase 2

      Default to push

      Flip the default once every supported supervisor sets supports_config_push. Old supervisors keep polling.

      Phase 3

      Remove polling

      Drop GetSandboxConfig and GetSandboxProviderEnvironment once old supervisors are out of support. The proto comments already call them polling projections.

      Phases 1 to 3 are my proposal for discussion. They are not part of this PR.

      Low level

      Where to spend review time

      Generated Go bindings and moved supervisor code make up most of the line count. The new logic sits in about 6k lines across four gateway files, and about half of session.rs is tests.

      File+−Read for

      Suggested order

      1. proto/openshell.proto, the contract.
      2. config_delivery/task.rs. Marks, lanes, permits. Short and well tested.
      3. config_delivery/mod.rs. Publishing, registration, builds, owner check, peer notify.
      4. config_delivery/session.rs. Bootstrap, acks, admission.
      5. supervisor_session.rs on the server, for the handshake.
      6. supervisor-process/supervisor_session.rs, then config_runtime.rs. Use git diff --color-moved for the latter.

      Open questions

      • Is FIFO fair enough across workspaces, or do we need round-robin now? One busy workspace can fill the shared queue.
      • Should the build limit be configurable? Today it follows a hard-coded DB pool and ignores credential driver latency.
      • Is 30 s the right reconcile interval for large fleets? It's a constant.
      • Should provider readiness reports move onto the supervisor stream after this lands?
      • Durable config update operations are deferred. #1731's acceptance criteria need an amendment.