Overview
The dispatcher is a thin, stateless controller that turns the
backend’s run queue into Kubernetes work. It
holds no durable state of its own — the backend’s job table is the source of truth
— and it does only one thing: claim a queued run and create one
driver Job to execute it.
It replaces the old long-lived worker pool. Where a worker was a registered, hand-scaled HTTP server that a console addressed directly, the dispatcher sits entirely behind the backend: a console enqueues a run at the backend and never talks to the dispatcher at all. Concurrency scales with the cluster (queue admission plus available capacity), not with a manually-sized set of workers, and there is no per-pod registration to manage.
What it does
Section titled “What it does”The dispatcher runs a single control loop forever:
- Claim the next claimable job from the backend (
POST /jobs/next), authenticating with a shared service token. The claim is atomic, so the backend hands each job to exactly one dispatcher. The backend — not the dispatcher — enforces each harness’s maximum parallelism here: it only hands back a job whose harness has fewer than its configured limit of runs already in flight, holding the rest in thependingstate until a slot frees (see Harnesses → per-harness configuration). The backend also holds back a game-jam job while another run of the same jam and model is in flight under any harness, so a model’s jam entries run one at a time and each is briefed with the previous one’s README (see Game jam → repeated runs). - Create one driver
Jobfor the claimed run through the Kubernetes API, with exactly the environment the driver reads: the backend URL, the job id and its per-job token, the serialized launch request,TCAB_DRIVER_RUNTIME=kubernetes, theTCAB_K8S_*sandbox-pod passthroughs, and theTCAB_CONTAINER_*run-image selection the driver resolves the sandbox image from (so a deployment pins the run images by:<git-sha>here, not via a Kubernetesimage:field) — plus the driver pod’s own IP from the downward API, so the driver can route a sandbox’s live-preview frames back to itself. - Watch the
Jobs it created, holding at mostTCAB_DISPATCHER_MAX_INFLIGHTin flight across all harnesses (this global cap composes with the backend’s per-harness limit from step 1), and report any driver-pod death the driver itself could not (POST /jobs/{id}/status), reading the dead pod’s logs for the failure detail. - Reap the sandbox pods a dead driver orphaned — see Sandbox reaping.
- Let each finished
Jobreap itself (ttlSecondsAfterFinished).
The dispatcher never executes a run, resolves a definition, or touches a record: all of that is the driver’s job. It is purely the bridge between the backend’s queue and the cluster’s scheduler.
Relationship to the others
Section titled “Relationship to the others”- The backend owns the queue; the dispatcher owns all
Jobcreation. This keeps the backend portable (HTTP + a database, no cluster dependency) and isolates the cluster RBAC in one small component. - The driver does the work. The dispatcher’s whole product is a
driver
Jobper run; see that page for how a run actually executes. - One service token, per-job tokens minted by the backend. The dispatcher
authenticates its claim with a shared service token (
TCAB_BACKEND_SERVICE_TOKEN, which the backend also holds); each driver authenticates its own streaming with the per-job token the backend minted at enqueue and the dispatcher passed in.
Sandbox reaping
Section titled “Sandbox reaping”The driver normally deletes its own sandbox pod, but
that cleanup is in-process: a driver killed by SIGKILL (OOM kill, eviction, node
drain, spot preemption) never runs it, and the orphaned sandbox has no
ownerReference to garbage-collect it — so it runs until something deletes it,
holding its requests against the node the entire time. Left alone this compounds:
the leaked requests crowd the node, which makes the next driver more likely to be
killed, which leaks another sandbox.
The dispatcher is the only component positioned to clean this up — it is long-lived
and already watches every driver Job it created — so when one fails terminally it
deletes that job’s sandbox pods, selecting on both the job-id label and the
driver’s managed-by label. Both are required: the driver Job’s own pod carries
the same job id, and matching it would destroy the logs the failure report reads.
The reap is deliberately independent of the death report. Reporting needs a
retained per-job token and a non-terminal backend job, neither of which is
guaranteed; a sandbox must be cleaned up regardless, so it is gated only on the
Job having failed. A failed reap is retried on the next tick rather than recorded
as done.
The driver pod also carries small resource requests
(TCAB_DISPATCHER_DRIVER_{CPU,MEMORY}_REQUEST) purely to keep it out of the
BestEffort QoS class, which is what made it the first thing evicted and
OOM-killed. Limits are deliberately unset by default: a memory limit would
re-introduce the same SIGKILL from the container’s own cgroup.
The dispatcher runs under its own ServiceAccount with a namespaced Role
granting exactly: batch/jobs create/get/list/watch/delete (to create and
reconcile driver Jobs), core/pods get/list/delete (to read a dead driver pod’s
status and to reap orphaned sandbox pods), and core/pods/log get (for the
failure detail). It creates no pods directly — the
driver does that, under its own identity. A
deployment that points the driver at a different sandbox namespace
(TCAB_K8S_NAMESPACE) must grant the same pod list/delete there, or reaping
fails in that namespace (logged, never fatal). The manifests are in
deployments/k8s/base/rbac.yaml
and Kubernetes: staging & prod.
Status
Section titled “Status”The dispatcher is implemented as the test-cabinet-dispatcher crate
(crates/dispatcher), with no HTTP server and no flags — its whole configuration
is environment variables, documented on its
config.rs.
It is deployed as a single-replica Deployment (a second replica would only race
the same atomic claim — wasted work, not a correctness risk) with no Service,
since it binds no socket. Local development runs the same manifests on
k3d, so a run schedules as a Job locally exactly as it
does in the cloud.