Operate over time
A passthrough instance is designed to run unattended for months. This page is what that costs you: what to watch, what to back up, what grows, and how to upgrade. Some of the tooling that will make parts of this easier is still ahead — those parts are badged below rather than hidden.
Probes and metrics
Section titled “Probes and metrics”| Path | Role |
|---|---|
/healthz |
Liveness. Answers as soon as the listener is up: the process is alive, not necessarily useful. |
/readyz |
Readiness. 503 until the store and the configuration are usable, and 503 again during the shutdown drain. |
/metrics |
OpenMetrics (FR-091). Behind the same authentication as every other surface — give your scraper a viewer account or token; the chart’s ServiceMonitor supports basic auth. |
Opening a large store takes time: the reference deployment uses a startup probe (30 × 5s) rather than a lax liveness threshold, so the instance gets time to come up without a hung process surviving minutes afterwards.
Tasks are the unit of observation
Section titled “Tasks are the unit of observation”Every synchronization and every one-off import runs as a tracked task
(FR-062): state, per-ingredient progress, and a log downloadable raw.
The task list is /tasks in the UI and GET /api/v1/tasks in the API
(both paginated); a task’s detail and log are
GET /api/v1/tasks/{id} and GET /api/v1/tasks/{id}/logs. Logs are
structured JSON with stable correlation keys — run ID, task ID, recipe,
ingredient, digest (FR-090) — so one synchronization is fully
reconstructable by filtering on its task ID. See
metrics and logs for the schema.
Finished tasks are bounded: the queue keeps the most recent
tasks.keepFinished (default 500; 0 keeps everything), purging older
entries together with their log files. Pending and running tasks are
never purged.
Interrupted transfers resume
Section titled “Interrupted transfers resume”A killed process or a cut connection does not restart a run from
scratch. A synchronization resumes from persisted state, and blobs
already stored are never re-downloaded (FR-029). Since v0.4.0, resume is
fine-grained inside large blobs too (R-29): above
transfer.resumeThreshold (default 64MiB) the bytes are spooled in the
state directory with their offset, and the next attempt asks for the
rest with an HTTP Range request — a download interrupted at 90 %
restarts at 90 %, surviving a killed process, not just a dropped
connection. Integrity stays blocking: the digest is computed over the
whole spool, resumed prefix included, before a byte reaches the store,
and a source that ignores Range is detected and restarted rather than
concatenated. The task detail shows per-blob progress, including whether
a transfer resumed.
The operational consequence: the state volume temporarily holds one copy
of each resumable blob in flight — the reference deployment sizes it at
20Gi for this reason. transfer.resumeThreshold: 0 disables the spool
and restores pure streaming.
What to alert on
Section titled “What to alert on”/readyznon-200 outside deploy windows.- Failed tasks — poll
GET /api/v1/tasksor alert on the task-failure metrics. - Policy refusals: allowlist rejections (FR-030) and signature verification failures (FR-033) are logged, audited, and counted in metrics. In a healthy zone these are zero; any step is either an attack or a configuration drift, and both deserve a page.
- A last-successful-sync age beyond a few multiples of
sync.interval— a proxy or credential silently broken shows up here first. - State-volume usage, since resumable spools live there.
The metric families are listed in metrics and logs — with the honest caveat that their names are not yet contractual.
Backup: the state directory
Section titled “Backup: the state directory”The state root is the backup target. It holds what nothing
recreates: accounts, tokens, the served TLS pair, the interval override.
It is small, so back it up like any precious directory — snapshots or
file copies of state.root, taken while the instance is stopped or from
a filesystem snapshot. The store needs no backup: everything in it can
be fetched again, and losing it costs bandwidth, not identity. Never
put the state on the store volume; Tobby refuses the nesting outright.
Store growth, stated plainly
Section titled “Store growth, stated plainly”By default, nothing cleans the store automatically in passthrough mode. Every synchronized version and every one-off import stays until an administrator removes unit-imported repositories by hand (FR-044) — recipe-managed content is not individually removable at all. A zone whose recipes track moving constraints accumulates every version they ever resolved.
That default is deliberate. A passthrough transit store is not a delivery unit, and an operator asking for fresher content has not asked for older content to be deleted: a reconciliation loop that quietly shrinks the store a zone pulls from is exactly the failure this default prevents.
Pruning to the Retriever
Section titled “Pruning to the Retriever”Set sync.prune: true (or TOBBY_SYNC_PRUNE=true) and every
reconciliation cycle removes the recipe-managed content the resolved
Retriever no longer references. Three kinds of content are never
eligible, because none of them is recipe-managed:
- unit imports (FR-023),
- the offline vulnerability database (FR-032),
- anything pushed through
/v2/outside managed namespaces (UC3 seeding).
Two safeguards are worth knowing. A cycle in which any recipe failed to resolve prunes nothing and says so in the run log: content of a recipe that did not resolve is indistinguishable from content the Retriever dropped, and deleting on the strength of a network failure is not a trade this product makes. And every removed item is named — repository, tag, digest, and the recipe that brought it — in the run log of the cycle, not merely counted.
Watching the volume
Section titled “Watching the volume”Set storage.occupancyThreshold (for example 500GiB) and the instance
says when it is over: a persistent warning on every UI page, the same
fact on GET /api/v1/content and GET /api/v1/retriever, and the
tobby_store_occupancy_exceeded metric. Crossing back under retracts all
three — a warning that appears and never clears is a warning operators
learn to ignore. Unset means unmonitored, which is reported as such and
never as “within limits”.
Seeing it coming
Section titled “Seeing it coming”tobby sync --dry-run reports what the next synchronization would do
without doing any of it: resolved versions, per-digest statuses, the
deduplicated volume to transfer, the projected store size against the
volume’s free space, and the content a prune would remove. It writes
nothing and does not touch the reconciliation schedule. The same report is
on POST /api/v1/plan and on the /recipes/plan screen, where a candidate
Retriever can be planned instead of the configured one — which is how you
review a Retriever change before adopting it.
Exit code 5 means “changes planned”, distinct from 0 (“nothing to do”),
so a pipeline can gate on it without treating a plan with work in it as a
broken build. See the CLI reference.
Upgrading
Section titled “Upgrading”Read the release notes first; the compatibility policy — what is stable already, what freezes at 1.0, and how store formats are carried across versions — lives in release process and compatibility.
- Packages: verify the new package, then install it over the old one
(
dpkg -i/rpm -U/apk add --allow-untrusted) and restart the service. Packages are scriptless; nothing runs on install. - Kubernetes:
helm upgrade tobby ./deploy/charts/tobby --namespace tobby --reuse-values --set image.tag=v0.4.2— pinimage.digestin production. The strategy isRecreatewith a single replica: the old pod releases the volumes before the new one starts, so there are never two writers on one store. Expect a short outage per upgrade; both PVCs carryhelm.sh/resource-policy: keep, so evenhelm uninstallleaves the data behind.
Graceful shutdown
Section titled “Graceful shutdown”On SIGTERM or SIGINT, the instance stops accepting new work, turns
/readyz to 503, and gives in-flight transfers
shutdown.gracePeriod (default 30s, --shutdown-grace-period) to
finish or checkpoint before exiting 0 (FR-093). Checkpointed transfers
resume on the next start. Whatever supervises the process must wait
longer than the grace period — terminationGracePeriodSeconds: 60 in
the reference deployment, TimeoutStopSec=60 in a systemd unit —
or the final kill lands mid-checkpoint.
That closes the passthrough journey. From here: write recipes to grow what the zone holds, or read how the same instance prepares for isolated zones.