Source: https://mayphus.org/operating-small-services/ Title: Operating a small service through change Metadata: {"kind":"page","language":"en","route":"/operating-small-services/","tags":["systems","infrastructure","operations","learning"],"type":"page"} # Operating a small service through change A service needs more than a successful installation. Someone must know what it depends on, how to tell whether it works, how to recover its data, and when the old copy can be retired. Those questions connect device provisioning, virtual machines, application releases and database changes. This is an educational guide with an invented example: a small equipment catalogue, an application, a PostgreSQL database and a reverse proxy. It is not a report of a past deployment or a claim of personal responsibility for a production system. The planning exercise below has not been executed as a migration. Use documentation for the versions actually installed before turning a plan into commands. ## Start with the service, then draw its dependencies For the fictional catalogue, a browser reaches the proxy, the proxy reaches the application, and the application reads and writes the database. A scheduled worker also writes to that database. Uploaded attachments live separately. This last detail changes both the backup boundary and the migration plan: a database restore alone would not recover the attachments. Make a compact service record before changing anything: | Record | Question it should answer | | --- | --- | | Purpose and owner | What user task must work, and who accepts the result? | | Running versions | Which application revision, runtime, database major version and OS release are in use? | | State | Where are database data, attachments and queued work stored? | | Dependencies | Which names, routes, certificates, repositories and external services must be available? | | Recovery | What can be restored, from which recovery point, and who can access it? | | Change boundary | Which writers and readers move together, and what observation stops the change? | Record immutable artifact identifiers where available. A version pin helps identify an input; it does not make that input supported, secure or compatible. Keep an update and review process alongside the pin. Separate secret references from the record itself: a handoff can identify the authorized credential store without embedding credentials. ## Provisioning ends with an accepted handoff Separate installation, configuration, application delivery and acceptance. Booting an installed OS establishes one milestone. It does not establish that the application starts after reboot, receives the correct configuration, or can recover after its network dependency disappears. For the catalogue, a useful acceptance check would create a synthetic entry, retrieve it through the public-facing route, restart the test instance and retrieve it again. A separate check would exercise attachment storage and the scheduled worker. These are proposed tests, not results obtained here. Agree who owns failed checks and how temporary provisioning access is removed before calling the device ready. The [local-media Ubuntu record](/local-media-ubuntu-install/) provides a concrete distinction between a documented historical workflow and a separately tested installation. Keep its evidence and limitations on that page. The broader lesson here is to name the remaining steps at handoff instead of treating an installer finishing as acceptance of the whole service. ## A virtual machine has more state than its disk A VM plan needs its disk images, domain definition, firmware state, CPU requirements, attached devices and network/storage dependencies. The exact pieces depend on the hypervisor and guest. Libvirt's [domain format](https://libvirt.org/formatdomain.html) documents these configuration surfaces. Distinguish moving a guest, copying its definition and restoring it from backup. Libvirt's [migration documentation](https://libvirt.org/migration.html) describes several transport and coordination modes; its offline migration transfers a definition and does not copy non-shared storage. A completed definition transfer therefore cannot prove that the destination has a bootable disk. Check storage transfer, destination compatibility and persistent configuration explicitly. A live migration also needs a plan for failure between the source and destination; it is not a backup strategy. For the fictional catalogue, write down which instance is allowed to accept writes at each stage. Keeping an old VM available for recovery must not accidentally leave two independent copies serving new writes. ## Prove recovery before relying on a backup Choose a recovery point objective—the acceptable amount of lost recent work—and a recovery time objective—the acceptable restoration delay. These are requirements to agree on and test, not numbers to invent after a successful backup job. PostgreSQL distinguishes [SQL dumps, filesystem backups and continuous archiving](https://www.postgresql.org/docs/18/backup.html). Select a method that fits the required recovery point and installed version. An ordinary copy of files from a running database is not automatically consistent. PostgreSQL's [filesystem-backup guidance](https://www.postgresql.org/docs/18/backup-file.html) explains shutdown and consistent-snapshot requirements, including coordination across volumes and the need for WAL data. In an isolated restore exercise, record the backup identifier, restore steps, elapsed time and application checks. For the catalogue, check relationships between entries and attachments, not just whether the database process starts. Keep the recovery copy outside the failure boundary it is meant to survive. Document retention, access and the last successful restore test; a green upload status answers a smaller question. ## Coordinate application and database changes A database major upgrade and an application schema change are different operations. PostgreSQL's [upgrade documentation](https://www.postgresql.org/docs/18/upgrading.html) distinguishes major-version migration paths and calls for testing application compatibility. Do not assume a new server can simply open the old data directory. For the fictional service, plan a maintenance-window migration in this order: 1. Rehearse restoration and compatibility checks in an isolated destination using synthetic data. List what the rehearsal cannot establish about real data volume or extensions. 2. Define the write-stop boundary, including background workers and queued jobs. Record how to confirm it took effect. 3. Capture the final recoverable state using the selected database method and coordinated attachment copy. Verify the destination before routing users to it. 4. Test through the intended route, then deliberately enable the destination's writers. Observe application errors and data behavior against the acceptance criteria. 5. Retain the old state for the agreed recovery window, with its writers disabled, before retirement. Rollback needs a decision point. Before destination writes begin, switching back may be straightforward if the source remained intact. After new writes or an incompatible schema change, routing back can lose or split data. The plan must specify whether recovery requires reconciliation, restoration or a forward fix. “Keep the old VM” does not answer that question. ## Observe the route that users actually take An application can pass a direct health check while a proxy sends users to the wrong upstream. A cache can also serve an older response after a successful application release. Define which responses may be cached, how their keys vary, how long they remain usable, and what invalidation means for the release. Treat personalized responses separately; do not assume a shared cache is safe for them. NGINX's [proxy module reference](https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_cache) documents cache keys, bypass and no-cache controls, and response-header handling. For the catalogue, inspect an anonymous listing and an authenticated view through the proxy with synthetic accounts. Compare expected content and cache behavior, not only status codes. Keep administrative checks distinct from the user path. Google's [monitoring chapter](https://sre.google/sre-book/monitoring-distributed-systems/) describes latency, traffic, errors and saturation as useful signals. Pair those with an end-to-end check of the catalogue's core task. Logs help explain individual failures; metrics help show patterns. Choose alerts with an owner and an action. Avoid logging credentials or unnecessary user data, and define retention so debugging does not become an unbounded archive. ## Make the next handoff easier A change record should name the intended behavior, exact deployed revision, checks performed, observed result and recovery decision. Keep proposed, attempted and verified steps visibly distinct. Review access when ownership changes; service accounts, recovery access and temporary migration permissions need owners too. Retirement is a final change: confirm dependencies have moved, stop old writers, remove obsolete routing and scheduled work, resolve retention requirements, then remove resources and access through the appropriate process. Update the service record so the next operator does not rediscover a retired system as an apparent dependency. For a paper exercise, add one failure to the fictional catalogue plan: the destination accepts new records, but attachment retrieval fails. State what you would stop, what evidence you would preserve, and why redirecting traffic back is no longer a complete rollback. No command execution is needed; the exercise tests whether the plan accounts for state and ownership. Continue with [Understanding computer systems](/understanding-computer-systems/) for the underlying files, processes, storage and networking concepts. The [site infrastructure record](/infrastructure/) shows a confirmed public example of resolving source versions and promoting the same checked release, without implying that this fictional VM migration has been performed.