devops-architecture/

What doing DevOps taught me about architecture

9 min read codestinger

I'm a software architect by profession, but I rarely stop at the diagram. I develop and manage teams. When a product needs it, I do the DevOps myself: pipelines, infrastructure and releases. For a while design and operations felt like two different jobs: one drew the boxes, the other kept the boxes running. Doing both changed how I design almost everything.

These are the lessons that changed my designs the most, with the concrete practices behind each one.

1. The deployment is part of the design#

If nobody has decided how a component is built, versioned, deployed and rolled back, the design isn't finished. I ask "how does this ship? How does it un-ship?" in the first architecture discussion, not the last.

Two practices follow from that question.

Build once, promote everywhere. The artefact that passed testing must be the exact artefact that reaches production. Build a wheel or image once, give it an immutable version and promote that same artefact from test to staging to production. Environment differences belong in configuration, never in rebuilt code.

Make every change reversible. Rolling back code is easy; rolling back a database is not. So schema changes follow the expand and contract pattern, where each step is backward compatible with the version before it:

-- Release 1 (expand): add the new column; old code simply ignores it
ALTER TABLE orders ADD COLUMN lot_number TEXT;

-- Release 1 (migrate): backfill from the old column, in batches on large tables
UPDATE orders SET lot_number = batch_code WHERE lot_number IS NULL;

-- Release 2: code reads and writes lot_number only

-- Release 3 (contract): once no deployed version uses the old column
ALTER TABLE orders DROP COLUMN batch_code;

At every point, the previous release still works against the current schema, which means rollback is always one deployment away.

2. Configuration is an interface, so validate it#

Most production incidents I've been pulled into were not code bugs. They were configuration: a missing variable, a typo in a URL, a timeout in milliseconds where seconds were expected. Treat configuration as a typed interface and validate all of it at start-up, reporting every problem at once:

import os
from dataclasses import dataclass, fields


class ConfigError(Exception):
    pass


@dataclass(frozen=True)
class Settings:
    erp_url: str
    batch_size: int
    request_timeout: float
    dry_run: bool

    @classmethod
    def from_env(cls, prefix="APP_", env=os.environ):
        values, problems = {}, []
        for field in fields(cls):
            name = prefix + field.name.upper()
            raw = env.get(name)
            if raw is None:
                problems.append(f"{name} is not set")
                continue
            try:
                if field.type is bool:
                    if raw.lower() not in {"true", "false"}:
                        raise ValueError("expected true or false")
                    values[field.name] = raw.lower() == "true"
                else:
                    values[field.name] = field.type(raw)
            except ValueError as exc:
                problems.append(f"{name}={raw!r} is invalid: {exc}")
        if problems:
            raise ConfigError("invalid configuration:\n  " + "\n  ".join(problems))
        return cls(**values)

A service that refuses to start with a clear list of what's wrong is far cheaper than one that starts and fails an hour later on the first real request.

3. Operability is a feature, so give it numbers#

"The system should be reliable" is not a requirement. A service level objective is. Agree on what "working" means (for example, 99.9% of order messages processed within five minutes) and the target turns into an error budget: the amount of failure you can afford before reliability work takes priority over features.

def error_budget_minutes(slo, days=30):
    return (1 - slo) * days * 24 * 60


def burn_rate(bad_minutes, elapsed_days, slo, days=30):
    """1.0 means the budget will run out exactly at the end of the period."""
    budget = error_budget_minutes(slo, days)
    return (bad_minutes / budget) * (days / elapsed_days)


error_budget_minutes(0.999)  # 43.2 minutes per 30 days
burn_rate(bad_minutes=20, elapsed_days=5, slo=0.999)  # ≈ 2.8: alert, the budget is burning too fast

Alert on the burn rate rather than on every individual error. The on-call rotation stops drowning in noise. Pair it with health checks that mean something: a liveness check says the process is running; a readiness check says it can actually reach its database, queue and downstream systems. Orchestrators need both. They should never be the same endpoint.

4. Boring technology wins, by explicit criteria#

Every new tool has a cost that isn't on the pricing page: learning, upgrading, securing and debugging it at night. I'd rather run a few well-understood things superbly than many interesting things badly. To keep that from becoming "we never adopt anything new", I ask the same questions of every proposal:

  • Which problem does it solve that our current stack cannot? How do we measure that?
  • Who on the team can operate it at 3 a.m.? Who is the second person?
  • How is it upgraded, backed up, monitored and secured?
  • What is the exit plan if it's abandoned or relicensed?

Technology that answers all four earns its place. Novelty alone doesn't.

5. The pipeline is the product's immune system#

Security and quality checks that depend on someone remembering a checklist will eventually be skipped. In the pipeline they run on every change. Nothing reaches production without passing them. A condensed version of the shape I use in GitLab CI:

stages: [test, security, build, deploy]

variables:
  UV_FROZEN: "1"  # never resolve new versions in CI: the lockfile is the truth

test:
  stage: test
  image: python:3.12
  script:
    - pip install uv
    - uv sync
    - uv run pytest --junitxml=report.xml
  artifacts:
    reports:
      junit: report.xml

audit:
  stage: security
  image: python:3.12
  script:
    - pip install uv
    - uv export --no-hashes --no-emit-project --format requirements-txt -o requirements.txt
    - uvx pip-audit -r requirements.txt
    - uvx bandit -r src -ll
    - uvx --from cyclonedx-bom cyclonedx-py requirements requirements.txt -o sbom.json
  artifacts:
    paths: [sbom.json]

build:
  stage: build
  image: python:3.12
  script:
    - pip install uv
    - uv build
  artifacts:
    paths: [dist/]

deploy-production:
  stage: deploy
  id_tokens:
    CLOUD_ID_TOKEN:
      aud: https://cloud.example.com
  script:
    - ./deploy.sh dist/  # exchanges CLOUD_ID_TOKEN for short-lived cloud credentials
  environment: production
  rules:
    - if: $CI_COMMIT_TAG
      when: manual

A few details carry most of the value:

  • The lockfile is enforced, so every build is reproducible and a surprise dependency upgrade can't sneak in.
  • Dependencies are audited and code is scanned on every change. A software bill of materials (SBOM) records exactly what shipped, which turns "are we affected by this new vulnerability?" into a search instead of a guess.
  • The artefact is built once. The deploy job uses that same dist/.
  • No long-lived cloud secrets exist. id_tokens gives the job a signed OIDC token that the cloud provider trades for short-lived credentials, scoped to this project and this environment.
  • Production needs a tag and a human, so every release is deliberate and traceable.
  • People get the least access they need. Routine changes go through the pipeline. Direct access to production is temporary, approved and logged. Power over production is a responsibility and the process should make that visible.

6. Own your supply chain#

Every dependency is code you run but didn't write. Beyond the pipeline checks, a few habits make that manageable: pin everything through lockfiles, host approved packages on an internal index so builds don't depend on the public internet being available and friendly, build native wheels once per platform in CI rather than on customer machines and sign release artefacts so a deployment can prove where it came from.

7. Design for the team you have, then pave the road#

The best architecture is one your team can actually own. A design that quietly depends on a specialist nobody has hired is a liability, however clever it is. Match the complexity to the people and grow both together. When a system genuinely needs something specialised, someone has to own it openly. Often I build that part myself and deliver it as a well-tested building block with a simple interface, so the team can rely on it without carrying its complexity.

The most effective way I've found to do that is a golden path: a project template that already contains the pipeline above, logging, configuration validation, health checks and documentation skeletons. Starting a new service the right way becomes the easiest option. Good practice spreads by default instead of by memo.

8. Write decisions down#

Most painful conversations in a codebase's life start with "why on earth did we do it this way?" A short written record answers that question for years. On my teams, every significant change gets a one-page Architecture Decision Record:

# ADR-012: Use a message queue between the order connector and the ERP

## Status
Accepted on 2026-03-18

## Context
The ERP accepts about 20 requests/second and goes down for maintenance weekly.
Order bursts reach several hundred per minute.

## Decision
Publish orders to a durable queue; a single consumer writes to the ERP with
retries and idempotency keys.

## Consequences
+ Bursts and ERP downtime no longer lose or delay orders at the source.
+ The consumer can be paused, replayed and scaled independently.
- One more component to deploy and monitor.
- Orders reach the ERP eventually, not instantly.

It takes ten minutes to write and saves hours of archaeology. The Consequences section is the most important part: every decision has downsides and naming them honestly is what makes the record trustworthy.

The short version#

Architecture and operations aren't separate disciplines. Build once and promote the same artefact. Make every change reversible. Validate configuration like an interface. Give reliability numbers and alert on budgets, not noise. Adopt technology by criteria, not fashion. Let the pipeline guard quality and the supply chain. Pave a golden path for the team. Write decisions down. Your future self (probably on call) will thank you.