Everyways · how-to

How Cloud Native Applications Run

"Cloud native" is an unhelpful name for a genuinely useful idea: build the software so that the machines underneath it are disposable, describe what should be running instead of doing the running yourself, and let a control loop keep reality matching the description. Here is what that actually means, one piece at a time.

The port behind this page is the system, and it follows along as you read. Drag the margins to move it — or use Hide the article, up in the corner, to push the words out of the way and have the whole harbour to yourself.

01What "cloud native" actually means

The phrase is marketing residue, and it has been attached to almost everything. Strip it back and there is one real idea underneath: stop treating the machines as the thing you are managing. Individual servers become interchangeable and short-lived, the application is packaged so it can start anywhere, and a piece of software — not a person with a runbook — is responsible for keeping the right number of copies alive in the right places.

Everything else follows from that. Containers exist because you need a package that runs identically on any machine. Registries exist because that package has to be somewhere. Declarative manifests exist because if a person has to type the deployment commands, the machines were not really disposable. Orchestrators exist to run the loop. Service discovery exists because addresses stop being stable once instances are.

The one idea to keep

You describe what should be true, and something else is permanently responsible for the difference between that and what is true. Every component on the map behind this page is either part of writing that description, part of closing that gap, or part of noticing that the gap exists.

The old joke — cattle, not pets — is crude but exactly right. A pet server has a name, an uptime you are proud of, and a configuration that exists nowhere but on the server. A herd of cattle has a count. When one is sick you replace it, and if replacing it is dramatic then you have not actually got a herd.

None of this makes your software good. It makes your software operable — which is a different and much narrower promise than the conference talks imply.

02The application: the part that has to cooperate

A platform cannot make a badly behaved program disposable. If your application writes user uploads to its own local disk, an orchestrator that moves it to another machine is not a feature — it is data loss with extra steps. The properties have to be in the application first.

The canonical list is the twelve-factor app, written for Heroku in 2011 and still the best short description of what a container-shaped program looks like. The parts that matter most in practice:

Stateless processes

Anything worth keeping goes to a database, a queue or object storage. Local disk is scratch space that may vanish mid-request.

Config from the environment

The same image runs in dev, staging and production. Everything that differs arrives as environment variables or mounted files.

Logs to stdout

The process writes an event stream and stops caring. Collecting, shipping and retaining it is the platform's job, not the app's.

Disposability

Start in seconds, and handle SIGTERM by finishing in-flight work and exiting. You will be killed routinely, not exceptionally.

Two more are worth calling out because they are usually where the pain lands. Backing services are attached resources: the database is reached by a URL from config, not compiled in, so it can be swapped for a local one or a managed one without changing code. And dev/prod parity: the further your laptop is from production, the more of your bugs are only discoverable by your users.

Graceful shutdown is the one everyone skips

When a pod is removed, it receives SIGTERM and then, after a grace period, SIGKILL. An application that ignores the first one will drop every request it was holding, every time it is rescheduled — which, on a healthy cluster, is often. Most mysterious deployment-time 502s are this.

03Containers: the unit that travels

A container is not a small virtual machine. There is no second kernel and no emulated hardware. It is an ordinary process on the host, started with a restricted view of the world: its own filesystem root, its own process table, its own network interfaces (namespaces), and a cap on how much CPU and memory it may use (cgroups). Both features are Linux kernel plumbing, and both predate the word "container" being interesting.

An image is what you build and ship: a stack of read-only filesystem layers plus a small JSON document saying which command to run, which user to run it as, and which ports it expects. Layers are content-addressed, so two images that share a base share the bytes — which is why pulling the fifth version of your service is fast and pulling the first one is not.

ThingWhat it isCommon confusion
ImageThe immutable package, sitting in a registryNot running. An image is a noun.
ContainerOne running instance of an imageNot a VM. Shares the host kernel.
LayerOne filesystem diff in the stackDeleting a file in a later layer does not remove it from the image.
DockerfileA recipe for producing layersOne of several ways in; buildpacks and Nix produce images too.
RuntimeWhat actually starts the process (containerd, CRI-O)Docker is a toolchain around this, not a requirement.

The discipline that makes images worth having is immutability. You do not patch a running container, and you do not build a new image for each environment. The bytes that passed your tests are the bytes that serve traffic, and the only way to change them is to build again and roll out again. It sounds restrictive. It is the reason "works on my machine" stopped being a sentence anyone says.

Smaller is safer, not just faster

Every package in your base image is code you are responsible for and a scanner will eventually shout about. Distroless and minimal bases cut the attack surface as much as the pull time. Add a non-root user, a read-only root filesystem and dropped capabilities, and a container escape needs considerably more than one bad dependency.

04Registries: where the truth is stored

A registry is a content-addressed store for images. You push, it deduplicates layers you already sent, and it hands back a digest — a SHA-256 hash of the manifest. That digest is the only genuinely stable name an image has.

A tag, by contrast, is a mutable pointer. app:2.4.0 can be pushed over. latest means nothing at all except "whatever was pushed most recently, by anyone". Deploying by tag means the thing running in production is whatever the tag pointed to at the moment each node happened to pull — which is how two replicas of the "same" version end up behaving differently.

# what you write, and what should end up in the manifest tag registry.example.com/app:2.4.0 mutable — a convenience digest registry.example.com/app@sha256:9c1f… immutable — the actual artefact # and the three things worth attaching to it ✓ signature cosign · keyless, tied to the CI identity ✓ SBOM every package that went in, so "are we affected?" is a query ✓ provenance which repo, which commit, which builder produced it

Signing matters because a registry is a supply chain, and a supply chain is worth attacking. Verifying at admission — refusing to run an image that is not signed by your own build system — is one of the higher-value, lower-effort controls available, and it closes the "someone pushed to the registry" hole entirely.

05Declarative: describing, not doing

The imperative way to deploy is a list of instructions: copy the artefact, stop the service, run the migration, start the service, check it. It works, and it has one fatal property — the instructions only describe the transition you were expecting. Run them on a machine that is in a state you did not anticipate and the result is undefined.

The declarative way is to write down the destination and let something else work out the route. Three replicas of this image, with this much memory, reachable on this port. Nothing about how to get there from here.

# the whole idea, in eight lines apiVersion: apps/v1 kind: Deployment spec: replicas: 3 template: spec: containers: - image: registry.example.com/app@sha256:9c1f…

This is what makes the machines disposable. The description does not mention a server, so it does not care which servers exist. Apply the same file to an empty cluster and you get three replicas; apply it to a cluster that already has three and nothing happens; apply it after two machines burn down and you get two replacements. The same input, in three different worlds, produces the right thing each time — because it describes a destination rather than a journey.

Level-triggered, not edge-triggered

Controllers do not react to events; they react to differences. A missed event in an edge-triggered system leaves it permanently wrong. A missed event in a level-triggered one is corrected on the next pass, because the next pass asks "what is the state?" rather than "what just happened?". This is unglamorous and it is the single most important design decision in the whole ecosystem.

Keep those descriptions in Git and you get GitOps: the repository is the intended state, an agent in the cluster continuously pulls and applies it, and any manual change made in a hurry gets reverted within the minute. Deployment becomes a merge; rollback becomes a revert; the audit log is git log. Access to production becomes access to a pull request.

06Orchestration: the loop that never stops

The control tower in the middle of the map has cables running to almost every other station, and it sits inside the circuit rather than on it, because the control plane is not a stage of anything. It is a set of processes running the same small loop, forever: read the desired state, read the actual state, take one step to reduce the difference.

ComponentJobIf it stops
API serverThe only way in. Validates, authenticates, and writes to the storeYou cannot change anything — but everything running keeps running
etcdThe consistent store of record for all cluster stateThis is the one to back up. Everything else can be rebuilt
SchedulerChooses which machine each new pod goes onNew pods stay Pending; existing ones are unaffected
ControllersOne loop per kind of thing — deployments, replica sets, jobs, endpointsReality slowly stops matching the description
KubeletOn every machine: starts containers, reports healthThat machine's pods are declared lost and rescheduled elsewhere

Notice what is not in that table: anything that handles a user request. The control plane is not in the data path. A cluster whose API server is down still serves traffic perfectly well; it simply cannot be changed. That separation is deliberate and it is why the failure modes are survivable.

The best part is that you can extend it

The loop is not special-cased for deployments. Define your own resource type — a Database, a TenantEnvironment, whatever your domain needs — write a controller that reconciles it, and it behaves exactly like the built-in kinds: same API, same access control, same declarative workflow. This is the operator pattern, and it is the most genuinely novel thing in the ecosystem.

Kubernetes is the one everyone means, but the pattern is bigger than it: Nomad, ECS and most serverless platforms are the same loop with different amounts of it hidden.

07Pods and quays: where things actually run

The unit that gets scheduled is not a container but a pod: one or more containers that share a network namespace, a set of volumes and a lifetime. They are always on the same machine and can talk over localhost. Most pods hold one real container plus, sometimes, a sidecar doing something adjacent — proxying traffic, shipping logs, refreshing a certificate.

Pods are deliberately cheap and deliberately temporary. You do not create them directly; you declare a Deployment and a controller creates them for you, replacing any that die and never reusing a name. A pod that is restarted is a new pod, with a new address, and anything that cached the old one is now wrong.

Requests

What the pod is guaranteed. The scheduler uses this — and only this — to decide whether a machine has room.

Limits

The ceiling. Over on CPU and you are throttled; over on memory and the kernel kills you outright, with no chance to explain.

Affinity & spread

Rules that keep replicas apart — different machines, different zones — so one failure cannot take all of them.

Taints & tolerations

Machines that repel pods unless a pod explicitly says it can cope: GPU nodes, spot instances, anything unusual.

Getting requests wrong is the most common source of mysterious cluster behaviour. Set them too high and you pay for machines that are two-thirds empty, because the scheduler believes them full. Set them too low and pods land on machines with nothing left to give, and the whole node degrades at once. Setting a memory limit far above the request is the classic way to build a cluster that works fine until the day it does not.

OOMKilled is not a crash

When a container exceeds its memory limit the kernel terminates it immediately — no signal it can handle, no shutdown hook, no log line explaining why. The evidence is an exit code of 137 and a terse event. If a service "restarts randomly under load", check this before anything else.

08Networking: stable names over moving parts

Every pod gets its own IP address, and every pod can reach every other pod without translation. That flat model is what makes the rest tractable — and it means pod addresses are useless to hold onto, because the set of them changes constantly.

A Service is the fix: one stable name and address in front of a set of pods that a controller keeps up to date. Ask DNS for checkout-svc and you get something that will still resolve after every pod behind it has been replaced twice. Traffic to it is load-balanced across whichever endpoints are currently ready.

LayerWhat it doesWhere it lives
ServiceA stable virtual address over a changing set of podsInside the cluster
Ingress / GatewayTerminates TLS, matches host and path, routes to a serviceThe one controlled way in
Load balancerThe cloud's own front door, pointing at the gatewayOutside, billed hourly
Service meshmTLS, retries, timeouts, traffic splitting, per-call telemetryBeside or beneath every pod
Network policyWhich pods may talk to which — deny by default, if you are sensibleEnforced by the CNI plugin

The mesh is the one that needs justifying. It moves concerns that every service would otherwise implement — mutual TLS, retry with backoff, circuit breaking, timeouts, traffic splitting for canaries, and a consistent trace for every call — out of the application and into infrastructure. That is a real win in a large estate with several languages. It is also a second network to debug, and in a system of six services it usually costs more than it returns.

The default worth changing on day one

Out of the box, every pod can talk to every other pod, across every namespace. That is a convenient default and a poor security posture: one compromised container can reach your database, your metrics store and everything else. A default-deny network policy per namespace, with explicit allowances, is cheap to add early and miserable to retrofit late.

09Configuration and secrets

One image, many environments — so everything that differs between them has to arrive from outside the image, at start-up. Non-sensitive settings go in a ConfigMap; credentials go in a Secret; both can be injected as environment variables or mounted as files.

The name "Secret" oversells it. A Kubernetes Secret is base64-encoded, which is an encoding, not encryption — anyone who can read the object can read the value. Making it a secret in the ordinary sense takes three further steps: encrypt etcd at rest, restrict read access with RBAC so that "can list secrets in this namespace" is a rare permission, and keep the real values in a dedicated store that syncs them in.

External secret stores

Vault, or the cloud's own manager, holding the truth. The cluster gets a synced copy with a short life and an audit trail behind it.

Workload identity

Better still: no long-lived credential at all. The pod proves who it is and receives a token that expires in minutes.

Never in the image

A secret baked into a layer is in the registry forever, for everyone who can pull it, including in the layer you thought you deleted it from.

Rotation, rehearsed

A credential you have never rotated is one you cannot rotate. Find out on a Tuesday, not during the incident that requires it.

Config changes are deployments. A ConfigMap edited by hand at 2am, with no corresponding change in Git, is exactly the kind of invisible drift the declarative model exists to prevent — and a GitOps agent will silently undo it, which is its own kind of confusing incident.

10State: the part that is not cattle

Everything so far assumes instances are interchangeable. Data is the place where that assumption breaks, and pretending otherwise is the most expensive mistake available in this whole area.

The primitives exist. A PersistentVolumeClaim asks for storage; a CSI driver provisions a real disk from the cloud; a StatefulSet gives pods stable identities and stable volumes, and starts and stops them in order. That is enough to run a database on Kubernetes, and plenty of people do.

Whether you should is a different question. Running a database well means backups you have restored from, replication with a tested failover, version upgrades that do not lose writes, and someone awake who understands its consistency model. None of that gets easier by being inside a cluster; several parts get harder, because now the orchestrator can also decide to move your primary at an inconvenient moment.

The honest default

Rent your databases, queues and object storage from a provider whose job it is. Run your stateless services on the cluster. Revisit only when you have a specific reason — cost at real scale, a regulator, a database with an unusually good operator — and not because homogeneity feels tidy. Most teams who run their own end up with a part-time database administrator who did not apply for the job.

Two properties are worth designing for whatever you choose. Idempotency: a retried request must not charge the card twice, because in a system with retries at three layers it will be retried. And backward-compatible migrations: during a rolling update both versions of your code run at once, against one schema, so any migration that only works after the old code is gone will break during every deploy.

11Scaling: two loops, at different speeds

Elasticity is the part most people buy this for, and it is genuinely good — provided you know there are two separate mechanisms, running at different speeds and costing different amounts.

The fast loop adds pods. A horizontal autoscaler watches a metric — CPU, or better, something meaningful like requests in flight or queue depth — and edits the replica count in the declared state. Everything after that is the ordinary reconciliation loop. Because the image is usually already on the machine, new capacity arrives in seconds.

The slow loop adds machines. When no node has room, a cluster autoscaler asks the cloud provider for another one: a minute or two to boot, join and pull images, and a bill that starts immediately. This is the loop that has to be tuned carefully, because it is the one with a price.

Scale on the right signal

CPU is a proxy for load and often a bad one. Queue depth, concurrent requests or latency track what users feel.

Scale down slowly

Aggressive scale-down plus a traffic wobble gives you a system that spends its life booting. Long cooldowns, short scale-up.

Scale to zero

Great for jobs and dev environments, painful for user-facing paths — someone pays for the cold start, and it is a user.

Set a maximum

An autoscaler with no ceiling will faithfully convert a retry storm or an infinite loop into an enormous invoice.

A caution worth stating plainly: autoscaling only helps with load that your application can actually spread. Nine replicas in front of one database that is already at capacity does not serve more traffic — it queues more of it, then fails harder. Scaling reveals the next bottleneck; it does not remove it.

12Resilience: expecting the machine to die

In a traditional deployment, a server failing is an incident. Here it is a Tuesday: the cloud reclaims spot instances, nodes get drained for kernel patches, hardware fails, and the cluster is upgraded a node at a time by design. The system is not built to prevent that. It is built so that it does not matter.

MechanismWhat it protects against
Liveness probeA process that is running but wedged — restart it
Readiness probeSending traffic to a pod that is up but not yet able to serve
Startup probeA slow-booting app being killed by its own liveness check
Replicas across zonesOne machine, rack or data centre taking everything with it
Pod disruption budgetMaintenance politely draining every replica at once
Retries with backoffA brief blip becoming a user-visible error

Probes are where the subtlety is. A liveness probe that checks a dependency will restart your healthy pod because the database is slow — turning someone else's degradation into your outage, in a loop. Liveness should ask only "is this process still capable of responding?". Readiness is the one that may consider dependencies, because its consequence is being taken out of rotation rather than killed.

Retries are a load amplifier

Three layers each retrying three times is twenty-seven requests to a service that is already struggling, which is how a slow dependency becomes a dead one. Retry at one layer, with a budget, jitter and a circuit breaker — and make sure the operation is idempotent before you retry it at all.

The honest framing of resilience is not "nothing goes wrong". It is a small, bounded amount of harm: a handful of requests that took a retry, contained within a blast radius you chose in advance. Watch the A quay goes under scenario behind this page — a few requests do fail. That is what success looks like.

13Progressive delivery: small steps, walked back

Because a deployment is an edit to a description, releasing gets to be gradual by default. The controller replaces pods a few at a time, waiting for each batch to pass its readiness probe before continuing, so at no point is all capacity running the new version — and at no point is it running none of the old one.

Rolling update

The default. Replace a few pods at a time, gated on readiness. Cheap, and enough for most changes.

Blue-green

Two full versions side by side; flip the service to the new one when you are satisfied. Rollback is flipping back.

Canary

A small share of real traffic to the new version, compared against the old one, widened automatically if the metrics agree.

Feature flags

Separates deploying from releasing entirely. The most flexible option, and the one that leaves the most debris behind.

The canary is worth the extra machinery for anything user-facing, and the reason is statistical rather than operational: it gives you a control group. "Error rate is 0.4%" means very little on its own. "Error rate is 0.4% on the new version and 0.01% on the old one, right now, under identical traffic" is conclusive in about ninety seconds.

Rollback then costs almost nothing. The previous image is still in the registry, the previous manifest is still in the history, and the control plane has no opinion about direction — it closes the gap between declared and actual, whichever way that points. Which means the important question about any release is not "are we confident?" but "how quickly can we undo it, and how many people are exposed while we decide?".

14Observability: the evidence outlives the pod

When instances are disposable, so is everything they knew. The pod that produced the error is gone, its filesystem with it, and there is nothing to log into. Telemetry has to be collected somewhere durable while the pod is still alive, or it does not exist.

Metrics

Cheap numbers over time: request rate, error rate, latency percentiles, saturation. Good at telling you something is wrong.

Logs

Structured events, shipped off the node. The most expensive of the three at volume, and the most useful when you already know where to look.

Traces

One request's path across every service, with timings. Usually the fastest way to find which hop is wrong.

Events

The cluster's own narration — scheduled, pulled, failed, evicted, OOMKilled. The first place to look when a pod will not start.

OpenTelemetry is the part worth committing to. It is a vendor-neutral way to produce all three signals, which means the instrumentation in your code survives changing your mind about where to send it — and you will change your mind, because observability pricing is where cloud native budgets quietly go.

Cardinality is the bill

A metric labelled with something unbounded — a user id, a request id, a full URL — creates a separate time series for every distinct value. A single careless label can multiply storage by a hundred thousand and take the metrics system down with it. High cardinality belongs in traces and logs; metrics want a handful of low-cardinality dimensions and nothing else.

Above all this sit service level objectives: a written target such as 99.9% of requests succeeding within 300 ms over 30 days. They convert an argument about whether things feel slow into arithmetic, and the failure they permit is an error budget — spend it on shipping while it lasts, spend it on reliability when it runs out. Alert on those symptoms, not on every twitch a dashboard makes.

15What it costs, and when not to

Everything on this page is real engineering with real benefits, and it is not free. The honest accounting looks like this: you have replaced a small number of servers you understood with a distributed system that has its own control plane, its own network, its own storage abstraction, its own failure modes and its own release cadence. You will spend real time operating the thing that operates your things.

You gainYou take on
Machines become disposable, and failure becomes routineA cluster to upgrade, patch and pay for whether or not it is busy
Deployments are declarative, reviewable and reversibleA large amount of YAML, and tooling to keep it from sprawling
Elasticity, in seconds, for the parts that can spreadCost controls, or an autoscaler that turns bugs into invoices
One deployment model for every service and languageNetworking, identity and storage abstractions your team must learn
A vast ecosystem of interoperable componentsChoosing from it, and owning the choices for years

So: when should you not? If you have one application, a handful of servers, no meaningful traffic variability and a team of four, a managed platform — App Runner, Cloud Run, Fly, Render, a good old PaaS — will give you most of the benefits above for a fraction of the surface area. Container images without an orchestrator is a perfectly respectable place to stop. The people who regret adopting Kubernetes are almost always the ones who adopted it before they had the problems it solves.

Platform engineering, in one paragraph

Once the platform is genuinely useful, the failure mode changes: every team must now understand all of it. The answer is a golden path — an opinionated, supported route from repository to production that covers eighty per cent of cases with almost no decisions, plus an escape hatch for the rest. The measure of a platform team is not how much the platform can do. It is how little a product team has to know.

✳Glossary

TermIn one line
Cloud nativeBuilding so that the machines underneath are disposable and the desired state is declared, not performed.
ContainerA process on the host kernel with its own view of the filesystem, network and process table, and a cap on resources.
ImageThe immutable, layered package a container is started from.
DigestThe content hash of an image — the only name for it that nobody can move.
RegistryThe store images are pushed to and pulled from.
ManifestA declaration of what should exist, usually YAML, usually in Git.
Control planeThe API server, store, scheduler and controllers that keep reality matching the manifests.
ReconciliationThe loop comparing desired state with actual state and taking one step to close the gap.
NodeA machine in the cluster, running a kubelet and a container runtime.
PodThe smallest schedulable unit: one or more containers sharing a network namespace and volumes.
DeploymentA declaration of how many replicas of a pod should exist, and how to roll them over.
ServiceA stable name and address in front of a changing set of pods.
Ingress / GatewayThe controlled entrance from outside: TLS, host and path routing.
Service meshInfrastructure providing mTLS, retries, timeouts and per-call telemetry between services.
ConfigMap / SecretEnvironment-specific settings and credentials, injected into pods at start-up.
PVCA persistent volume claim: a request for storage that outlives the pod using it.
StatefulSetPods with stable identities, stable volumes and ordered start-up, for the things that cannot be interchangeable.
HPAHorizontal pod autoscaler: edits the replica count in response to a metric.
ProbeA periodic check deciding whether a pod should be restarted (liveness) or sent traffic (readiness).
OperatorA custom resource type plus a controller that reconciles it — the platform's own extension mechanism.
GitOpsKeeping the declared state in Git and having an agent continuously apply it.
CanaryExposing a new version to a small share of real traffic and comparing it against the old.
SBOMA software bill of materials: the inventory of everything that went into an image.
SLOA written reliability target, and the budget of failure it implies.

→Go deeper

  • Kubernetes Concepts — the official documentation, and unusually good for official documentation.
  • The Twelve-Factor App — still the clearest description of what a container-shaped program looks like.
  • CNCF — the foundation behind most of the components named on this page, and the landscape map that will overwhelm you.
  • OCI Image Spec — what an image actually is, in about twenty readable pages.
  • OpenTelemetry — vendor-neutral metrics, logs and traces.
  • Google SRE Books — service level objectives, error budgets and incident response, free online.
  • SLSA — supply chain integrity and build provenance for the images you publish.
  • How Software Gets Made — the previous instalment: how a change reaches the pipeline in the first place.
  • How the Internet Works — and the one before that: how the request found your cluster at all.
1

Loading… Starting the simulation.