Docs Navigation
Docs/Security & Reliability

High Availability & Replication

Running multiple stateless KOLMOS replicas over shared S3/R2 storage.

Last Updated: August 2026

1. HA Architecture: Peers, Not Leader/Follower

KOLMOS does not implement Raft, leader election, or a `--role leader/follower` flag. High availability instead comes from running multiple stateless, equal `kolmos serve` processes pointed at the **same** S3/R2 bucket (`--store s3`). Object storage already provides the durability and erasure coding you'd otherwise build replication for, so KOLMOS's job is just making sure every replica sees new commits promptly. Note this model requires `--store s3` — a `--store local` deployment is inherently single-node, since local disk isn't shared between processes.

2. The Background Refresh Worker

Each replica runs a background worker that periodically catches its in-memory catalog up to the bucket's current HEAD, picking up commits any peer has written. Control the cadence with `--refresh-interval-secs` on `serve` (default `5`; pass `0` to disable). This makes cross-replica writes visible — the underlying compare-and-swap on the HEAD pointer already made concurrent writes from multiple replicas *safe*; the refresh worker is what makes them *visible* elsewhere. The consistency model is eventual between ticks, not linearizable: a client hitting replica B may not immediately see a commit just written on replica A.

bashKOLMOS Reference
kolmos --root /data serve --store s3 --s3-bucket kolmos-prod \
  --host 0.0.0.0 --refresh-interval-secs 5

3. Health Checks for Load Balancers

`GET /healthz` on the `--metrics-port` listener returns 200 while the bucket is reachable and the refresh worker isn't failing, and 503 otherwise — wire it up as your load balancer or Kubernetes readiness/liveness probe so traffic drains away from a replica that's lost bucket connectivity.

bashKOLMOS Reference
curl http://127.0.0.1:9090/healthz

4. Known Limitations (v1)

- No leader lease: background vacuum runs redundantly on every replica. Safe, but wastes some compute — expect this to be optimized in a future release. - Consistency is eventual, not linearizable — don't assume read-your-own-write across two different replica connections without routing sticky sessions. - `--store local` cannot be scaled this way; multi-replica HA requires shared S3/R2-compatible storage.

5. Example: 3-Replica Kubernetes Deployment

yamlKOLMOS Reference
apiVersion: apps/v1
kind: Deployment
metadata:
  name: kolmos
spec:
  replicas: 3
  selector:
    matchLabels: { app: kolmos }
  template:
    metadata:
      labels: { app: kolmos }
    spec:
      containers:
        - name: kolmos
          image: kolmos/kolmos:latest
          args: ["--root", "/data", "serve", "--host", "0.0.0.0",
                 "--store", "s3", "--s3-bucket", "kolmos-prod",
                 "--metrics-port", "9090"]
          readinessProbe:
            httpGet: { path: /healthz, port: 9090 }
          livenessProbe:
            httpGet: { path: /healthz, port: 9090 }
---
apiVersion: v1
kind: Service
metadata:
  name: kolmos
spec:
  selector: { app: kolmos }
  ports:
    - { name: postgres, port: 5432, targetPort: 5432 }