Декларативная платформа оркестрации резервного копирования — control-plane запускает бэкапы через доверенные внешние утилиты (pg_dump, rclone, age и т.д.), не переизобретая протоколы.
  • Go 99.7%
  • Makefile 0.3%
Find a file
Repository files (latest commit first)
Filename Latest commit message Latest commit date
Vlad Samoylik 19c10604d9 Harden job lifecycle, agent durability, and audit follow-ups.
Require lease for progress/result, add lease renewal heartbeats, fix PollJob
payload error loops, async outbox replay with quarantine on stale results,
and close gaps in metrics, SFTP mkdir, secret resolution, and CLI defaults.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-07-12 15:11:53 +03:00
cmd Fix job lifecycle bugs, add watchdog, and enforce structured JSON logging. 2026-05-29 19:30:18 +03:00
configs Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
deployments/systemd Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
internal Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
pkg/client Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
test Initial commit: Archivist backup orchestration platform. 2026-05-20 02:16:11 +03:00
testdata/programs Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
.gitignore Initial commit: Archivist backup orchestration platform. 2026-05-20 02:16:11 +03:00
.golangci.yml Initial commit: Archivist backup orchestration platform. 2026-05-20 02:16:11 +03:00
go.mod Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
go.sum Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
Makefile Fix agent job status race and tighten review follow-ups. 2026-05-20 17:26:31 +03:00
README.md Harden job lifecycle, agent durability, and audit follow-ups. 2026-07-12 15:11:53 +03:00
README.ru.md Add upload reliability, atomic publish, and in-memory job queue. 2026-06-05 15:57:49 +03:00

Archivist

Declarative backup orchestration platform — schedules and runs backups through trusted external tools (pg_dump, s5cmd, sftp, age, etc.) instead of reimplementing protocols.

Russian: README.ru.md.


What Archivist is

Archivist is a backup control-plane, not another backup engine. It does not dump databases, speak S3, or implement encryption — it orchestrates tools you already trust (pg_dump, mysqldump, zstd, age, s5cmd, OpenSSH sftp) through a fixed, auditable pipeline.

The outcome is reproducible artifacts: every backup produces an artifact, a .sha256 sidecar, and a .manifest.json that records which tools, profiles, and args were used — without secrets.

How it differs

Typical approach Archivist
Cron + shell scripts Validated YAML, structured jobs, audit trail, manifest index
Monolithic backup suites (Bacula, proprietary agents) Thin orchestrator; dump/encrypt/upload logic stays in standard CLI tools
Deduplicating backup tools (restic, borg) Orchestration layer; compress/encrypt/storage backends chosen via profiles
Cloud-only operators (Velero, managed DB backups) Agent on the host, S3/SFTP/local storage

Security model

  • Bearer token auth — the server API requires Authorization: Bearer <token> from operators (server.operator_tokens) and agents (agents.<id>.token); tokens are compared in constant time.
  • Agent pull model — agents connect outbound to the server and poll for work; no inbound port on the agent host (works behind NAT, in Kubernetes, etc.).
  • Server-only TLS — all communication over HTTPS; agents and operators verify the server certificate (custom CA or system roots); no client certificates.
  • Least privilege on agent — optional non-root enforcement, shell forbidden in module profiles, only allowlisted template variables in args.
  • Integrity at write time — SHA-256 of the final artifact is computed before upload and stored in the manifest.
  • Secrets stay out of logs — passwords resolved from env: refs on the server; credentials sent to the agent per-job only; audit JSONL and manifests never contain them.
  • Path safety — manifest paths, output filenames, and storage remote keys are validated against directory traversal.

Reliability

  • Deterministic pipeline — dump → compress → [encrypt] → sha256 → upload.
  • Sidecar checksums — .sha256 written alongside every artifact for out-of-band verification.
  • Manifest as contract — records tools, profile names, extensions, args, and storage metadata for operational replay and debugging.
  • Startup validation — binaries, template syntax, and profile uniqueness checked before the agent accepts any work.
  • Job recovery on restart — running jobs are marked failed on server startup; pending jobs remain in queue and will be picked up on next agent poll.

Components

Binary Role
archivist-agent Restricted execution runtime on backup hosts; polls server, runs backup pipeline
archivist-server Control-plane: cron scheduler, pending job queue, file-backed state, HTTPS API
archivistctl Operator CLI: validate configs, list resources, trigger backups, inspect agents

Build

make build
go test ./...

Quick start

# 1. Generate a TLS cert for the server (or use an existing one)
openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:P-256 \
  -keyout server.key -out server.crt -days 365 -nodes \
  -subj "/CN=archivist-server" \
  -addext "subjectAltName=IP:127.0.0.1,DNS:localhost"

# 2. Generate random tokens
OPERATOR_TOKEN=$(openssl rand -hex 32)
AGENT_TOKEN=$(openssl rand -hex 32)

# 3. Fill in configs
cp configs/server.example.yaml server.yaml   # set tls paths, tokens
cp configs/agent.example.yaml  agent.yaml    # set server URL, token, workdir, modules

# 4. Start server then agent
./archivist-server -config server.yaml
./archivist-agent  -config agent.yaml

# 5. Use the CLI
cp configs/ctl.example.yaml ~/.config/archivist/ctl.yaml
# edit: set url and token
./archivistctl -profile local list targets
./archivistctl -profile local trigger backup -target app-postgresql

TLS setup

The server needs a certificate and key. Set server.tls.cert_file and server.tls.key_file. Agents and archivistctl verify it:

  • Custom CA: distribute ca.crt to agents (agent.tls.ca_file) and CLI users (servers.<name>.ca_file).
  • Public CA (Let's Encrypt, etc.): leave ca_file empty; system roots are used.
  • Development only: agent.tls.insecure_skip_verify: true skips verification.

Integration tests use ephemeral certs from internal/testpki — not for production.


Server

The server schedules backups from targets[].schedule (5-field cron) and exposes an HTTPS API protected by operator bearer tokens. State is file-backed under server.state.dir (jobs, audit JSONL, manifest index).

Server config (key fields)

server:
  listen: 0.0.0.0:8443
  tls:
    cert_file: /etc/archivist-server/tls/server.crt
    key_file:  /etc/archivist-server/tls/server.key
  operator_tokens:
    - "your-operator-token"           # generate with: openssl rand -hex 32
  state:
    dir: /var/lib/archivist-server
  dispatcher:
    agent_job_timeout: 30m            # default; min 1m
  schedule_timezone: UTC              # optional; IANA name for cron evaluation
  metrics:
    listen: 127.0.0.1:9060            # plain HTTP, separate from API

agents:
  db-prod-01:
    token: "your-agent-token"         # must match agent.token in agent config

targets:
  - name: app-postgresql
    agent: db-prod-01
    module:
      type: postgresql
      profile: pg-prod-logical
    source:
      kind: database                  # database | directory
    archive_profile: zstd-balanced
    encryption_profile: age-prod      # optional; omit for plaintext
    storage_profile: prod-s3
    storage:
      driver: s3
      prefix: backups/app-pg
      bucket: my-backups
      region: ru-central1
      endpoint: https://storage.yandexcloud.net
      access_key_id_ref: env:S3_ACCESS_KEY_ID
      secret_access_key_ref: env:S3_SECRET_ACCESS_KEY
    database:
      host: 127.0.0.1
      port: 5432
      name: myapp
      username: backup_user
      password_ref: env:BACKUP_PASSWORD
    schedule: "0 2 * * *"

See configs/server.example.yaml for full annotated example.

How jobs flow

  1. Schedule — cron triggers StartBackup → job enters pending state.
  2. Poll — agent sends POST /v1/agent/poll; server atomically claims one pending job for that agent → running.
  3. Progress — agent periodically posts POST /v1/agent/jobs/{id}/progress.
  4. Result — agent posts POST /v1/agent/jobs/{id}/result with succeeded or failed (retried from a local outbox until acknowledged).
  5. Timeout — agent_job_timeout applies to execution time after dispatch; progress renews the lease. Expired leases are re-queued instead of lost on server restart.

Agent

The agent connects outbound to archivist-server using a bearer token. No inbound port needed.

Agent config (key fields)

agent:
  id: db-prod-01
  server: https://backup.example.com:8443
  token: "your-agent-token"
  poll_interval: 10s       # default 10s; range 200ms–5m
  tls:
    ca_file: /etc/archivist-agent/server-ca.crt   # optional
    insecure_skip_verify: false                    # dev only

workdir:
  root: /var/lib/archivist-agent

limits:
  max_concurrent_jobs: 2

upload:
  retry:                     # retry-with-backoff for the upload step
    max_attempts: 5          # default 5
    initial_backoff: 1s      # default 1s
    max_backoff: 30s         # default 30s
  verify:
    mode: size               # none | size | checksum (default size)

modules:
  postgresql:
    profiles:
      pg-prod-logical:
        backup:
          binary: /usr/bin/pg_dump
          args:
            - "--host={{ .database.host }}"
            - "--port={{ .database.port }}"
            - "--username={{ .database.username }}"
            - "--format=plain"
            - "{{ .database.name }}"
          env:
            PGPASSWORD: "{{ .database.password }}"
          output:
            filename: dump.sql

archive:
  profiles:
    zstd-balanced:
      binary: /usr/bin/zstd
      args: ["-6", "-T0", "-c"]
      extension: ".zst"

encryption:
  profiles:                          # optional; omit entirely if no targets use encryption
    age-prod:
      binary: /usr/bin/age
      args: ["--encrypt", "--recipient-file", "/etc/archivist-agent/keys/age-prod.pub"]
      extension: ".age"

storages:
  prod-s3:
    driver: s3
    binary: /usr/bin/s5cmd
  local-archive:
    driver: local
    path: /var/lib/archivist-agent/storage

See configs/agent.example.yaml for full annotated example.

Startup validation

On start, the agent validates:

  • all binary paths are executable
  • template variables in args are in the allowlist ({{ .database.host }}, etc.)
  • workdir.root is an absolute path

archivistctl

Server profiles live in ~/.config/archivist/ctl.yaml:

default: local                         # optional

servers:
  local:
    url: https://127.0.0.1:8443
    token: "your-operator-token"
    ca_file: /etc/archivist/server-ca.crt   # optional
  prod:
    url: https://archivist.example.com:8443
    token: "prod-operator-token"

Global flags (before the command): -profile <name>, -config <path>, -url, -token, -ca.

Environment fallbacks (after CLI flags, before ctl profile): ARCHIVIST_TOKEN, ARCHIVIST_SERVER_URL, ARCHIVIST_CA_FILE. The ctl config file is optional when url and token are provided via flags/env.

Command Description
(no args) List configured servers
version Print CLI version
validate -config PATH [-kind agent|server|ctl] Validate a config file
health GET /v1/health
list targets GET /v1/targets
list jobs GET /v1/jobs
list agents GET /v1/agents
list audit [-limit N] GET /v1/audit?limit=N
list manifests [-target T] GET /v1/manifests?target=T
get job -id ID GET /v1/jobs/{id}
trigger backup -target T [-database D] [-no-wait] [-timeout D] [-json] POST /v1/jobs/backup then wait
inspect modules -agent ID GET /v1/agents/{id}/capabilities

trigger backup waits until succeeded or failed (default timeout 10m) and returns exit code 1 on failure. Use -no-wait to enqueue and return immediately.

cp configs/ctl.example.yaml ~/.config/archivist/ctl.yaml
# edit: set url and token per server

./archivistctl
./archivistctl version
./archivistctl validate -config agent.yaml -kind agent
./archivistctl -profile prod health
./archivistctl -profile prod list agents
./archivistctl -profile prod list jobs
./archivistctl -profile prod trigger backup -target app-postgresql
./archivistctl -profile prod trigger backup -target app-postgresql -database staging -no-wait
./archivistctl -profile prod inspect modules -agent db-prod-01
./archivistctl -profile prod list audit -limit 100

API reference

All operator routes require Authorization: Bearer <operator_token>.

Method Path Description
GET /v1/health Health check (no auth)
GET /v1/targets List configured targets
GET /v1/agents List agents with last-seen time and capabilities
GET /v1/agents/{id}/capabilities Agent capabilities reported on last poll
GET /v1/jobs List all jobs
GET /v1/jobs/{id} Get one job
GET /v1/audit Audit log (?limit=N)
GET /v1/manifests Manifest index (?target=T)
POST /v1/jobs/backup Enqueue backup job → 202 Accepted
# Enqueue a backup
curl --cacert server.crt \
  -H 'Authorization: Bearer your-operator-token' \
  -H 'Content-Type: application/json' \
  -d '{"target":"app-postgresql"}' \
  https://127.0.0.1:8443/v1/jobs/backup

# Override database for this run
curl ... -d '{"target":"app-postgresql","database_name":"staging"}' ...

# Poll status
curl --cacert server.crt \
  -H 'Authorization: Bearer your-operator-token' \
  https://127.0.0.1:8443/v1/jobs/srv-1234567890

POST /v1/jobs/backup accepts target (required) and optional database_name (overrides target's configured DB for this run). Returns the created ServerJobRecord in pending state.


Backup pipeline

dump → compress → [encrypt] → sha256 → upload
                                         ↓
                               artifact + artifact.sha256 + artifact.manifest.json

The encrypt step runs only when the target has encryption_profile set. Manifests record tool names, profile names, args (without secrets), and storage metadata.

Live progress fields

While a job runs, the agent sends incremental progress updates visible via GET /v1/jobs/{id}:

Field Description
progress.current_step Active pipeline stage: dump, compress, encrypt, checksum, upload, done
progress.database_name Database being backed up
progress.step_durations_sec Completed step durations (map)
progress.artifact_size_bytes Encrypted artifact size once known
progress.storage_errors Upload failure count
progress.storage_last_error Last upload error message

Target structure

A target combines four orthogonal blocks:

Block Purpose Required when
source { kind, name? } Identity used in remote path and manifest Always
module { type, profile } Which backup tool and profile Always
database { host, port, name, username, password_ref } DB connection + auth source.kind: database
filesystem { path } Absolute directory on agent host source.kind: directory

source.name auto-fills from database.name when not set, so most database targets need no duplication. Set source.name explicitly when the DB name contains non-FS-safe characters or you need a stable path alias.


Remote artifact layout

{prefix}/{target}/{source.name}/{YYYY}/{MM}/{source.name}-{ts}.{archExt}[{encExt}]
{prefix}/{target}/{source.name}/{YYYY}/{MM}/{source.name}-{ts}.{archExt}[{encExt}].sha256
{prefix}/{target}/{source.name}/{YYYY}/{MM}/{source.name}-{ts}.{archExt}[{encExt}].manifest.json
  • ts is always UTC in the FS-safe form YYYY-MM-DDTHH-MM-SSZ (no colons), assigned by the server at dispatch time.
  • encExt (e.g. .age) is absent when encryption_profile is not set.
  • target, source.name, and the final filename are validated as single path segments — no traversal, no slashes.

Examples:

  • backups/app-pg/app-postgresql/myapp/2026/05/myapp-2026-05-25T14-30-00Z.zst.age
  • archives/docs-archive/docs/2026/05/docs-2026-05-25T03-00-00Z.zst

Storage drivers

Driver Agent config Server target config
local path driver: local, optional prefix
sftp binary (sftp executable) host, port, username, private_key_ref, host_key_ref, prefix
s3 binary (s5cmd-compatible) bucket, region, endpoint, access_key_id_ref, secret_access_key_ref, session_token_ref, prefix

SFTP uses SSH key authentication only (no password). The server resolves _ref fields from env:VAR_NAME at dispatch time and sends credentials to the agent per-job.

Upload reliability

The agent publishes every backup as a set (artifact, .sha256, manifest, optional encrypted manifest) using a stage-then-finalize flow:

  1. Each object is uploaded to a temporary .partial key.
  2. The artifact is verified (upload.verify.mode).
  3. Objects are renamed into their final names, with manifest.json committed last as the commit marker. Consumers must ignore backup sets without a final manifest.json; artifact files may become visible before the manifest during finalize.

Configured via the agent upload section:

  • upload.retry — the upload step is retried with exponential backoff and jitter on transient errors (max_attempts, initial_backoff, max_backoff).
  • upload.verify.mode:
    • none — no post-upload check.
    • size — compare the remote object size against the local artifact (cheap; default). Best-effort for S3/SFTP since it parses s5cmd ls / sftp ls -l output.
    • checksum — re-download the artifact and compare its SHA256 (strongest; adds egress).

A crash mid-upload can leave orphaned multipart fragments and stray .partial objects. Add a bucket lifecycle rule so the object store cleans them up automatically:

{
  "Rules": [
    {
      "ID": "abort-incomplete-multipart",
      "Status": "Enabled",
      "Filter": { "Prefix": "" },
      "AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 3 }
    }
  ]
}

Optionally add an expiration rule for .partial suffixed objects to reclaim space from failed verifications.


Monitoring (Prometheus)

Enable a plain HTTP metrics listener in server config:

server:
  metrics:
    listen: 127.0.0.1:9060

Scrape http://127.0.0.1:9060/metrics. This listener is separate from the HTTPS API.

Key metrics (prefix archivist_):

Metric Type Labels
server_jobs_total Counter target, agent_id, kind, trigger, result
server_job_duration_seconds Histogram target, agent_id, kind, trigger
server_jobs_in_flight Gauge kind
target_last_success_timestamp Gauge target, agent_id
target_last_backup_size_bytes Gauge target, agent_id
target_last_failure_timestamp Gauge target, agent_id, kind
pipeline_step_duration_seconds Histogram target, agent_id, kind, step
storage_errors_total Counter target, agent_id
server_http_requests_total Counter route, method, code
server_http_request_duration_seconds Histogram route, method

trigger label: schedule (cron-initiated) or manual (API/CLI-initiated).

Example alerts:

- alert: ArchivistBackupOverdue
  expr: time() - archivist_target_last_success_timestamp > 26 * 3600
- alert: ArchivistStorageErrors
  expr: increase(archivist_storage_errors_total[5m]) > 0

Repository layout

cmd/                     binaries (server, agent, ctl)
internal/agent/          poll loop, pipeline, validation
internal/auth/           bearer token middleware
internal/controller/     server: scheduler, dispatcher, job store, HTTP handlers
internal/api/            shared types and route constants
internal/config/         config structs and validation
internal/ctl/            archivistctl commands
internal/modules/        profile-based module runner
internal/archive/        archive profile runner
internal/crypto/         encryption profile runner
internal/storage/        local, sftp, s3 drivers
internal/tlsconfig/      TLS config helpers
internal/testpki/        ephemeral certs for tests
internal/schedule/       cron parser (robfig/cron/v3) with schedule_timezone
configs/                 example YAML files
testdata/programs/       fake binaries for integration tests

Development

Integration tests use internal/testpki for ephemeral TLS certs and fake pg_dump/zstd/age binaries — no real database or manual PKI required.

go test ./...
make test-race          # recommended before releases
golangci-lint run --config .golangci.yml ./...