- Go 99.7%
- Makefile 0.3%
| Filename | Latest commit message | Latest commit date |
|---|---|---|
Require lease for progress/result, add lease renewal heartbeats, fix PollJob payload error loops, async outbox replay with quarantine on stale results, and close gaps in metrics, SFTP mkdir, secret resolution, and CLI defaults. Co-authored-by: Cursor <cursoragent@cursor.com> |
||
| cmd | ||
| configs | ||
| deployments/systemd | ||
| internal | ||
| pkg/client | ||
| test | ||
| testdata/programs | ||
| .gitignore | ||
| .golangci.yml | ||
| go.mod | ||
| go.sum | ||
| Makefile | ||
| README.md | ||
| README.ru.md | ||
Archivist
Declarative backup orchestration platform — schedules and runs backups through trusted external tools (pg_dump, s5cmd, sftp, age, etc.) instead of reimplementing protocols.
Russian: README.ru.md.
What Archivist is
Archivist is a backup control-plane, not another backup engine. It does not dump databases, speak S3, or implement encryption — it orchestrates tools you already trust (pg_dump, mysqldump, zstd, age, s5cmd, OpenSSH sftp) through a fixed, auditable pipeline.
The outcome is reproducible artifacts: every backup produces an artifact, a .sha256 sidecar, and a .manifest.json that records which tools, profiles, and args were used — without secrets.
How it differs
| Typical approach | Archivist |
|---|---|
| Cron + shell scripts | Validated YAML, structured jobs, audit trail, manifest index |
| Monolithic backup suites (Bacula, proprietary agents) | Thin orchestrator; dump/encrypt/upload logic stays in standard CLI tools |
| Deduplicating backup tools (restic, borg) | Orchestration layer; compress/encrypt/storage backends chosen via profiles |
| Cloud-only operators (Velero, managed DB backups) | Agent on the host, S3/SFTP/local storage |
Security model
- Bearer token auth — the server API requires
Authorization: Bearer <token>from operators (server.operator_tokens) and agents (agents.<id>.token); tokens are compared in constant time. - Agent pull model — agents connect outbound to the server and poll for work; no inbound port on the agent host (works behind NAT, in Kubernetes, etc.).
- Server-only TLS — all communication over HTTPS; agents and operators verify the server certificate (custom CA or system roots); no client certificates.
- Least privilege on agent — optional non-root enforcement, shell forbidden in module profiles, only allowlisted template variables in args.
- Integrity at write time — SHA-256 of the final artifact is computed before upload and stored in the manifest.
- Secrets stay out of logs — passwords resolved from
env:refs on the server; credentials sent to the agent per-job only; audit JSONL and manifests never contain them. - Path safety — manifest paths, output filenames, and storage remote keys are validated against directory traversal.
Reliability
- Deterministic pipeline — dump → compress → [encrypt] → sha256 → upload.
- Sidecar checksums —
.sha256written alongside every artifact for out-of-band verification. - Manifest as contract — records tools, profile names, extensions, args, and storage metadata for operational replay and debugging.
- Startup validation — binaries, template syntax, and profile uniqueness checked before the agent accepts any work.
- Job recovery on restart —
runningjobs are markedfailedon server startup;pendingjobs remain in queue and will be picked up on next agent poll.
Components
| Binary | Role |
|---|---|
archivist-agent |
Restricted execution runtime on backup hosts; polls server, runs backup pipeline |
archivist-server |
Control-plane: cron scheduler, pending job queue, file-backed state, HTTPS API |
archivistctl |
Operator CLI: validate configs, list resources, trigger backups, inspect agents |
Build
make build
go test ./...
Quick start
# 1. Generate a TLS cert for the server (or use an existing one)
openssl req -x509 -newkey ec -pkeyopt ec_paramgen_curve:P-256 \
-keyout server.key -out server.crt -days 365 -nodes \
-subj "/CN=archivist-server" \
-addext "subjectAltName=IP:127.0.0.1,DNS:localhost"
# 2. Generate random tokens
OPERATOR_TOKEN=$(openssl rand -hex 32)
AGENT_TOKEN=$(openssl rand -hex 32)
# 3. Fill in configs
cp configs/server.example.yaml server.yaml # set tls paths, tokens
cp configs/agent.example.yaml agent.yaml # set server URL, token, workdir, modules
# 4. Start server then agent
./archivist-server -config server.yaml
./archivist-agent -config agent.yaml
# 5. Use the CLI
cp configs/ctl.example.yaml ~/.config/archivist/ctl.yaml
# edit: set url and token
./archivistctl -profile local list targets
./archivistctl -profile local trigger backup -target app-postgresql
TLS setup
The server needs a certificate and key. Set server.tls.cert_file and server.tls.key_file. Agents and archivistctl verify it:
- Custom CA: distribute
ca.crtto agents (agent.tls.ca_file) and CLI users (servers.<name>.ca_file). - Public CA (Let's Encrypt, etc.): leave
ca_fileempty; system roots are used. - Development only:
agent.tls.insecure_skip_verify: trueskips verification.
Integration tests use ephemeral certs from internal/testpki — not for production.
Server
The server schedules backups from targets[].schedule (5-field cron) and exposes an HTTPS API protected by operator bearer tokens. State is file-backed under server.state.dir (jobs, audit JSONL, manifest index).
Server config (key fields)
server:
listen: 0.0.0.0:8443
tls:
cert_file: /etc/archivist-server/tls/server.crt
key_file: /etc/archivist-server/tls/server.key
operator_tokens:
- "your-operator-token" # generate with: openssl rand -hex 32
state:
dir: /var/lib/archivist-server
dispatcher:
agent_job_timeout: 30m # default; min 1m
schedule_timezone: UTC # optional; IANA name for cron evaluation
metrics:
listen: 127.0.0.1:9060 # plain HTTP, separate from API
agents:
db-prod-01:
token: "your-agent-token" # must match agent.token in agent config
targets:
- name: app-postgresql
agent: db-prod-01
module:
type: postgresql
profile: pg-prod-logical
source:
kind: database # database | directory
archive_profile: zstd-balanced
encryption_profile: age-prod # optional; omit for plaintext
storage_profile: prod-s3
storage:
driver: s3
prefix: backups/app-pg
bucket: my-backups
region: ru-central1
endpoint: https://storage.yandexcloud.net
access_key_id_ref: env:S3_ACCESS_KEY_ID
secret_access_key_ref: env:S3_SECRET_ACCESS_KEY
database:
host: 127.0.0.1
port: 5432
name: myapp
username: backup_user
password_ref: env:BACKUP_PASSWORD
schedule: "0 2 * * *"
See configs/server.example.yaml for full annotated example.
How jobs flow
- Schedule — cron triggers
StartBackup→ job enterspendingstate. - Poll — agent sends
POST /v1/agent/poll; server atomically claims onependingjob for that agent →running. - Progress — agent periodically posts
POST /v1/agent/jobs/{id}/progress. - Result — agent posts
POST /v1/agent/jobs/{id}/resultwithsucceededorfailed(retried from a local outbox until acknowledged). - Timeout —
agent_job_timeoutapplies to execution time after dispatch; progress renews the lease. Expired leases are re-queued instead of lost on server restart.
Agent
The agent connects outbound to archivist-server using a bearer token. No inbound port needed.
Agent config (key fields)
agent:
id: db-prod-01
server: https://backup.example.com:8443
token: "your-agent-token"
poll_interval: 10s # default 10s; range 200ms–5m
tls:
ca_file: /etc/archivist-agent/server-ca.crt # optional
insecure_skip_verify: false # dev only
workdir:
root: /var/lib/archivist-agent
limits:
max_concurrent_jobs: 2
upload:
retry: # retry-with-backoff for the upload step
max_attempts: 5 # default 5
initial_backoff: 1s # default 1s
max_backoff: 30s # default 30s
verify:
mode: size # none | size | checksum (default size)
modules:
postgresql:
profiles:
pg-prod-logical:
backup:
binary: /usr/bin/pg_dump
args:
- "--host={{ .database.host }}"
- "--port={{ .database.port }}"
- "--username={{ .database.username }}"
- "--format=plain"
- "{{ .database.name }}"
env:
PGPASSWORD: "{{ .database.password }}"
output:
filename: dump.sql
archive:
profiles:
zstd-balanced:
binary: /usr/bin/zstd
args: ["-6", "-T0", "-c"]
extension: ".zst"
encryption:
profiles: # optional; omit entirely if no targets use encryption
age-prod:
binary: /usr/bin/age
args: ["--encrypt", "--recipient-file", "/etc/archivist-agent/keys/age-prod.pub"]
extension: ".age"
storages:
prod-s3:
driver: s3
binary: /usr/bin/s5cmd
local-archive:
driver: local
path: /var/lib/archivist-agent/storage
See configs/agent.example.yaml for full annotated example.
Startup validation
On start, the agent validates:
- all
binarypaths are executable - template variables in
argsare in the allowlist ({{ .database.host }}, etc.) workdir.rootis an absolute path
archivistctl
Server profiles live in ~/.config/archivist/ctl.yaml:
default: local # optional
servers:
local:
url: https://127.0.0.1:8443
token: "your-operator-token"
ca_file: /etc/archivist/server-ca.crt # optional
prod:
url: https://archivist.example.com:8443
token: "prod-operator-token"
Global flags (before the command): -profile <name>, -config <path>, -url, -token, -ca.
Environment fallbacks (after CLI flags, before ctl profile): ARCHIVIST_TOKEN, ARCHIVIST_SERVER_URL, ARCHIVIST_CA_FILE. The ctl config file is optional when url and token are provided via flags/env.
| Command | Description |
|---|---|
| (no args) | List configured servers |
version |
Print CLI version |
validate -config PATH [-kind agent|server|ctl] |
Validate a config file |
health |
GET /v1/health |
list targets |
GET /v1/targets |
list jobs |
GET /v1/jobs |
list agents |
GET /v1/agents |
list audit [-limit N] |
GET /v1/audit?limit=N |
list manifests [-target T] |
GET /v1/manifests?target=T |
get job -id ID |
GET /v1/jobs/{id} |
trigger backup -target T [-database D] [-no-wait] [-timeout D] [-json] |
POST /v1/jobs/backup then wait |
inspect modules -agent ID |
GET /v1/agents/{id}/capabilities |
trigger backup waits until succeeded or failed (default timeout 10m) and returns exit code 1 on failure. Use -no-wait to enqueue and return immediately.
cp configs/ctl.example.yaml ~/.config/archivist/ctl.yaml
# edit: set url and token per server
./archivistctl
./archivistctl version
./archivistctl validate -config agent.yaml -kind agent
./archivistctl -profile prod health
./archivistctl -profile prod list agents
./archivistctl -profile prod list jobs
./archivistctl -profile prod trigger backup -target app-postgresql
./archivistctl -profile prod trigger backup -target app-postgresql -database staging -no-wait
./archivistctl -profile prod inspect modules -agent db-prod-01
./archivistctl -profile prod list audit -limit 100
API reference
All operator routes require Authorization: Bearer <operator_token>.
| Method | Path | Description |
|---|---|---|
GET |
/v1/health |
Health check (no auth) |
GET |
/v1/targets |
List configured targets |
GET |
/v1/agents |
List agents with last-seen time and capabilities |
GET |
/v1/agents/{id}/capabilities |
Agent capabilities reported on last poll |
GET |
/v1/jobs |
List all jobs |
GET |
/v1/jobs/{id} |
Get one job |
GET |
/v1/audit |
Audit log (?limit=N) |
GET |
/v1/manifests |
Manifest index (?target=T) |
POST |
/v1/jobs/backup |
Enqueue backup job → 202 Accepted |
# Enqueue a backup
curl --cacert server.crt \
-H 'Authorization: Bearer your-operator-token' \
-H 'Content-Type: application/json' \
-d '{"target":"app-postgresql"}' \
https://127.0.0.1:8443/v1/jobs/backup
# Override database for this run
curl ... -d '{"target":"app-postgresql","database_name":"staging"}' ...
# Poll status
curl --cacert server.crt \
-H 'Authorization: Bearer your-operator-token' \
https://127.0.0.1:8443/v1/jobs/srv-1234567890
POST /v1/jobs/backup accepts target (required) and optional database_name (overrides target's configured DB for this run). Returns the created ServerJobRecord in pending state.
Backup pipeline
dump → compress → [encrypt] → sha256 → upload
↓
artifact + artifact.sha256 + artifact.manifest.json
The encrypt step runs only when the target has encryption_profile set. Manifests record tool names, profile names, args (without secrets), and storage metadata.
Live progress fields
While a job runs, the agent sends incremental progress updates visible via GET /v1/jobs/{id}:
| Field | Description |
|---|---|
progress.current_step |
Active pipeline stage: dump, compress, encrypt, checksum, upload, done |
progress.database_name |
Database being backed up |
progress.step_durations_sec |
Completed step durations (map) |
progress.artifact_size_bytes |
Encrypted artifact size once known |
progress.storage_errors |
Upload failure count |
progress.storage_last_error |
Last upload error message |
Target structure
A target combines four orthogonal blocks:
| Block | Purpose | Required when |
|---|---|---|
source { kind, name? } |
Identity used in remote path and manifest | Always |
module { type, profile } |
Which backup tool and profile | Always |
database { host, port, name, username, password_ref } |
DB connection + auth | source.kind: database |
filesystem { path } |
Absolute directory on agent host | source.kind: directory |
source.name auto-fills from database.name when not set, so most database targets need no duplication. Set source.name explicitly when the DB name contains non-FS-safe characters or you need a stable path alias.
Remote artifact layout
{prefix}/{target}/{source.name}/{YYYY}/{MM}/{source.name}-{ts}.{archExt}[{encExt}]
{prefix}/{target}/{source.name}/{YYYY}/{MM}/{source.name}-{ts}.{archExt}[{encExt}].sha256
{prefix}/{target}/{source.name}/{YYYY}/{MM}/{source.name}-{ts}.{archExt}[{encExt}].manifest.json
tsis always UTC in the FS-safe formYYYY-MM-DDTHH-MM-SSZ(no colons), assigned by the server at dispatch time.encExt(e.g..age) is absent whenencryption_profileis not set.target,source.name, and the final filename are validated as single path segments — no traversal, no slashes.
Examples:
backups/app-pg/app-postgresql/myapp/2026/05/myapp-2026-05-25T14-30-00Z.zst.agearchives/docs-archive/docs/2026/05/docs-2026-05-25T03-00-00Z.zst
Storage drivers
| Driver | Agent config | Server target config |
|---|---|---|
local |
path |
driver: local, optional prefix |
sftp |
binary (sftp executable) |
host, port, username, private_key_ref, host_key_ref, prefix |
s3 |
binary (s5cmd-compatible) |
bucket, region, endpoint, access_key_id_ref, secret_access_key_ref, session_token_ref, prefix |
SFTP uses SSH key authentication only (no password). The server resolves _ref fields from env:VAR_NAME at dispatch time and sends credentials to the agent per-job.
Upload reliability
The agent publishes every backup as a set (artifact, .sha256, manifest, optional encrypted manifest) using a stage-then-finalize flow:
- Each object is uploaded to a temporary
.partialkey. - The artifact is verified (
upload.verify.mode). - Objects are renamed into their final names, with
manifest.jsoncommitted last as the commit marker. Consumers must ignore backup sets without a finalmanifest.json; artifact files may become visible before the manifest during finalize.
Configured via the agent upload section:
upload.retry— the upload step is retried with exponential backoff and jitter on transient errors (max_attempts,initial_backoff,max_backoff).upload.verify.mode:none— no post-upload check.size— compare the remote object size against the local artifact (cheap; default). Best-effort for S3/SFTP since it parsess5cmd ls/sftp ls -loutput.checksum— re-download the artifact and compare its SHA256 (strongest; adds egress).
S3 lifecycle (recommended)
A crash mid-upload can leave orphaned multipart fragments and stray .partial objects. Add a bucket lifecycle rule so the object store cleans them up automatically:
{
"Rules": [
{
"ID": "abort-incomplete-multipart",
"Status": "Enabled",
"Filter": { "Prefix": "" },
"AbortIncompleteMultipartUpload": { "DaysAfterInitiation": 3 }
}
]
}
Optionally add an expiration rule for .partial suffixed objects to reclaim space from failed verifications.
Monitoring (Prometheus)
Enable a plain HTTP metrics listener in server config:
server:
metrics:
listen: 127.0.0.1:9060
Scrape http://127.0.0.1:9060/metrics. This listener is separate from the HTTPS API.
Key metrics (prefix archivist_):
| Metric | Type | Labels |
|---|---|---|
server_jobs_total |
Counter | target, agent_id, kind, trigger, result |
server_job_duration_seconds |
Histogram | target, agent_id, kind, trigger |
server_jobs_in_flight |
Gauge | kind |
target_last_success_timestamp |
Gauge | target, agent_id |
target_last_backup_size_bytes |
Gauge | target, agent_id |
target_last_failure_timestamp |
Gauge | target, agent_id, kind |
pipeline_step_duration_seconds |
Histogram | target, agent_id, kind, step |
storage_errors_total |
Counter | target, agent_id |
server_http_requests_total |
Counter | route, method, code |
server_http_request_duration_seconds |
Histogram | route, method |
trigger label: schedule (cron-initiated) or manual (API/CLI-initiated).
Example alerts:
- alert: ArchivistBackupOverdue
expr: time() - archivist_target_last_success_timestamp > 26 * 3600
- alert: ArchivistStorageErrors
expr: increase(archivist_storage_errors_total[5m]) > 0
Repository layout
cmd/ binaries (server, agent, ctl)
internal/agent/ poll loop, pipeline, validation
internal/auth/ bearer token middleware
internal/controller/ server: scheduler, dispatcher, job store, HTTP handlers
internal/api/ shared types and route constants
internal/config/ config structs and validation
internal/ctl/ archivistctl commands
internal/modules/ profile-based module runner
internal/archive/ archive profile runner
internal/crypto/ encryption profile runner
internal/storage/ local, sftp, s3 drivers
internal/tlsconfig/ TLS config helpers
internal/testpki/ ephemeral certs for tests
internal/schedule/ cron parser (robfig/cron/v3) with schedule_timezone
configs/ example YAML files
testdata/programs/ fake binaries for integration tests
Development
Integration tests use internal/testpki for ephemeral TLS certs and fake pg_dump/zstd/age binaries — no real database or manual PKI required.
go test ./...
make test-race # recommended before releases
golangci-lint run --config .golangci.yml ./...