Production deployment
Deploy Observal as a distributed, highly available stack on AWS or GCP using Terraform. Best for SLA-bound deployments, teams over 50 users, and any environment where uptime and horizontal scaling matter.
End state: a fully managed Observal install with autoscaling compute, managed databases with automated failover, encrypted secrets, centralized logging, automated backups, and HTTPS on your domain — provisioned with a single terraform apply.
When to use this vs. single-node
Best for
Small/mid teams, internal use, POCs
Enterprise, SLA-bound, high-traffic
Infra
1 VM, any cloud or on-prem
~100 managed cloud resources
Cost
$20–150/mo
~$255/mo (AWS)
HA
No — single point of failure
Yes — Multi-AZ databases, autoscaling compute
Failover
Manual (reboot / redeploy)
Automatic (managed services handle it)
Scaling
Vertical (bigger VM)
Horizontal (add more containers)
Backups
Cron + S3
Automated (RDS snapshots, systemd timers, S3 lifecycle)
Secrets
Server-package files, or direct .env values for source Compose
SSM Parameter Store / Secret Manager (encrypted, auditable)
Logging
docker compose logs
CloudWatch / Cloud Logging (centralized, retained, searchable)
Time to deploy
10 minutes
20–30 minutes
Choose your cloud
Observal ships Terraform modules for both AWS and GCP. They provision equivalent architectures using each cloud's native managed services.
AWS
Compute (stateless)
ECS Fargate
api, web, worker as separate services
Postgres
RDS Postgres 16
Multi-AZ on prod, encrypted, automated backups
Redis
ElastiCache Redis 7
2-node replication, automatic failover on prod
ClickHouse
EC2 + EBS gp3
Self-hosted; option to use ClickHouse Cloud
Load balancer
ALB
HTTPS via ACM, path-based routing
Secrets
SSM Parameter Store
Encrypted with KMS, injected into ECS tasks
Logging
CloudWatch
Per-service log groups
Backups
S3 + RDS snapshots
Lifecycle: Standard → IA → Glacier → expire
DNS
Route 53
Optional; works without a custom domain
Full walkthrough: AWS deployment with Terraform
Quick start:
Apply takes 20–30 minutes (RDS Multi-AZ provisioning dominates). Baseline cost: **$255/mo** in us-east-1.
GCP
Compute (stateless)
Cloud Run v2
api, web, worker as separate services
Migrations
Cloud Run Jobs
One-shot init task
Postgres
Cloud SQL
HA configuration available
Redis
Memorystore
Managed Redis
ClickHouse
GCE instance
Docker Compose on a single VM; option for ClickHouse Cloud
Load balancer
Global HTTPS LB
Managed SSL certificate
Secrets
Secret Manager
Injected into Cloud Run at start
Logging
Cloud Logging
Built-in, no config needed
Backups
GCS
Versioned bucket with lifecycle
DNS
Cloud DNS
Optional
Shell access
IAP SSH tunnel
No public SSH, no SSH keys
Full walkthrough: GCP deployment with Terraform
Quick start:
Day-2 operations
These apply to both AWS and GCP. Cloud-specific commands are in the respective guides.
Upgrade to a new release
The init task re-runs migrations. Compute services roll over with zero downtime.
Roll back
Set image_tag to the previous version and terraform apply. Database data is not affected. If the failed version ran a destructive migration (rare, always called out in the CHANGELOG), restore from the pre-upgrade database backup.
Scale up
Adjust variables in terraform.tfvars:
ClickHouse Cloud (for HA)
The self-hosted ClickHouse is a single instance. For real high availability:
The EC2/GCE data host is skipped entirely. You become responsible for Grafana hosting (AWS Managed Grafana or a separate Cloud Run service).
Tear down
On prod, databases have deletion protection enabled. Disable manually before destroy if you really mean it.
Cost comparison
AWS (us-east-1, on-demand)
Fargate api 2×
0.5 vCPU / 1 GB
$30
Fargate web 2×
0.25 vCPU / 0.5 GB
$15
Fargate worker 1×
0.5 vCPU / 1 GB
$15
EC2 data host
t3.large
$60
RDS Postgres Multi-AZ
db.t4g.small
$50
ElastiCache Redis 2×
cache.t4g.micro
$25
ALB
—
$20
NAT Gateway
—
$33 + egress
EBS gp3 100 GB
—
$8
S3 backups
~1 GB cold
$0.10
Total
~$255
Set environment = "staging" to halve the bill (single-AZ RDS, one Redis node, no deletion protection).
GCP (us-central1, on-demand)
Cloud Run api
1 vCPU / 512 MB, min 1 instance
$25
Cloud Run web
1 vCPU / 256 MB, min 1 instance
$15
Cloud Run worker
1 vCPU / 512 MB, min 1 instance
$25
Cloud SQL Postgres
db-f1-micro (shared)
$10
Memorystore Redis
1 GB basic
$35
GCE data host
e2-standard-2
$50
Global HTTPS LB
—
$18
Persistent Disk 100 GB
—
$4
GCS backups
~1 GB cold
$0.02
Total
~$180
GCP is typically cheaper due to Cloud Run's per-request billing and no NAT Gateway equivalent charge.
Production hardening checklist
Applies to both clouds. See the cloud-specific guides for implementation details.
Next
AWS deployment with Terraform — full AWS walkthrough
GCP deployment with Terraform — full GCP walkthrough
Configuration — environment variables reference
Upgrades — upgrade and rollback procedures
Backup and restore — detailed restore procedures
Last updated
Was this helpful?