Slashing AWS Infra Costs by 68% with Hetzner K3s and WireGuard: A Lagos Fintech Case Study
When the Nigerian Naira devalued in late 2023, AWS bills tripled overnight for mid-stage startups. Here is how we migrated a high-volume payments client from AWS EKS to a self-managed Hetzner K3s cluster behind Cloudflare, saving 68% monthly without sacrificing reliability.
In late 2023, the sudden floating of the Nigerian Naira turned managed cloud infrastructure into a financial emergency. A mid-sized Lagos fintech handling 1.8 million monthly transaction events saw its monthly AWS bill jump from ₦1.9 million to over ₦6.3 million practically overnight. The underlying USD spend hadn't changed—it was running around $4,200 per month—but local revenue could not absorb a 3x multiplier on fixed infrastructure overhead.
Management gave us a clear directive: reduce infrastructure costs by at least 50% within six weeks without increasing transaction processing latency or breaking strict uptime requirements.
This case study details how we moved the client's production workloads off Elastic Kubernetes Service (EKS) and RDS onto a multi-region Hetzner Cloud footprint running K3s, secured with WireGuard, and routed via Cloudflare.
The Starting Architecture and the Cost Bottleneck
When we audited the existing footprint, the application was typical of modern fintech startups built on AWS defaults:
- Compute: AWS EKS running two
m5.largeworker nodes for stateless microservices ineu-west-1(Ireland). - Database: Managed Multi-AZ PostgreSQL on AWS RDS (
db.r6g.xlarge). - Caching & Queues: AWS ElastiCache Redis (
cache.m6g.large) and SQS queues. - Ingress & Networking: AWS Application Load Balancer (ALB) with NAT Gateways handling outbound egress across two Availability Zones.
The system performed reliably, but the bill breakdown revealed major inefficiencies:
| Service Component | Monthly Cost (USD) | Primary Cost Driver |
| :--- | :--- | :--- |
| EKS Control Plane | $73.00 | Flat fee per cluster |
| EC2 Worker Nodes | $288.00 | $0.096/hr per m5.large instance |
| RDS PostgreSQL (Multi-AZ) | $745.00 | High RAM reservation + Provisioned IOPS (gp3) |
| ElastiCache Redis | $196.00 | Dedicated instance hours |
| NAT Gateway & Egress Data | $1,120.00 | Data processing ($0.045/GB) + Egress rates |
| CloudWatch & GuardDuty | $480.00 | Ingestion volume and log retention |
| Support & Misc (S3, ECR) | $1,300.00 | Cross-region transfers, enterprise backup retention |
| Total | $4,202.00 | |
Data transfer costs alone were crippling. The application processed continuous webhook payloads from local switches and banks, generating over 18 Terabytes of cross-region transfer and NAT processing monthly. Every webhook hit the AWS ALB, went through the NAT Gateway, triggered DB writes, and sent out telemetry.
Worse, AWS af-south-1 (Cape Town) wasn't an easy fix: bandwidth egress out of Cape Town is significantly more expensive than European regions, and latency to major Nigerian ISPs like MainOne or Glo fluctuated wildly due to transit routing via Europe anyway, a phenomenon we explored when analyzing telemetry resilience in When MainOne Goes Down: Buffering Vector and Grafana Loki Pipelines for Flaky West African Transit.
Constraints and Design Criteria
We couldn't just throw everything onto a single cheap VPS and call it a day. Financial workflows demand durability, quick recovery, and zero data loss.
- Zero Downtime Migration: The system processed payment disbursements and webhooks 24/7. Downtime meant dropped webhooks and lost revenue.
- Low Latency for Webhook Ingestion: Payment processors like Paystack and Flutterwave expect HTTP 200 responses within 2,000ms. High round-trip time (RTT) leads to aggressive retries, which frequently trigger race conditions like those discussed in our analysis of Preventing Double-Crediting in Paystack Webhooks: Postgres Advisory Locks vs. Redis Redlock.
- Data Loss Prevention: RPO (Recovery Point Objective) had to be zero for committed transactions, and RTO (Recovery Time Objective) under 15 minutes.
- Automated Deployments: The engineering team was accustomed to pushing code via GitHub Actions to ECR and triggering rolling updates. The new setup could not compromise developer velocity.
What Failed: The Managed Cloud Optimization Fallacy
Our initial attempt focused on standard cloud optimization inside AWS. We tried:
- 3-Year Compute Savings Plans: This offered a ~35% discount on EC2 and EKS worker nodes, but locked the client into USD-denominated contracts during extreme currency volatility. It was a severe balance sheet risk.
- Graviton Migrations (
m5tom6g): Dropped compute costs by ~15%, but did nothing to address NAT Gateway processing fees or RDS storage unit economics. - Self-Hosting Redis on EC2: Saved $150/month, but required managing automated failover manually.
These tweaks brought the monthly bill down from $4,200 to around $3,450. A 18% reduction was nowhere near enough. Managed AWS services carry an implicit tax that small-to-midscale African fintechs simply cannot afford when local currency drops 60% against the dollar.
The Architecture That Worked: Hetzner Cloud + K3s + WireGuard Overlay
We pivoted to bare-metal/cloud VPS provider Hetzner, hosting nodes in Falkenstein and Frankfurt, paired with Cloudflare Enterprise for ingress and Edge DDoS protection.

1. Compute Layer: K3s over Full Kubernetes
Instead of heavy upstream Kubernetes or managed EKS, we deployed lightweight K3s. K3s bundles control plane components into a single process, consuming under 512MB RAM per node.
We provisioned four Hetzner CX42 dedicated vCPU instances (8 vCPU, 32 GB RAM, 240 GB NVMe SSD) at roughly €24/month each.
- Node 1 & 2: Control plane + Stateful services (Redis, Vector collectors).
- Node 3 & 4: Worker nodes running stateless API services and worker pools.
2. Networking and Security: WireGuard Mesh
Because Hetzner Cloud servers lack private VPC isolation out of the box across different datacenters, we established an encrypted overlay network using WireGuard. K3s supports Flannel with WireGuard backend backend (flannel-backend=wireguard-native), encrypting all pod-to-pod traffic between nodes across public networks without measurable CPU overhead.
To manage external ingress without exposing raw VPS IPs to public scans, Cloudflare sits in front of the cluster. Cloudflare handles SSL termination, Web Application Firewall (WAF), and Anycast routing, directing traffic straight to Traefik ingress controllers running on our K3s nodes.
3. Database Layer: Self-Hosted PostgreSQL with Automated S3 WAL-G Backups
Instead of RDS, we deployed PostgreSQL 16 on dedicated Hetzner AX52 bare-metal servers (AMD Ryzen 7 7700, 64 GB DDR5 RAM, 2x 1 TB NVMe SSDs in RAID 1) costing €69/month.
To ensure durability:
- We configured point-in-time recovery (PITR) using
WAL-Gstreaming continuous write-ahead logs to an S3-compatible bucket (Backblaze B2, costing $6/month per TB). - A standby replica was set up in a separate Hetzner availability zone running asynchronous streaming replication.
bash # Sample WAL-G archiving configuration in postgresql.conf archive_mode = on archive_command = 'wal-g wal-push %p' archive_timeout = 60 max_wal_senders = 10 hot_standby = on
4. Telemetry and Logging
We stripped out AWS CloudWatch entirely. In its place, we deployed lightweight Vector agents shipping logs to a self-hosted VictoriaMetrics and Quickwit instance running on an auxiliary VPS. This setup handled 50GB of log ingestion daily for under $30/month in total compute costs.
The Migration Playbook: Executing Without Downtime
We executed the migration across four phased sprints over three weeks.
Step 1: Infrastructure Provisioning with Ansible
We used Ansible to provision the Hetzner instances, set up OS-level firewall rules (ufw), configure sysctl network kernel tuning, and join the nodes to the K3s cluster.
Here is a snippet from our K3s control-plane bootstrap playbook:
- name: Install K3s Master Node
hosts: k3s_primary
become: true
tasks:
- name: Download and execute K3s installer
shell: |
curl -sfL https://get.k3s.io | K3S_TOKEN="{{ k3s_cluster_secret }}" sh -s - server \
--disable servicelb \
--disable traefik \
--flannel-backend=wireguard-native \
--node-ip={{ private_ip }} \
--tls-san={{ public_ip }}
args:
creates: /usr/local/bin/k3s
- name: Ensure K3s service is active and running
systemd:
name: k3s
state: started
enabled: yes
Step 2: Database Data Synchronization
Migrating PostgreSQL out of RDS required zero-downtime logical replication:
- Created a read replica on the target Hetzner Postgres instance using
pg_dumpsnapshot baseline. - Enabled logical replication from AWS RDS (publisher) to Hetzner Postgres (subscriber) to stream delta updates in real-time.
- Verified sequence sync and replication lag using
pg_stat_replicationuntil lag stayed consistently at 0ms.
Step 3: Traffic Shadowing and Dual-Writing
Before switching primary DNS, we deployed the microservices to K3s and routed 10% of incoming production webhooks to both environments using Cloudflare Workers. We logged HTTP response statuses, payloads, and execution timings to verify that the K3s environment produced identical outputs to the AWS cluster.
Step 4: The Cutover
At 02:00 WAT on a Sunday, during low transaction volumes:
- Placed API gateway in maintenance mode for 45 seconds.
- Promoted the Hetzner PostgreSQL subscriber to primary (
pg_replication_origin_advance). - Switched Cloudflare DNS records to point ingress to the Hetzner cluster IP endpoints.
- Re-enabled live traffic.
Total service interruption during cutover was exactly 38 seconds.
Results and Cost Metrics
Three months post-migration, the metrics demonstrate massive operational improvements along with cost reduction.
Monthly Expenditure Comparison
| Item | AWS Infrastructure (USD) | Hetzner + Cloudflare Setup (USD) | Savings | | :--- | :--- | :--- | :--- | | Compute Nodes | $361.00 (EKS + EC2) | $104.00 (4x Hetzner CX42) | 71% | | Database | $745.00 (RDS PostgreSQL) | $75.00 (Bare-Metal AX52) | 90% | | Caching | $196.00 (ElastiCache) | $0.00 (In-cluster Valkey/Redis) | 100% | | Networking / Egress | $1,120.00 (NAT + Transfer) | $0.00 (20TB Free Hetzner Egress) | 100% | | Log Management | $480.00 (CloudWatch) | $28.00 (VictoriaMetrics VPS) | 94% | | Backups & Security | $1,300.00 (S3, ECR, GuardDuty) | $137.00 (Cloudflare + Backblaze B2) | 89% | | Total Monthly Spend | $4,202.00 | $1,344.00 | 68% Net Savings |
In Naira terms, despite further local currency depreciation, monthly infrastructure costs dropped from over ₦6.3 million down to approximately ₦2.0 million, placing the platform back on sustainable unit economics.
Performance Metrics
- API P99 Latency: Dropped from 280ms (AWS Ireland) to 195ms (Hetzner Frankfurt) for West African end-users due to direct Cloudflare Anycast peering via Lagos IXPN (Internet Exchange Point of Nigeria).
- Deployment Build Speed: GitHub Actions deployment pipelines now build and push container images to a self-hosted GitHub Container Registry (GHCR) mirror in 2 minutes 10 seconds, down from 5 minutes 45 seconds.
- Uptime: Sustained 99.99% uptime over the trailing 90 days with zero unplanned outages.
Practical Rules for VPS Migrations in Emerging Markets
If you are planning to migrate away from managed hyper scalers to cut costs, keep these engineering principles in mind:
- Do Not Skip Egress Limits: VPS providers like Hetzner offer generous egress allowances (20TB per instance), but overages can surprise you. Enforce bandwidth caps and track traffic metrics using Prometheus export metrics.
- Never Expose Storage or Databases Directly: Managed cloud hides security misconfigurations behind IAM policies. When using bare VPS servers, lock down ports with
iptablesorufwso that database ports (5432,6379) listen only on private WireGuard interfaces. - Test Restore Pipelines Weekly: Without RDS automated backups, your backup pipeline is only as good as your last restore test. Write an automated cron job that downloads a
WAL-Gbackup once a week, spins up a temporary staging container, restores the dataset, and runs integrity assertions. - Keep Containers Small: When running on lighter compute nodes, base image size matters. Multi-stage Alpine or Distroless Docker builds reduce node disk pressure and make rolling restarts instantaneous when network bandwidth dips.
Neobot Engineering Standard
Every system deployed by Neobot Tech incorporates enterprise baseline practices. We continuously audit our database topologies, REST API query paths, and frontend modular bundles to prevent latency spikes and ensure top-tier security posture.
Discussion
Comments Coming Soon
We are currently migrating our discussion engine to a new real-time database schema. Check back shortly to join the conversation.