Smart Payment Routing in Go: Recovering ₦140M in Dropped Card Transactions
When single-processor payment infrastructure failed during peak evening traffic in Lagos, card success rates plunged below 70%. Here is how Neobot Tech designed an open-source payment routing engine in Go with sliding-window circuit breakers, recovering over ₦140M in dropped card transactions.
If you process card transactions in Nigeria, you already know the painful truth: single-processor architecture is a ticking time bomb. Every evening between 5:00 PM and 8:30 PM West Africa Time, transaction success rates across major payment gateways plummet. It isn't always the processor's fault—often, the underlying switch operated by NIBSS or an upstream issuing bank's core banking application is throwing 504 Gateway Timeouts or silent drops.
When a user taps 'Pay' and sees a spinning wheel for 45 seconds followed by 'Transaction Failed', two things happen. First, you lose immediate revenue. Second, thirty seconds later, the user receives an SMS debit alert from their bank anyway. Now your support team is buried under tickets while your engineering team struggles to figure out if the transaction failed at the processor, the switch, or the issuer.
Last year, Neobot Tech engaged with a high-volume mobility and digital wallet platform processing over ₦1.2 billion monthly across Lagos and Abuja. During peak commuting hours, their card charge success rate crashed to 68.4%. By re-architecting their payment engine with a dynamic multi-acquirer router written in Go, we brought that success rate up to 91.2% and recovered ₦140M+ in previously lost monthly volume.
Here is how we built it, the architectural traps we fell into, and the playbook you can use for your own infrastructure.
The Problem: The High Cost of Static Acquirer Binding
When we audited the platform's legacy backend, the payment implementation looked like what 90% of Nigerian fintech startups deploy: a single wrapper around Paystack's API. For standard e-commerce, this is fine. For a high-throughput on-demand service, it is a single point of failure.
When Paystack experienced temporary degrading on GTBank Mastercard routes or Zenith Visa authorization switches, the application kept mindlessly throwing 100% of its traffic at that broken pipe. Worse, when an API call timed out after 30 seconds, the client application simply displayed a failure banner and prompted the user to try again.
Desperate users would tap 'Try Again' three times. When the gateway network finally recovered ten minutes later, three separate pending hold charges would clear simultaneously, debiting the user three times for a single ride.
We faced four hard constraints while refactoring this system:
- Sub-500ms Routing Overhead: The routing decision had to happen inline before presenting the payment page or initiating the charge token without adding perceived UI lag.
- Card Token Heterogeneity: Paystack reusable authorization codes (
AUTH_xxxx), Flutterwave card tokens (flw-t12-xxxx), and Squad card tokens cannot be used interchangeably. A card tokenized on Processor A cannot be charged directly on Processor B. - Strict Zero Double-Charge Guarantee: We could never blindly fall back to a second processor if the status on the primary processor was indeterminate (
pendingortimeout). - CBN Settlement and Reconciliation Compliance: Every transaction attempt across every processor had to map back to a unified internal ledger entry for daily CBN reporting.
What Didn't Work: Naive Retries and Client-Side Fallbacks
Before building a dynamic server-side router, we tested two simpler approaches. Both failed in production.
Failed Approach 1: Client-Side Fallback JS SDKs
We initially attempted to catch processor failures inside the mobile web view and re-initialize the modal using a secondary gateway (Flutterwave).
This was a disaster. If a user's network connection dropped midway through the first gateway's webview handoff, the app assumed failure and triggered the second modal. The user ended up paying on Gateway B, while Gateway A's webhook fired 4 minutes later confirming the original charge. Resolving these double-debits required manual refund requests across multiple dashboard interfaces, burning dozens of engineering hours.
Failed Approach 2: Static Priority Failover
We next tried a server-side static rule: Route to Paystack first; if Paystack returns an HTTP status code >= 500, failover to Monnify or Squad.
This failed because Nigerian gateway outages rarely manifest as neat HTTP 500 errors. Instead, processors return HTTP 200 OK with a body indicating status: false or `gateway_response:
Neobot Engineering Standard
Every system deployed by Neobot Tech incorporates enterprise baseline practices. We continuously audit our database topologies, REST API query paths, and frontend modular bundles to prevent latency spikes and ensure top-tier security posture.
Discussion
Comments Coming Soon
We are currently migrating our discussion engine to a new real-time database schema. Check back shortly to join the conversation.