AI/ML

WhatsApp AI Agents Under Meta's 20-Second Timeout: Asynchronous Task Queues with FastAPI, Redis, and Pidgin Intent Routing

A
Adebayo FalojuPrincipal Systems Architect
October 8, 202618 min read
WhatsApp AI Agents Under Meta's 20-Second Timeout: Asynchronous Task Queues with FastAPI, Redis, and Pidgin Intent Routing

When deploying LLM agents on WhatsApp in Nigeria, slow tool-calling and high inference latency trigger Meta's webhook retries, resulting in duplicate messages and ballooning API costs. Here is how we decoupled webhook ingestion with Redis Streams, FastAPI, and lightweight Pidgin intent routing.

A retail client built a WhatsApp customer support agent powered by GPT-4o and custom tool-calling endpoints. Within three hours of public rollout in Lagos, their engineering lead called us with two critical problems: customers were receiving the same reply three or four times, and OpenAI billing was ticking up at ₦80,000 per hour.

The root cause was straightforward. Meta's WhatsApp Cloud API enforces a strict HTTP response timeout on webhooks. If your endpoint does not return a 200 OK within 20 seconds (and ideally under 3 seconds to stay safe against network jitter on African transit routes), Meta assumes delivery failed. Meta then initiates backoff retries, hitting your webhook again with the exact same payload.

Because the team ran their tool calls—fetching order statuses, checking account balances, and hitting LLMs—synchronously inside the FastAPI handler, requests took 12 to 18 seconds. When an LLM model stalled or tool execution added latency, the processing crossed the 20-second threshold. FastAPI had not responded yet, Meta retried the message, and a second thread started processing the duplicate message. The system spiraled into a cascade of duplicated API runs, race conditions, and confused customers.

Here is how we solved this in production: we decoupled Meta's webhook delivery from agent execution using an asynchronous event queue in Redis Streams, implemented deterministic idempotency, and placed a fast local Pidgin intent router in front of our heavy LLM calls.


The Architecture: Decoupling Ingestion from Execution

To keep webhook delivery under 100 milliseconds, your webhook handler must do exactly three things and nothing more:

  1. Verify Meta's SHA-256 HMAC signature.
  2. Deduplicate the incoming wamid (WhatsApp Message ID) in Redis.
  3. Push the payload into a Redis Stream and return an immediate HTTP 200 OK with an empty JSON body.

All heavy lifting—language normalization, retrieval-augmented generation (RAG), tool calling, and response generation—happens downstream in worker processes. When a worker completes a response, it pushes the message to Meta's /messages REST endpoint asynchronously.

WhatsApp Webhook Async Architecture Diagram

In our benchmarks across local infrastructure choices, as discussed in our analysis of Self-Hosted vLLM vs. Groq vs. Together AI: Token Costs, P99 Latency, and FX Risk for Nigerian AI Workloads, inference latency alone can fluctuate by several seconds depending on token count and provider load. You simply cannot tie HTTP webhook responses to live inference loops.


Step 1: FastAPI Webhook Receiver with Strict Idempotency

Meta sends every incoming user message with a unique ID inside entry[0].changes[0].value.messages[0].id formatted as wamid.HBgM.... We use this ID as a distributed lock key in Redis.

If Meta retries a request because of a transient network drops before our ACK reached their servers, our service checks Redis for the wamid. If it exists, we immediately return 200 OK without re-queuing the task.

Here is the production-tested FastAPI ingestion endpoint using redis-py and standard Python hashlib for signature verification.

import hmac
import hashlib
import os
import json
import redis.asyncio as aioredis
from fastapi import FastAPI, Request, HTTPException, status, Response

app = FastAPI()
REDIS_URL = os.getenv("REDIS_URL", "redis://localhost:6379/0")
APP_SECRET = os.getenv("WHATSAPP_APP_SECRET", "your_app_secret_here")

redis_client = aioredis.from_url(REDIS_URL, decode_responses=True)

def verify_signature(payload: bytes, signature: str) -> bool:
    if not signature or not signature.startswith("sha256="):
        return False
    expected_sig = hmac.new(
        APP_SECRET.encode("utf-8"),
        msg=payload,
        digestmod=hashlib.sha256
    ).hexdigest()
    return hmac.compare_digest(f"sha256={expected_sig}", signature)

@app.post("/webhook/whatsapp")
async def whatsapp_webhook(request: Request):
    raw_body = await request.body()
    signature = request.headers.get("X-Hub-Signature-256", "")
    
    if not verify_signature(raw_body, signature):
        raise HTTPException(status_code=403, detail="Invalid HMAC signature")
    
    payload = json.loads(raw_body.decode("utf-8"))
    
    # Extract message details from payload safely
    try:
        entry = payload["entry"][0]["changes"][0]["value"]
        messages = entry.get("messages", [])
        if not messages:
            # Status updates (read receipts, delivery receipts) land here
            return Response(content="OK", status_code=200)
        
        message = messages[0]
        wamid = message["id"]
        from_number = message["from"]
    except (KeyError, IndexingError):
        return Response(content="OK", status_code=200)

    # Idempotency check: Set NX with 24-hour expiration
    is_new = await redis_client.set(f"idempotency:{wamid}", "processing", nx=True, ex=86400)
    if not is_new:
        # Duplicate delivery attempt from Meta. Acknowledge without re-processing.
        return Response(content="OK", status_code=200)

    # Enqueue work into Redis Stream
    event_data = {
        "wamid": wamid,
        "phone": from_number,
        "payload": json.dumps(message)
    }
    await redis_client.xadd("whatsapp_incoming_stream", event_data)
    
    # Fast ACK under 50ms
    return Response(content="OK", status_code=200)

Step 2: Routing Local Pidgin and Code-Switched Queries Fast

Many Nigerian WhatsApp users blend Nigerian Pidgin, Yoruba, and English in a single message (e.g., "I sharp sharp transfer ₦15k since morning, my account never credit, wetin dey sup?").

If you send every single text straight to GPT-4o or Claude 3.5, you burn unnecessary input tokens on conversational filler and run into high P99 latency. Worse, standard base models often struggle with local financial context unless specifically prompt-engineered with extensive few-shot examples.

Instead, we run incoming messages through a multi-tier local intent router before hitting any external LLM service:

  1. Exact/Regex Match: Checks for deterministic standard keywords ("balance", "statement", "talk to human").
  2. Pidgin Normalization & Lightweight Classifier: Uses a fast set of regular expressions and keyword maps tuned for local phrasing ("wetin dey sup", "I wan buy", "abeg fast", "money hanging").
  3. LLM Agent Fallback: Handles complex, multi-step queries that require contextual reasoning and external tool calls.
import re
from typing import Tuple, Dict, Any

PIDGIN_DICTIONARY = {
    r"\b(wetin dey sup|how far|how body)\b": "greeting",
    r"\b(i wan check|show me|wetin be) (my balance|acct balance|money)\b": "check_balance",
    r"\b(money never enter|hanging|debit error|deduct my money)\b": "failed_transaction",
    r"\b(i wan talk to|connect me to) (human|agent|person)\b": "agent_escalation",
}

class LocalIntentRouter:
    def __init__(self):
        self.patterns = {
            intent: re.compile(pattern, re.IGNORECASE) 
            for pattern, intent in PIDGIN_DICTIONARY.items()
        }

    def route(self, text: str) -> Tuple[str, float]:
        """
        Returns (intent_name, confidence_score).
        Confidence is 1.0 for exact local pattern matches.
        """
        cleaned_text = text.strip().lower()
        
        for pattern, intent in self.patterns.items():
            if pattern.search(cleaned_text):
                return intent, 1.0
                
        return "llm_fallback", 0.0

By intercepting routine balance checks or transactional status queries with deterministic code, you resolve roughly 40% of incoming customer interactions in under 100 milliseconds without incurring LLM charges.


Step 3: Asynchronous Stream Worker Execution

The background process consumes from the Redis Stream, uses the intent router, processes tool calls asynchronously, and dispatches the final response back to Meta's WhatsApp API via HTTP client (httpx).

import os
import json
import asyncio
import httpx
import redis.asyncio as aioredis
from intent_router import LocalIntentRouter

REDIS_URL = os.getenv("REDIS_URL", "redis://localhost:6379/0")
WHATSAPP_TOKEN = os.getenv("WHATSAPP_TOKEN")
PHONE_NUMBER_ID = os.getenv("WHATSAPP_PHONE_NUMBER_ID")

router = LocalIntentRouter()

async def send_whatsapp_message(to_phone: str, text_body: str):
    url = f"https://graph.facebook.com/v19.0/{PHONE_NUMBER_ID}/messages"
    headers = {
        "Authorization": f"Bearer {WHATSAPP_TOKEN}",
        "Content-Type": "application/json",
    }
    payload = {
        "messaging_product": "whatsapp",
        "recipient_type": "individual",
        "to": to_phone,
        "type": "text",
        "text": {"preview_url": False, "body": text_body}
    }
    async with httpx.AsyncClient(timeout=10.0) as client:
        resp = await client.post(url, headers=headers, json=payload)
        resp.raise_for_status()

async def process_event(redis_client, message_id, data):
    phone = data["phone"]
    message_payload = json.loads(data["payload"])
    text = message_payload.get("text", {}).get("body", "")

    if not text:
        # Skip media or unhandled message types for this example
        return

    # 1. Fast Intent Check
    intent, confidence = router.route(text)
    
    if intent == "greeting":
        response_text = "How far! How I fit help you today? You fit check your balance or ask about your last transaction."
    elif intent == "check_balance":
        # Local DB call for balance
        response_text = "Your current balance na ₦45,200.00."
    elif intent == "failed_transaction":
        response_text = "E pele. If your transfer hang, abeg wait 15 minutes. Bank network dey small delay today."
    else:
        # 2. Heavy LLM Invocation with Tool Calling (Fallback)
        response_text = await run_llm_agent(phone, text)

    # 3. Dispatch reply to user
    await send_whatsapp_message(phone, response_text)

async def run_llm_agent(phone: str, user_text: str) -> str:
    # Call OpenAI / Anthropic / Local vLLM instance here
    # Simulate tool calls and retrieval
    await asyncio.sleep(2.0) # Simulated LLM processing latency
    return f"I have processed your request regarding: '{user_text}'. Is there anything else I can assist with?"

async def start_worker():
    redis = aioredis.from_url(REDIS_URL, decode_responses=True)
    group_name = "whatsapp_workers"
    consumer_name = "worker_node_1"

    try:
        await redis.xgroup_create("whatsapp_incoming_stream", group_name, id="0", mkstream=True)
    except Exception:
        pass # Consumer group already exists

    print("Worker connected. Polling Redis Stream...")
    while True:
        try:
            entries = await redis.xreadgroup(
                groupname=group_name,
                consumername=consumer_name,
                streams={"whatsapp_incoming_stream": ">"},
                count=10,
                block=2000
            )
            for stream_name, messages in entries:
                for msg_id, data in messages:
                    try:
                        await process_event(redis, msg_id, data)
                        await redis.xack("whatsapp_incoming_stream", group_name, msg_id)
                    except Exception as e:
                        print(f"Failed to process message {msg_id}: {e}")
        except Exception as e:
            print(f"Stream reading error: {e}")
            await asyncio.sleep(1)

if __name__ == "__main__":
    asyncio.run(start_worker())

Comparison: Synchronous Webhook vs. Asynchronous Stream Architecture

The operational differences between processing requests inline versus using asynchronous streams are stark, especially under latency spikes or high request volume.

| Operational Metric | Synchronous Webhook Architecture | Asynchronous Redis Stream Architecture | | :--- | :--- | :--- | | P99 Webhook Response Time | 8,000ms – 22,000ms | 15ms – 45ms | | Meta Delivery Retries | Frequent (Triggers double runs) | Near-zero (Handled by strict idempotency ACK) | | Max Concurrent Users | Bottlenecked by web server worker count | Scalable across distributed Redis consumers | | LLM Spend Efficiency | High (Duplicate requests rerun models) | Optimized (Local routing drops costs up to 40%) | | Flaky Network Resiliency | Poor (Dropped requests fail silently) | High (Queue retains work until acknowledged) |

For mobile-first applications deployed in environments with flaky connections, maintaining queue durability is critical. If your application also supports offline functionality on the frontend, check out our guide on Offline-First React Native: Building an Idempotent SQLite Mutation Queue for Flaky 2G Networks.


Common Pitfalls

1. In-Memory State Retention on Workers

Do not store conversation history in-memory on the worker process. When running multiple consumer nodes behind Redis Streams, a customer's subsequent message will likely be processed by a different worker instance. Store session histories in Redis using key formats like chat_history:{phone} or a centralized Postgres instance.

2. Failing to Handle WhatsApp Out-of-Order Delivery

If a user sends three short messages in rapid succession ("Hello", "My money never drop", "Abeg check am"), Meta might send webhooks slightly out of sequence over distinct TCP connections. Ensure your background process sorts or buffers incoming user messages by timestamp per user (phone) if strict sequence order matters for your retrieval context.

3. Ignoring Meta Rate Limits

WhatsApp enforces tier limits on outgoing business messages (e.g., 80 messages per second per phone number). If your consumer queue attempts to push thousands of replies at once following a local network recovery, Meta will respond with HTTP 429 Too Many Requests. Wrap your outgoing HTTP client calls with a token bucket rate limiter.

4. Over-engineering Simple Verification Checks

When setting up the webhook endpoint in the Meta Developer Portal, Meta sends a GET request with a hub.challenge query parameter. Do not wrap this verification endpoint in heavy authentication middlewares or signature validation. Match hub.verify_token against your environment variable and return the raw hub.challenge plain text string instantly.


Frequently Asked Questions

What happens if the worker process crashes while processing a message?

Because we use Redis Streams consumer groups with explicit XACK confirmation, any unacknowledged message remains in the Pending Entries List (PEL). If a worker process dies unexpectedly, another worker can inspect the PEL via XPENDING and claim the abandoned task using XCLAIM after a specified timeout (e.g., 60 seconds).

Why use Redis Streams instead of Celery or RabbitMQ?

Redis Streams gives you zero additional infrastructure overhead if you are already using Redis for caching or idempotency key storage. It supports consumer groups, message replay, and offset tracking out of the box with negligible resource footprint on modest cloud instances.

How do we handle voice notes in Nigerian Pidgin?

Meta sends WhatsApp voice notes as media objects containing a media_id. Your worker downloads the audio binary using Meta's media API, routes it through an optimized speech-to-text model (like OpenAI Whisper or a local model hosted via vLLM/Triton), and passes the transcribed text into the intent router. Never process media downloads directly inside your FastAPI webhook handler.

How do we verify webhook security beyond signature checking?

Besides HMAC-SHA256 signature verification, restrict inbound traffic to your webhook route at the load balancer or security group level using Meta's publicly published IPv4 and IPv6 address ranges. Additionally, for financial transactions triggered via chat agents, always require out-of-band verification or device binding, as detailed in our technical breakdown on Bypassing SMS OTP: Implementing WebAuthn Device Binding and Telco SIM-Swap Checks in Node.js.


To keep WhatsApp AI agents stable in production, never let external LLM latencies block web servers. Decouple message ingestion from processing using asynchronous task queues, implement strict idempotency on wamid values, and route routine queries through fast local rule engines before incurring LLM charges.

Neobot Engineering Standard

Every system deployed by Neobot Tech incorporates enterprise baseline practices. We continuously audit our database topologies, REST API query paths, and frontend modular bundles to prevent latency spikes and ensure top-tier security posture.

Tags:#AI/ML#WhatsApp Cloud API#FastAPI#Redis#LLM#Python

Discussion

Comments Coming Soon

We are currently migrating our discussion engine to a new real-time database schema. Check back shortly to join the conversation.