The probe
The API gateway is a cross-cutting infrastructure component that every FAANG-tier system uses. The question tests whether you understand authentication, rate limiting, routing, and load balancing as composable middleware layers — and what the operational challenges are when this component is in the critical path of every API call.
Step 1 — Clarify
- What features: auth, rate limiting, routing, load balancing, caching, request transformation, observability?
- Scale: 1M req/sec aggregate across all upstream services
- Multi-tenant (different customers, different routing rules) or single-tenant? - SLO: gateway latency overhead < 5ms p99 (adds < 5ms to every request)
Step 2 — Architecture
The middleware chain (executed in order for every request):
1. TLS termination — decrypt HTTPS at the gateway; upstream services communicate in plaintext within the datacenter
2. Authentication — validate API key or JWT token. Token validation: symmetric JWT (verify signature with shared secret, ~0.5ms) or asymmetric (fetch public key from auth service — cache aggressively, 24h TTL)
3. Rate limiting — per API key, per endpoint, sliding window counter in Redis (see Rate Limiter walkthrough)
4. Request routing — match URL pattern to upstream service. Stored in a routing table (Redis or in-memory hash map, updated via config push)
5. Load balancing — pick a healthy upstream instance (round-robin, least-connections, or consistent hash for sticky sessions)
6. Request forwarding — proxy the request to the upstream; handle upstream timeout and retry
7. Response transformation — add CORS headers, strip internal headers, compress response
8. Observability — emit request log, latency metric, status code metric to monitoring pipeline
Configuration plane: Routing rules, rate limits, auth policies are stored in a config database (etcd or Postgres) and pushed to gateway instances via pub-sub. Gateways cache config in-memory; config staleness of 30 seconds is acceptable.
Health checking: Gateway maintains a health map of all upstream instances. Active health checks (HTTP probe every 5s) + passive health checks (circuit breaker on error rate). Unhealthy instances removed from the load balancing pool.

