Companion to smtp-egress-paths.html (current state). Goal: one enforcement point — the HOU gate — so the kill switch, allowlist, and engine filters physically live in one place. Remote sites keep only a dumb classifier ("is this port 25?") and a pipe to HOU.
Policy is already centralized — enforcement is not. All 4 gates read one Redis at HOU (23.165.104.28): the kill switch, allowlist, limits, and reputation are a single brain today. What is distributed is the data plane: 4 gate VMs doing the dropping, and 4 routers each carrying a 4-term SMTP-INTERCEPT filter whose T2 (DSCP af11) term is a proven bypass hole (#561) at every site. Centralizing the data plane buys: one gate to patch/audit, T2 deleted at 3 sites (hole closed), 3 VMs decommissioned, and BOS/DC come up safe with zero new gate builds. The price: a backhaul path, and HOU becomes a fleet-wide blast radius that must be engineered fail-closed.
Example shown for a LAX customer; KC and MCO are identical. HOU-local customers keep today's exact path.
Customer VM (LAX, e.g. 108.165.12.202)
│ :25 SYN, dst = external MX
▼
LAX-EDGE SMTP-INTERCEPT on customer IFLs — now only 3 terms:
T1 WHITELIST (Postal 23.165.104.147, infra IPs) → accept
T3 SMTP-REDIRECT tcp dst-port 25 → instance smtp-redirect
T4 DEFAULT → accept
T2 (dscp af11 accept) DELETED — #561 self-mark hole closed here
│
│ routing-instance smtp-redirect:
│ 0/0 → gr-0/0/x.25 (GRE → HOU-EDGE lo0) ← primary
│ 0/0 → discard, preference 250 ← tunnel down = FAIL-CLOSED
▼
═══════ GRE tunnel (or backbone-wave L3VPN later) · LAX→HOU ~35 ms ═══════
▼
HOU-EDGE gr- unit terminates INSIDE its smtp-redirect instance
│ 0/0 → 136.0.41.253 (gate) + qualified-next-hop gate-B (HA)
│ 0/0 → discard, preference 250 ← gate dead = FAIL-CLOSED
▼
SMTP-HOU gate 136.0.41.253 — THE gate. Redis is now LAN-local.
├─ blocked / ratelimited ──► DROP / REJECT in-kernel
├─ not allowlisted ──► @killswitch {25} ──► DROP ← today's default
└─ allowlisted ──► nfqueue → engine verdict
│ PERMIT: dscp af11, back out gw 136.0.41.1 = HOU-EDGE
▼
re-enters HOU's filtered IFL → T2 accept (af11) ← T2 survives ONLY
│ on HOU (or die too
▼ via plan-C ded. IFL)
HOU transit ──► internet MX
Return path: MX → SYN-ACK → dst 108.165.12.202 → normal BGP → LAX → VM directly.
Asymmetric on purpose — the gate only ever sees the outbound leg, exactly like today
(SMTP-INTERCEPT matches dst-port 25 on customer IFLs; return traffic is src-port 25
arriving on transit IFLs and is never diverted). No new statefulness anywhere.
Two ways to handle permitted traffic: (A) egress HOU transit directly — chosen above — or (B) tunnel the verdict back to the origin site and egress locally. (B) preserves egress locality but doubles backhaul, needs per-site return tunnels with policy routing on the gate, and forces every remote edge to keep a T2-equivalent "accept returning gate traffic" term — i.e. the #561 hole stays alive fleet-wide. (A) deletes T2 at 3 sites and is trivial. Cost of (A): LAX/KC/MCO-sourced packets exit via HOU transit — fine within one ASN, but verify HOU upstreams accept the whole DartNode aggregate as source (IRR AS-SET prefix filters / uRPF mode on the HOU transit sessions). Latency +15–35 ms on outbound SMTP is irrelevant.
| Component | Verdict | Detail |
|---|---|---|
| Shared Redis @ HOU | STAYS | Already the single policy brain (23.165.104.28). Becomes LAN-local to the only gate — one less WAN dependency, not one more. |
| SMTP-HOU gate VM | STAYS | Becomes THE fleet gate. Its customer resolution is src-IP → Redis and site-agnostic, so LAX/KC/MCO source IPs resolve identically — likely zero daemon changes; verify with one test flow at cutover. |
| Vader admin, SmtpGateService, allowlist ticket flow | STAYS | Zero PHP changes. Dashboard's fleet view simply shows 1 node instead of 4 (heartbeats). |
| Postal path (T1) | STAYS | Transactional mail keeps its per-edge whitelist term; never touches the gate. |
| Per-edge SMTP-INTERCEPT filter | STAYS shrinks | The classifier stays distributed (it must — it's where customer packets enter). But it loses T2 and its FBF next-hop stops being site-local. |
| FBF next-hop @ LAX/KC/MCO | CHANGES | smtp-redirect 0/0: local gate IP → GRE unit to HOU, plus 0/0 discard backup (tunnel down must blackhole :25, never fall open — the Jul/Aug SBL lesson). |
| T2 DSCP term @ LAX/KC/MCO | DELETED | Nothing legitimate re-enters those filters marked af11 anymore → the #561 self-mark bypass dies at 3 of 4 sites. (HOU keeps T2 until plan-C moves the gate to a dedicated unfiltered IFL.) |
| SMTP-LAX / SMTP-KC / SMTP-MCO VMs | DECOMMISSIONED | After per-site soak. Optionally keep one as a cold spare image. |
| GRE tunnels ×3 (→ ×5 with BOS/DC) | ADDED | Edge lo0 ↔ HOU-EDGE lo0, terminated inside the smtp-redirect instances. MX204 needs tunnel-services enabled (reserves PFE bandwidth). GRE over public internet works even where backbone waves are down (kc-bos is hard down today); migrate onto a backbone L3VPN VRF later if wanted — same FBF design, different transport. |
| MSS clamp for :25 into tunnels | ADDED | GRE eats ~24 bytes; clamp TCP MSS on the gr- units (or run jumbo on the underlay) so mail with full-size segments doesn't blackhole. |
| 2nd HOU gate (gate-B) | ADDED | Qualified-next-hop in HOU's smtp-redirect instance. Without it one VM reboot pauses fleet-wide allowlisted mail (killswitch'd traffic just drops, which is policy anyway). |
| Tunnel / FBF-route monitoring | ADDED | Alarm when any smtp-redirect instance falls to its discard backup, plus GRE keepalives. The existing smtpgate:node heartbeat covers the gate itself. |
nfqueue flags bypass on the gate | MUST CHANGE | The #26 fail-open (gate process dies → :25 accepted) was a one-site risk; centralized it becomes fleet-wide. Remove the bypass flag (fail-closed) as part of this project — it stops being a deferrable card. |
| Failure | Today (4 gates) | Centralized (design above) |
|---|---|---|
| Gate process dies | that site fails OPEN (nfqueue bypass flag, #26) | gate-B takes over; if both die, HOU discard backup → fleet fail-CLOSED (requires removing bypass flag) |
| Backhaul tunnel down | n/a | that site's :25 → discard = fail-closed, alarm fires; customers see what the kill switch already shows them today |
| Redis down | gates keep last-known snapshot | same behavior, but Redis is now LAN-local to the only consumer — strictly better |
| HOU site fully down | HOU customers only | fleet-wide :25 outage (closed) — acceptable while global kill switch is ON anyway; revisit if SMTP ever reopens broadly |
| VM self-marks af11 (#561) | bypass works at all 4 sites | dead at LAX/KC/MCO; HOU-only residual until plan-C dedicated gate IFL |
BOS/DC-EDGE: replace the cloned SMTP-INTERCEPT's dead donor next-hop with the
SAME pattern as every other remote site:
T1 whitelist → T3 redirect → T4 accept (no T2, no local gate, no new VM)
smtp-redirect: 0/0 → gr- to HOU · 0/0 discard backup
No gate build, no fail-open day-one risk — the cloned-filter hazard flagged in the
current-state page becomes a 5-line config fix instead of a new site deployment.
tunnel-services on all MX204s; the 2-command IFL check at LAX/KC/MCO (already flagged in the current-state page).