DartNode · smtp-gate topology · 2026-09-29

SMTP egress path per PoP — customer VM → internet

Pre-implementation picture. HOU is fully verified (including the 09-28 loop repro); LAX/KC/MCO are verified at the gate level; BOS/DC edge state comes from the card #609 clone work. Nothing here changes anything.

Legend / shared building blocks

Every live site uses the same 5-stage design; only the wiring differs per site.

[VM]      customer VM (qemu on a PVE node)
[BR]      Proxmox bridge (vmbr0) — pure L2, no filtering
[EDGE]    site MX204, customer-facing IFL carries input filter SMTP-INTERCEPT:
            T1 WHITELIST      src/dst = gate, Postal 23.165.104.147, 149.112.84.2-5 → accept
            T2 GATE-FORWARDED dscp af11 (0x0a)                                     → accept   ← the #561 hole
            T3 SMTP-REDIRECT  tcp dst-port 25 → routing-instance smtp-redirect (FBF → gate)
            T4 DEFAULT        accept
[GATE]    smtp-gate VM (transparent, one-armed hairpin; dst IP never rewritten):
            nft forward: blocked_customers DROP → ratelimited REJECT
                         → @global_allow → nfqueue(daemon verdict)
                         → dport @killswitch(={25}) DROP        ← kill switch, ON today
                         → dport 25 → nfqueue
            on ACCEPT: nft marks packet dscp af11, routes back out its default gw = EDGE
[TRANSIT] site upstream → internet

HOU fully verified — incl. 09-28 loop · the plan-C site

 Customer VM (e.g. HOU-GOLD-3, 38.134.40.20)
      │ :25 SYN, dst = external MX
      ▼
 PVE node vmbr0 ──L2──► HOU-EDGE et-0/0/1.0  "HOU-40-1 Trunk"
                        (ONE untagged unit, ~90 subnets INCLUDING the
                         gate's 136.0.41.0/24 — filter + gate share this IFL)
                        input filter SMTP-INTERCEPT
      │ T3 match: dst-port 25 → FBF routing-instance smtp-redirect
      │           (static 0/0 next-hop 136.0.41.253; L2 rewrite only)
      ▼
 SMTP-HOU gate 136.0.41.253  (v6 twin: 2602:f9f3:0:2::3df)
      │
      ├─ not allowlisted ──► @killswitch {25} ──► DROP in-kernel   ← today: ~all customers
      │
      ├─ allowlisted (7 blocks in @global_allow) ──► nfqueue → engine verdict
      │       │ PERMIT: nft sets dscp af11, forwards unchanged
      │       ▼
      │  back out gate gw 136.0.41.1 = HOU-EDGE, re-enters the same
      │  filtered IFL et-0/0/1.0 → hits T2 GATE-FORWARDED (af11) → accept
      │       │                      ▲
      │       │                      └── T2 is LOAD-BEARING here: delete it and this
      │       │                          leg matches T3 again → edge⇄gate ping-pong
      │       │                          until TTL death (proven 09-28)
      │       ▼
      │  HOU transit (23.26.125.1 → 216.227.254.249 → …) ──► internet MX
      │
      └─ ANY VM self-marking af11 ──► matches T2 at the edge, never
         reaches the gate ──► transit ──► internet        ← the #561 bypass hole
                                                            (re-proven live 09-24)

 Exempt flows (T1): VM → Postal 23.165.104.147:25 goes straight through (transactional mail).

LAX verified live, v4 + v6IFL sharing unverified

 Customer VM (e.g. 108.165.12.202 / 2602:f9f3:3000::108)
      ▼
 LAX-EDGE (MAC 78:4f:9b:4b:48:0f), SMTP-INTERCEPT on customer IFL(s)
      │ FBF divert (v4 AND v6 — v6 divert proven live here)
      ▼
 SMTP-LAX gate 108.165.12.15, gw 108.165.12.1
      ├─ killswitch DROP (today's default)
      └─ allowlisted PERMIT → af11 → back via LAX-EDGE T2 → Cogent (et-0/0/2) → internet

 Same T2 self-mark hole as HOU. Whether the gate shares the filtered IFL is
 UNVERIFIED at LAX — assume HOU-like (loop risk) until the arp/interface check is run.

KC divert verified at L2unfiltered-IFL theory unverified

This is where the FBF mechanism was reverse-engineered.

 Customer VM (e.g. 23.165.105.6 — note: different /24 than the gate)
      ▼
 KC-EDGE (MAC 78:4f:9b:4e:70:13, gw 108.165.120.1), SMTP-INTERCEPT input
      │ FBF: only L2 next-hop rewritten, dst IP untouched
      ▼
 SMTP-KC gate 108.165.120.91
      ├─ killswitch DROP  (tcpdump-proven: customer SYN retransmits same-seq, no reply)
      └─ allowlisted PERMIT → af11 → KC-EDGE → KC transit → internet

 09-18 inference said KC's gate may sit on an UNFILTERED IFL ("how KC avoids the
 hairpin") — if true, T2 deletion is clean at KC. Explicitly NOT verified; needs the
 same arp + show-interfaces check as HOU before touching anything.

MCO working since day oneleast-inspected edge

 Customer VM ──► MCO-EDGE (hostname wrongly says edge.ord) ──FBF──►
 SMTP-MCO gate 108.165.123.134
      ├─ killswitch DROP
      └─ allowlisted PERMIT → af11 → MCO-EDGE T2 → MCO transit → internet

 Same filter design assumed (fail-closed gate build deployed + enforcement observed);
 edge config never dumped — IFL layout unknown.

BOS & DC — today no gate exists · cloned filter risk · unverified

Today there is no SMTP path at BOS/DC — and that's not automatically safe. Only 4 gate boxes exist (HOU/LAX/KC/MCO). The edges at BOS/DC are live MX204s built as clones: BOS ≈ MCO clone, DC = full LAX clone (card #609). That means their configs very likely contain a cloned SMTP-INTERCEPT whose FBF next-hop points at the donor site's gate IP (108.165.123.134 / 108.165.12.15), which is unreachable from there:

 TODAY (if a customer VM ever lands there before this is fixed):
 VM ──► BOS/DC-EDGE ──SMTP-INTERCEPT(cloned)──► FBF 0/0 → donor-site gate IP
            │
            ├─ next-hop unresolved → Junos hides the 0/0 → FBF falls through
            │     → :25 goes STRAIGHT TO TRANSIT, ungated  (fail-open — a repeat
            │       of the Jul 29–Aug 10 SBL condition on day one)
            └─ or resolves via backbone wave → 25 hairpins cross-country into
                  another site's gate (wrong, and kc-bos wave is hard down anyway)

 Status: UNVERIFIED — needs `show configuration | display set | match SMTP` on both
 edges before either site takes customers. DC additionally has no egress of its own
 today (LAX clone stripped → wave-only, no owned space).

BOS & DC — target design greenfield · plan-C end-state

Greenfield = free chance to build the plan-C end-state that HOU can't retrofit:

 Customer VM
      ▼
 BOS/DC-EDGE  customer IFLs: SMTP-INTERCEPT with only
              T1 WHITELIST → T3 REDIRECT → T4 ACCEPT      ← NO DSCP term at all
      │ FBF divert
      ▼
 SMTP-BOS / SMTP-DC gate on its OWN dedicated unit/VLAN
      │              (gate subnet NOT on the customer trunk — the thing HOU got wrong)
      ├─ killswitch/engine DROP
      └─ PERMIT ──► returns to edge on the gate's UNFILTERED IFL ──► no re-divert
                    possible, no loop, no af11 needed ──► transit ──► internet
                    (BOS: own transit; DC: bos-dc wave → BOS until DC has real
                     allocations + its own uplink)

 Result at BOS/DC: self-mark hole doesn't exist, allowlist lives 100% in Redis/Vader,
 zero per-change router edits — the end-state plan C only approximates at HOU via
 prefix-list-sync or GRE.

Site-by-site summary

SiteEdgeGateDivertReturn leg (allowlisted)Self-mark holeGate on filtered IFL?
HOUHOU-EDGE, et-0/0/1.0136.0.41.253v4 live, v6 attached/unprobedvia T2 dscp af11 — load-bearingyes (proven)yes (verified) — why plan C exists
LAXLAX-EDGE108.165.12.15v4+v6 livevia T2yes (same design)unverified
KCKC-EDGE108.165.120.91v4 livevia T2yes (same design)inferred no — unverified
MCOMCO-EDGE108.165.123.134livevia T2yes (same design)unknown
BOSBOS-EDGE (MCO clone)nonecloned filter, dead next-hop?n/an/ato be built clean
DCDC-EDGE (LAX clone)nonecloned filter, dead next-hop?n/an/ato be built clean

Worth acting on regardless of which plan-C variant is picked

  1. The cloned SMTP-INTERCEPT state on BOS/DC edges should be checked and neutered before those sites take customers — it's a one-command check per edge: show configuration | display set | match SMTP.
  2. The "gate on filtered IFL?" column is the deciding fact per site for whether T2 deletion is safe, and it's only actually verified at HOU. KC/LAX/MCO each need the same two commands before repeating anything there: show arp no-resolve | match <gate-ip> + show configuration interfaces <ifl>.