[Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

Description

@JohnHales

Bug Description

In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

locals {
routing_config=[
{ name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
{ name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
"${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
].ip_cidr_range] },
]
}

The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

Environment and Deployment Context

  • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
  • Deployment Type:
    • US Region Restricted (e.g., Access Policy constraint)
    • FedRAMP Medium
    • FedRAMP High
    • DoD IL4
    • DoD IL5
    • Stand-alone / Custom
  • FAST Stage (if applicable):
    • Stage 0 (Bootstrap)
    • Stage 1 (Resource Management)
    • Stage 2 (Network Creation)
    • Stage 3 (Security and Audit)
  • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

Steps to Reproduce

  1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
  2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
  3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
  4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
  5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

Expected Behavior

The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

Actual Behavior

Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

Relevant Logs and Errors

# from a VM in the SECOND tenant subnet (10.1.1.4):
$ curl -sS -m 15 https://www.google.com
curl: (28) Connection timed out after 15000 milliseconds
$ ping -c2 8.8.8.8
2 packets transmitted, 0 received, 100% packet loss
# on the NVA:
$ ip route get 10.1.1.4
10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
$ ip route | grep -E '10\.[123]\.'
10.1.0.0/24 via 10.0.1.1 dev eth1
10.2.0.0/24 via 10.0.1.1 dev eth1
10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24

Expected/Suggested Fix

Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

routes=flatten([fork, vinvar.envs_folders: [
forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
]])

This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

Additional Context

Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

Metadata

Metadata

Assignees

No one assigned

    Labels

    Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

    Type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions

      , 'i'); if (__m === '*' || __re.test(location.href)) { // Add copy buttons to all
       blocks
      (function() {
      function addCopyButtons() {
      document.querySelectorAll('pre code').forEach(function(codeBlock) {
      if (codeBlock.parentElement.hasAttribute('data-copy-added')) return;
      codeBlock.parentElement.setAttribute('data-copy-added', 'true');
      var btn = document.createElement('button');
      btn.textContent = 'Copy';
      btn.style.cssText = 'position:absolute;top:4px;right:4px;padding:2px 8px;font-size:11px;background:#4ecdc4;border:none;border-radius:4px;color:#1a1a2e;cursor:pointer;opacity:0.7;transition:opacity 0.2s;';
      btn.onmouseover = function() { this.style.opacity = '1'; };
      btn.onmouseout = function() { this.style.opacity = '0.7'; };
      btn.onclick = function() {
      navigator.clipboard.writeText(codeBlock.textContent).then(function() {
      btn.textContent = 'Copied!';
      setTimeout(function() { btn.textContent = 'Copy'; }, 1500);
      });
      };
      codeBlock.parentElement.style.position = 'relative';
      codeBlock.parentElement.appendChild(btn);
      });
      }
      addCopyButtons();
      // Re-run on dynamic content
      var observer = new MutationObserver(addCopyButtons);
      observer.observe(document.body, { childList: true, subtree: true });
      })();
      }
      } catch(__e) { console.warn('[Userscript:Add Copy Buttons to Code Blocks]', __e); }
      })();
      (function(){
      try {
      var __m = "github.com";
      var __re = new RegExp('^' + "github\\.com" + '
      
      Skip to content

      [Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

      Description

      @JohnHales

      Bug Description

      In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

      locals {
      routing_config=[
      { name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
      { name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
      "${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
      ].ip_cidr_range] },
      ]
      }

      The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

      The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

      Environment and Deployment Context

      • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
      • Deployment Type:
        • US Region Restricted (e.g., Access Policy constraint)
        • FedRAMP Medium
        • FedRAMP High
        • DoD IL4
        • DoD IL5
        • Stand-alone / Custom
      • FAST Stage (if applicable):
        • Stage 0 (Bootstrap)
        • Stage 1 (Resource Management)
        • Stage 2 (Network Creation)
        • Stage 3 (Security and Audit)
      • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

      Steps to Reproduce

      1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
      2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
      3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
      4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
      5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

      Expected Behavior

      The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

      Actual Behavior

      Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

      Relevant Logs and Errors

      # from a VM in the SECOND tenant subnet (10.1.1.4):
      $ curl -sS -m 15 https://www.google.com
      curl: (28) Connection timed out after 15000 milliseconds
      $ ping -c2 8.8.8.8
      2 packets transmitted, 0 received, 100% packet loss
      # on the NVA:
      $ ip route get 10.1.1.4
      10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
      $ ip route | grep -E '10\.[123]\.'
      10.1.0.0/24 via 10.0.1.1 dev eth1
      10.2.0.0/24 via 10.0.1.1 dev eth1
      10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24
      

      Expected/Suggested Fix

      Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

      routes=flatten([fork, vinvar.envs_folders: [
      forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
      ]])

      This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

      Additional Context

      Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

      Metadata

      Metadata

      Assignees

      No one assigned

        Labels

        Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

        Type

        Projects

        No projects

          Milestone

          No milestone

          Relationships

          None yet

          Development

          No branches or pull requests

          Issue actions

          , 'i'); if (__m === '*' || __re.test(location.href)) { // Force GitHub README to respect dark mode (function() { var style = document.createElement('style'); style.textContent = ' .markdown-body { color-scheme: dark light; } .markdown-body pre { background: #161b22 !important; } .markdown-body code { background: rgba(110, 118, 129, 0.4) !important; } .markdown-body table th, .markdown-body table td { border-color: #30363d !important; } .markdown-body img { background: #0d1117; } .markdown-body blockquote { border-left-color: #8b949e; } .markdown-body hr { border-color: #30363d; } '; document.head.appendChild(style); })(); } } catch(__e) { console.warn('[Userscript:GitHub Dark Mode README Fix]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
          Skip to content

          [Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

          Description

          @JohnHales

          Bug Description

          In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

          locals {
          routing_config=[
          { name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
          { name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
          "${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
          ].ip_cidr_range] },
          ]
          }

          The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

          The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

          Environment and Deployment Context

          • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
          • Deployment Type:
            • US Region Restricted (e.g., Access Policy constraint)
            • FedRAMP Medium
            • FedRAMP High
            • DoD IL4
            • DoD IL5
            • Stand-alone / Custom
          • FAST Stage (if applicable):
            • Stage 0 (Bootstrap)
            • Stage 1 (Resource Management)
            • Stage 2 (Network Creation)
            • Stage 3 (Security and Audit)
          • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

          Steps to Reproduce

          1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
          2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
          3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
          4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
          5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

          Expected Behavior

          The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

          Actual Behavior

          Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

          Relevant Logs and Errors

          # from a VM in the SECOND tenant subnet (10.1.1.4):
          $ curl -sS -m 15 https://www.google.com
          curl: (28) Connection timed out after 15000 milliseconds
          $ ping -c2 8.8.8.8
          2 packets transmitted, 0 received, 100% packet loss
          # on the NVA:
          $ ip route get 10.1.1.4
          10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
          $ ip route | grep -E '10\.[123]\.'
          10.1.0.0/24 via 10.0.1.1 dev eth1
          10.2.0.0/24 via 10.0.1.1 dev eth1
          10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24
          

          Expected/Suggested Fix

          Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

          routes=flatten([fork, vinvar.envs_folders: [
          forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
          ]])

          This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

          Additional Context

          Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

          Metadata

          Metadata

          Assignees

          No one assigned

            Labels

            Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

            Type

            Projects

            No projects

              Milestone

              No milestone

              Relationships

              None yet

              Development

              No branches or pull requests

              Issue actions

              , 'i'); if (__m === '*' || __re.test(location.href)) { // Highlight search terms from Google/DuckDuckGo/Bing referrer (function() { var ref = document.referrer; var terms = []; if (ref.includes('google.com') || ref.includes('duckduckgo.com') || ref.includes('bing.com')) { var url = new URL(ref); var q = url.searchParams.get('q') || url.searchParams.get('p'); if (q) { terms = q.split(/\s+/).filter(function(t) { return t.length > 2; }); } } if (terms.length === 0) return; var style = document.createElement('style'); style.textContent = '.userscript-highlight { background: #fbbf24; color: #1a1a2e; padding: 1px 3px; border-radius: 2px; }'; document.head.appendChild(style); function highlight(node) { if (node.nodeType === 3) { // text node var text = node.textContent; var found = false; terms.forEach(function(term) { var regex = new RegExp('(' + term.replace(/[.*+?^${}()|[\]\\]/g, '\\') + ')', 'gi'); if (regex.test(text)) { found = true; var frag = document.createDocumentFragment(); var parts = text.split(regex); parts.forEach(function(part, i) { if (i % 2 === 0) { frag.appendChild(document.createTextNode(part)); } else { var span = document.createElement('span'); span.className = 'userscript-highlight'; span.textContent = part; frag.appendChild(span); } }); node.parentNode.replaceChild(frag, node); } }); } else if (node.nodeType === 1 && node.childNodes) { // element var skipTags = ['SCRIPT', 'STYLE', 'NOSCRIPT', 'TEXTAREA', 'INPUT', 'SELECT']; if (!skipTags.includes(node.tagName)) { Array.from(node.childNodes).forEach(highlight); } } } highlight(document.body); // Re-highlight on dynamic content var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1 || node.nodeType === 3) highlight(node); }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Highlight Search Terms]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
              Skip to content

              [Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

              Description

              @JohnHales

              Bug Description

              In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

              locals {
              routing_config=[
              { name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
              { name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
              "${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
              ].ip_cidr_range] },
              ]
              }

              The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

              The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

              Environment and Deployment Context

              • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
              • Deployment Type:
                • US Region Restricted (e.g., Access Policy constraint)
                • FedRAMP Medium
                • FedRAMP High
                • DoD IL4
                • DoD IL5
                • Stand-alone / Custom
              • FAST Stage (if applicable):
                • Stage 0 (Bootstrap)
                • Stage 1 (Resource Management)
                • Stage 2 (Network Creation)
                • Stage 3 (Security and Audit)
              • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

              Steps to Reproduce

              1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
              2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
              3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
              4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
              5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

              Expected Behavior

              The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

              Actual Behavior

              Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

              Relevant Logs and Errors

              # from a VM in the SECOND tenant subnet (10.1.1.4):
              $ curl -sS -m 15 https://www.google.com
              curl: (28) Connection timed out after 15000 milliseconds
              $ ping -c2 8.8.8.8
              2 packets transmitted, 0 received, 100% packet loss
              # on the NVA:
              $ ip route get 10.1.1.4
              10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
              $ ip route | grep -E '10\.[123]\.'
              10.1.0.0/24 via 10.0.1.1 dev eth1
              10.2.0.0/24 via 10.0.1.1 dev eth1
              10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24
              

              Expected/Suggested Fix

              Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

              routes=flatten([fork, vinvar.envs_folders: [
              forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
              ]])

              This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

              Additional Context

              Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

              Metadata

              Metadata

              Assignees

              No one assigned

                Labels

                Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

                Type

                Projects

                No projects

                  Milestone

                  No milestone

                  Relationships

                  None yet

                  Development

                  No branches or pull requests

                  Issue actions

                  , 'i'); if (__m === '*' || __re.test(location.href)) { // Strip utm_, fbclid, gclid, etc. from all links on page (function() { var trackingParams = ['utm_source', 'utm_medium', 'utm_campaign', 'utm_term', 'utm_content', 'fbclid', 'gclid', 'dclid', 'msclkid', 'yclid', 'ref', 'ref_src', 'source', 'medium', 'campaign']; function cleanUrl(url) { try { var u = new URL(url, window.location.origin); var changed = false; trackingParams.forEach(function(p) { if (u.searchParams.has(p)) { u.searchParams.delete(p); changed = true; } }); return changed ? u.toString() : url; } catch (e) { return url; } } function cleanLinks() { document.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } cleanLinks(); var observer = new MutationObserver(function(mutations) { mutations.forEach(function(m) { m.addedNodes.forEach(function(node) { if (node.nodeType === 1) { if (node.tagName === 'A') cleanLinks(); node.querySelectorAll('a[href]').forEach(function(a) { var clean = cleanUrl(a.href); if (clean !== a.href) a.href = clean; }); } }); }); }); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:Remove Tracking Parameters from Links]', __e); } })(); (function(){ try { var __m = "youtube.com"; var __re = new RegExp('^' + "youtube\\.com" + '
                  Skip to content

                  [Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

                  Description

                  @JohnHales

                  Bug Description

                  In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

                  locals {
                  routing_config=[
                  { name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
                  { name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
                  "${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
                  ].ip_cidr_range] },
                  ]
                  }

                  The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

                  The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

                  Environment and Deployment Context

                  • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
                  • Deployment Type:
                    • US Region Restricted (e.g., Access Policy constraint)
                    • FedRAMP Medium
                    • FedRAMP High
                    • DoD IL4
                    • DoD IL5
                    • Stand-alone / Custom
                  • FAST Stage (if applicable):
                    • Stage 0 (Bootstrap)
                    • Stage 1 (Resource Management)
                    • Stage 2 (Network Creation)
                    • Stage 3 (Security and Audit)
                  • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

                  Steps to Reproduce

                  1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
                  2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
                  3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
                  4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
                  5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

                  Expected Behavior

                  The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

                  Actual Behavior

                  Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

                  Relevant Logs and Errors

                  # from a VM in the SECOND tenant subnet (10.1.1.4):
                  $ curl -sS -m 15 https://www.google.com
                  curl: (28) Connection timed out after 15000 milliseconds
                  $ ping -c2 8.8.8.8
                  2 packets transmitted, 0 received, 100% packet loss
                  # on the NVA:
                  $ ip route get 10.1.1.4
                  10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
                  $ ip route | grep -E '10\.[123]\.'
                  10.1.0.0/24 via 10.0.1.1 dev eth1
                  10.2.0.0/24 via 10.0.1.1 dev eth1
                  10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24
                  

                  Expected/Suggested Fix

                  Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

                  routes=flatten([fork, vinvar.envs_folders: [
                  forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
                  ]])

                  This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

                  Additional Context

                  Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

                  Metadata

                  Metadata

                  Assignees

                  No one assigned

                    Labels

                    Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

                    Type

                    Projects

                    No projects

                      Milestone

                      No milestone

                      Relationships

                      None yet

                      Development

                      No branches or pull requests

                      Issue actions

                      , 'i'); if (__m === '*' || __re.test(location.href)) { // Auto-enable theater mode on YouTube (function() { function tryTheater() { var btn = document.querySelector('button[aria-label="Theater mode"], ytd-player #player button[title="Theater mode"]'); if (btn && !btn.classList.contains('activated')) { btn.click(); } } // Try immediately tryTheater(); // Try after navigation (SPA) var lastUrl = location.href; setInterval(function() { if (location.href !== lastUrl) { lastUrl = location.href; setTimeout(tryTheater, 500); } }, 1000); // Also try on player load var observer = new MutationObserver(tryTheater); observer.observe(document.body, { childList: true, subtree: true }); })(); } } catch(__e) { console.warn('[Userscript:YouTube Theater Mode Default]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                      Skip to content

                      [Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

                      Description

                      @JohnHales

                      Bug Description

                      In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

                      locals {
                      routing_config=[
                      { name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
                      { name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
                      "${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
                      ].ip_cidr_range] },
                      ]
                      }

                      The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

                      The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

                      Environment and Deployment Context

                      • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
                      • Deployment Type:
                        • US Region Restricted (e.g., Access Policy constraint)
                        • FedRAMP Medium
                        • FedRAMP High
                        • DoD IL4
                        • DoD IL5
                        • Stand-alone / Custom
                      • FAST Stage (if applicable):
                        • Stage 0 (Bootstrap)
                        • Stage 1 (Resource Management)
                        • Stage 2 (Network Creation)
                        • Stage 3 (Security and Audit)
                      • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

                      Steps to Reproduce

                      1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
                      2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
                      3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
                      4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
                      5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

                      Expected Behavior

                      The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

                      Actual Behavior

                      Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

                      Relevant Logs and Errors

                      # from a VM in the SECOND tenant subnet (10.1.1.4):
                      $ curl -sS -m 15 https://www.google.com
                      curl: (28) Connection timed out after 15000 milliseconds
                      $ ping -c2 8.8.8.8
                      2 packets transmitted, 0 received, 100% packet loss
                      # on the NVA:
                      $ ip route get 10.1.1.4
                      10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
                      $ ip route | grep -E '10\.[123]\.'
                      10.1.0.0/24 via 10.0.1.1 dev eth1
                      10.2.0.0/24 via 10.0.1.1 dev eth1
                      10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24
                      

                      Expected/Suggested Fix

                      Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

                      routes=flatten([fork, vinvar.envs_folders: [
                      forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
                      ]])

                      This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

                      Additional Context

                      Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

                      Metadata

                      Metadata

                      Assignees

                      No one assigned

                        Labels

                        Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

                        Type

                        Projects

                        No projects

                          Milestone

                          No milestone

                          Relationships

                          None yet

                          Development

                          No branches or pull requests

                          Issue actions

                          , 'i'); if (__m === '*' || __re.test(location.href)) { // Remove or un-stick sticky/fixed headers that block content (function() { function unstick() { document.querySelectorAll('header, nav, [role="banner"], .header, .navbar, .sticky, .fixed-top, [style*="position: fixed"], [style*="position:sticky"]').forEach(function(el) { if (el.style.position === 'fixed' || el.style.position === 'sticky' || getComputedStyle(el).position === 'fixed' || getComputedStyle(el).position === 'sticky') { el.style.position = 'static'; el.style.top = 'auto'; el.style.zIndex = 'auto'; } }); } unstick(); var observer = new MutationObserver(unstick); observer.observe(document.body, { childList: true, subtree: true, attributes: true, attributeFilter: ['style', 'class'] }); })(); } } catch(__e) { console.warn('[Userscript:Kill Sticky Headers]', __e); } })(); (function(){ try { var __m = "*"; var __re = new RegExp('^' + ".*" + '
                          Skip to content

                          [Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

                          Description

                          @JohnHales

                          Bug Description

                          In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

                          locals {
                          routing_config=[
                          { name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
                          { name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
                          "${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
                          ].ip_cidr_range] },
                          ]
                          }

                          The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

                          The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

                          Environment and Deployment Context

                          • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
                          • Deployment Type:
                            • US Region Restricted (e.g., Access Policy constraint)
                            • FedRAMP Medium
                            • FedRAMP High
                            • DoD IL4
                            • DoD IL5
                            • Stand-alone / Custom
                          • FAST Stage (if applicable):
                            • Stage 0 (Bootstrap)
                            • Stage 1 (Resource Management)
                            • Stage 2 (Network Creation)
                            • Stage 3 (Security and Audit)
                          • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

                          Steps to Reproduce

                          1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
                          2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
                          3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
                          4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
                          5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

                          Expected Behavior

                          The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

                          Actual Behavior

                          Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

                          Relevant Logs and Errors

                          # from a VM in the SECOND tenant subnet (10.1.1.4):
                          $ curl -sS -m 15 https://www.google.com
                          curl: (28) Connection timed out after 15000 milliseconds
                          $ ping -c2 8.8.8.8
                          2 packets transmitted, 0 received, 100% packet loss
                          # on the NVA:
                          $ ip route get 10.1.1.4
                          10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
                          $ ip route | grep -E '10\.[123]\.'
                          10.1.0.0/24 via 10.0.1.1 dev eth1
                          10.2.0.0/24 via 10.0.1.1 dev eth1
                          10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24
                          

                          Expected/Suggested Fix

                          Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

                          routes=flatten([fork, vinvar.envs_folders: [
                          forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
                          ]])

                          This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

                          Additional Context

                          Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

                          Metadata

                          Metadata

                          Assignees

                          No one assigned

                            Labels

                            Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

                            Type

                            Projects

                            No projects

                              Milestone

                              No milestone

                              Relationships

                              None yet

                              Development

                              No branches or pull requests

                              Issue actions

                              , 'i'); if (__m === '*' || __re.test(location.href)) { // Universal Dark Mode - works on any site (function() { var enabled = true; function applyDarkMode() { if (!enabled) return; // Create style element if it doesn't exist var style = document.getElementById('universal-dark-mode-style'); if (!style) { style = document.createElement('style'); style.id = 'universal-dark-mode-style'; document.head.appendChild(style); } // Dark mode CSS - inverts colors but preserves images/video style.textContent = ' /* Invert everything except media */ html { filter: invert(1) hue-rotate(180deg) !important; background: #1a1a2e !important; } /* Restore images, videos, iframes, canvas */ img, video, iframe, canvas, svg, picture, [style*="background-image"] { filter: invert(1) hue-rotate(180deg) !important; } /* Preserve specific elements that should not be inverted */ .no-dark-mode, .no-dark-mode *, [data-theme="light"], [data-theme="light"], .ace_editor, .ace_editor *, .CodeMirror, .CodeMirror *, .monaco-editor, .monaco-editor *, .markdown-body pre, .markdown-body pre *, .highlight, .highlight *, pre code, pre code * { filter: none !important; } /* Fix common UI elements */ .modal, .popup, .dropdown-menu, .tooltip, .popover { filter: invert(1) hue-rotate(180deg) !important; background: #2d2d44 !important; border-color: #444 !important; } /* Scrollbars */ ::-webkit-scrollbar { background: #1a1a2e !important; } ::-webkit-scrollbar-thumb { background: #444 !important; } ::-webkit-scrollbar-thumb:hover { background: #555 !important; } /* Selection */ ::selection { background: #4ecdc4 !important; color: #1a1a2e !important; } ::-moz-selection { background: #4ecdc4 !important; color: #1a1a2e !important; } '; } function removeDarkMode() { var style = document.getElementById('universal-dark-mode-style'); if (style) style.remove(); } // Toggle with Alt+Shift+D document.addEventListener('keydown', function(e) { if (e.altKey && e.shiftKey && e.key === 'D') { e.preventDefault(); enabled = !enabled; if (enabled) { applyDarkMode(); console.log('[Universal Dark Mode] Enabled'); } else { removeDarkMode(); console.log('[Universal Dark Mode] Disabled'); } } }); // Apply on load applyDarkMode(); // Re-apply on dynamic content var observer = new MutationObserver(function(mutations) { if (enabled && !document.getElementById('universal-dark-mode-style')) { applyDarkMode(); } }); observer.observe(document.head, { childList: true }); console.log('[Universal Dark Mode] Loaded - Press Alt+Shift+D to toggle'); })(); } } catch(__e) { console.warn('[Userscript:Universal Dark Mode]', __e); } })(); })();
                              Skip to content

                              [Bug] 2-networking-a-fedramp-high: NVA return-route generation only routes the FIRST tenant subnet per environment ([0] index) — additional tenant subnets get no NVA return route and lose ALL internet egress #146

                              Description

                              @JohnHales

                              Bug Description

                              In fast/stages-aw/2-networking-a-fedramp-high/nva.tf, the NVA cloud-config's trusted-side ("landing") return routes are built from a per-environment comprehension that selects only the FIRST tenant subnet of each environment:

                              locals {
                              routing_config=[
                              { name ="dmz", enable_masquerading =true, routes = [var.gcp_ranges.gcp_dmz_primary] },
                              { name ="landing", routes = [fork, vinvar.envs_folders:module.env-spoke-vpc[k].subnets[
                              "${var.regions.primary}/${[forsintry(var.subnets[lower(k)], []) :s.nameifs.tenant!=null][0]}"
                              ].ip_cidr_range] },
                              ]
                              }

                              The trailing […][0] means each environment contributes exactly ONE tenant subnet CIDR to the NVA's return routes. But var.subnets is typed map(list(object(... name, ip_cidr_range, tenant = optional(string) ...))) — a LIST of subnets per network — so an environment can legitimately hold multiple tenant subnets. Any tenant subnet that is not the first in its environment's list receives no return route on the NVA.

                              The NVA (the simple-nva cloud-config, enable_masquerading = true on the DMZ interface) masquerades outbound workload traffic on its untrusted/DMZ interface and forwards it correctly. But when the reply returns, the NVA looks up the original client IP, finds no trusted-side route for that subnet, falls through to its own default route, and sends the reply back out the DMZ/untrusted interface toward the internet gateway instead of back to the trusted spoke. The connection hangs. Net effect for VMs in the un-routed subnet: name resolution still works (the metadata resolver needs no egress) but 100% of real internet egress (TCP and ICMP) times out — a silent, hard-to-diagnose failure.

                              Environment and Deployment Context

                              • Stellar Engine Version/Commit:main (fast/stages-aw/2-networking-a-fedramp-high/nva.tf, local.routing_config "landing" routes)
                              • Deployment Type:
                                • US Region Restricted (e.g., Access Policy constraint)
                                • FedRAMP Medium
                                • FedRAMP High
                                • DoD IL4
                                • DoD IL5
                                • Stand-alone / Custom
                              • FAST Stage (if applicable):
                                • Stage 0 (Bootstrap)
                                • Stage 1 (Resource Management)
                                • Stage 2 (Network Creation)
                                • Stage 3 (Security and Audit)
                              • Affected Component:fast/stages-aw/2-networking-a-fedramp-high/nva.tflocal.routing_config, the landingroutes comprehension ([for s in try(var.subnets[lower(k)], []) : s.name if s.tenant != null][0]); the value is consumed by module.nva-cloud-config (modules/cloud-config-container/simple-nva) and pushed to both NVAs as their trusted-side static routes.

                              Steps to Reproduce

                              1. Deploy an SE FRH landing zone whose prod environment spoke has TWO tenant subnets — e.g. a base tenant subnet (10.1.0.0/24) and a second tenant subnet (10.1.1.0/24), both with tenant set.
                              2. Bring up a VM in the SECOND subnet (10.1.1.x), no external IP.
                              3. From that VM: curl -m 15 https://www.google.com and ping -c2 8.8.8.8 — both time out; DNS still resolves (metadata resolver at 169.254.169.254).
                              4. On either NVA: ip route get 10.1.1.4 → resolves via <dmz-gw> dev eth0 (the default/untrusted interface), NOT toward the trusted spoke; ip route shows landing return routes only for each environment's FIRST tenant subnet (10.1.0.0/24, 10.2.0.0/24, 10.3.0.0/24) and none for 10.1.1.0/24.
                              5. Add the missing route on both NVAs: sudo ip route replace 10.1.1.0/24 via <landing-gw> dev eth1 — egress immediately works (HTTP 200, ping 0% loss). This confirms the NVA return-route omission is the sole cause.

                              Expected Behavior

                              The NVA is given return routes for ALL tenant subnets in every environment, so every tenant subnet's VMs egress correctly through the NVA + Cloud NAT.

                              Actual Behavior

                              Only the first tenant subnet per environment gets a return route; VMs in any additional tenant subnet lose all internet egress (silently — DNS still resolves), while the NVA misroutes their return traffic out its untrusted interface.

                              Relevant Logs and Errors

                              # from a VM in the SECOND tenant subnet (10.1.1.4):
                              $ curl -sS -m 15 https://www.google.com
                              curl: (28) Connection timed out after 15000 milliseconds
                              $ ping -c2 8.8.8.8
                              2 packets transmitted, 0 received, 100% packet loss
                              # on the NVA:
                              $ ip route get 10.1.1.4
                              10.1.1.4 via 10.0.0.1 dev eth0 src 10.0.0.20 # <- out the DMZ/default iface, WRONG WAY
                              $ ip route | grep -E '10\.[123]\.'
                              10.1.0.0/24 via 10.0.1.1 dev eth1
                              10.2.0.0/24 via 10.0.1.1 dev eth1
                              10.3.0.0/24 via 10.0.1.1 dev eth1 # <- no route for 10.1.1.0/24
                              

                              Expected/Suggested Fix

                              Replace the single-subnet [0] pick with a flatten over ALL tenant subnets per environment:

                              routes=flatten([fork, vinvar.envs_folders: [
                              forsintry(var.subnets[lower(k)], []) :module.env-spoke-vpc[k].subnets["${var.regions.primary}/${s.name}"].ip_cidr_rangeifs.tenant!=null
                              ]])

                              This routes every tenant subnet back through the NVA, so adding a second tenant subnet no longer silently breaks its egress.

                              Additional Context

                              Verified live on an SE FRH deployment: before the fix, egress from the subnet was 100% broken; adding the missing NVA return route on both NVAs restored it (confirmed HTTP 200 + a Cloud NAT public egress IP + ping 0% loss). Belongs with the other subnet-handling gaps in the gemini ingress path (#105 hardcoded LB subnet, #111 fragile network resolution), but this one is upstream in Stage-2 networking, not the gemini blueprint. NOTE: the ephemeral ip route replace fix does NOT survive an NVA reboot / MIG replacement / networking-stage redeploy — the durable fix is the source change above (or a persistent route baked into the NVA cloud-config).

                              Metadata

                              Metadata

                              Assignees

                              No one assigned

                                Labels

                                Level of Effort - LowQuick, well-defined tasks with no unknowns; takes a few hours up to one day to completePriority - MediumStandard features and non-blocking bugs; important for the current milestone but not urgentbugSomething isn't working

                                Type

                                Projects

                                No projects

                                  Milestone

                                  No milestone

                                  Relationships

                                  None yet

                                  Development

                                  No branches or pull requests

                                  Issue actions