Skip to content

feat: support multi-tunnel RouteConfiguration and FirewallConfiguration using traffic marking - #3363

Open
MircoBarone wants to merge 6 commits into
liqotech:masterfrom
MircoBarone:PR14-multitunnel-gw-node-trafficmarking
Open

MircoBarone wants to merge 6 commits into
liqotech:masterfrom
MircoBarone:PR14-multitunnel-gw-node-trafficmarking

Conversation

@MircoBarone

@MircoBarone MircoBarone commented Aug 27, 2026 •

Copy link
Copy Markdown
Contributor

Part of the multi-tunnel WireGuard implementation (8th PR related to #3225).


RouteConfiguration Changes

This PR replaces what was done in #3338 (now closed). The main idea of that PR was to use a placeholder mechanism (liqo-tunnel*) inside the gw-node and extcidr RouteConfiguration resources, which was then expanded inside the gateway based on the actual number of interfaces.

Instead, this PR adopts a traffic-marking strategy, extending the pattern introduced in #3357.

The main idea is to tag packets arriving on WireGuard tunnel interfaces (liqo-tunnel*) with a specific fwmark (0xFE00). As done in #3357, this is applied in the prerouting chain at mangle priority.

In the RouteConfiguration resources, matching on iif: liqo-tunnel is replaced with the FwMark match. This eliminates the need to expand rules inside the gateway, avoiding the creation of $M \times N$ policy routing rules (where $M$ is the number of WireGuard ingress interfaces and $N$ is the number of target nodes). Instead, the system now maintains just 1 rule per target node, regardless of the interface count.


FirewallConfiguration Changes

This PR also extends the traffic-marking strategy to the FirewallConfiguration resources responsible for PodCIDR and ExternalCIDR remapping.

As a brief recap: outgoing Liqo traffic must be DNATted, which is handled in the prerouting chain; incoming Liqo traffic must be SNATted, which is handled in the postrouting chain.

Outgoing Path

In the single-interface baseline, the DNAT rule requires that traffic is not entering from eth0 nor from liqo-tunnel. Extending this to multiple WireGuard interfaces would have required enumerating all active interfaces in a set and using a not-in operator, an approach that ties rule complexity to the number of interfaces.

Thanks to traffic marking, this can be simplified: instead of excluding a set of interfaces, we can simply require that traffic carries the Geneve mark (GwExtMark = 0xFF00), which is equivalent but scales with no additional complexity.

OutgoingPath drawio

Example of an nft rule for the PodCIDR, following the example depicted in the picture above:

chain outgoing {
    type nat hook prerouting priority dstnat; policy accept;
    ip daddr 10.61.0.0/16 meta mark 0x0000ff00 dnat prefix to 10.200.0.0/16 comment "d2cb4786-9122-4fad-b287-3054fdb30122"
}

Incoming Path

In the single-interface baseline, the SNAT rule requires that traffic is arriving from the WireGuard interface (liqo-tunnel). Extending this to multiple interfaces would again have required a set-based in operator.

Here too, traffic marking provides a cleaner solution: rather than matching on a set of interfaces, we require that traffic carries the WireGuard mark (GwNodeMark = 0xFE00), assigned in prerouting when packets arrive on any liqo-tunnel* interface.

IncomingPath drawio

Here is the rule for the PodCIDR, referring to the scenario depicted in the picture above:

chain incoming {
    type nat hook postrouting priority srcnat; policy accept;
    oifname != "eth0" ip saddr 10.200.0.0/16 meta mark 0x0000fe00 snat prefix to 10.61.0.0/16 comment "1bbb6e2f-8669-4ec0-bacd-110d165e703f"
}

Additional Fixes

Mark Collision in conntrack-mark-to-meta-mark

Liqo uses the first address of the local External CIDR to identify traffic coming from NodePort services. This traffic does not have a well-defined source IP: it can arrive from any Geneve interface, since a NodePort service is reachable on every node of the cluster. Liqo automatically assigns the first IP of the local External CIDR to this kind of traffic.

For this type of traffic, in the outgoing path, in the forward chain, an nft rule sets the ct mark on the conntrack entry based on the Geneve interface the traffic is coming from. Each node is associated with a mark, which is recorded in the conntrack entry. An example of this rule, in a setup with two nodes, is:

table ip service-nodeport-routing {
    chain mark-to-conntrack {
        type filter hook forward priority filter; policy accept;
        ip saddr 10.70.0.0 iifname "liqo.n556fkng95" ct mark set 0x00000001 comment "rome-worker"
        ip saddr 10.70.0.0 iifname "liqo.t4wpt9k7mb" ct mark set 0x00000002 comment "rome-control-plane"
    }
}

This mark must be restored in the incoming path and written into the packet's meta mark, so that the routing decision can determine which Geneve interface the flow should return through. An example of a rule that does this is:

table ip service-nodeport-routing {
    chain conntrack-mark-to-meta-mark {
        type filter hook prerouting priority filter; policy accept;
        ip daddr 10.70.0.0 meta mark 0x0000fe00 meta mark set ct mark comment "conntrack-mark-to-meta-mark"
    }
}

The problem, as depicted in the figure below, is that the meta mark set by this rule overwrites the original mark (0xFE) that was set because the traffic was coming from WireGuard.

As a result, the problem manifests in postrouting: the rule performing SRC-NAT requires the packet to carry the 0xFE mark in order to be SNATed. If this mark is overwritten, SRC-NAT cannot occur. The conclusion is that traffic coming from a NodePort service cannot be NATed on return.

MarkProblem drawio

Solutions

A first solution is to combine the two marks using a bitwise OR, since they operate on different bit ranges (high vs. low bits). Matching rules would then need to use an appropriate bitmask to isolate the relevant bits. This solution is somewhat complicated, as it requires using masks both in the routing decision and in the postrouting chain.

The adopted solution (Implemented in ac375fc) consists of an additional rule in the postrouting chain, forcing the remapping for traffic destined to 10.70.0.0. An example of this rule, for the PodCIDR, is:

table ip remap-podcidr {
    chain incoming {
        type nat hook postrouting priority srcnat; policy accept;

        oifname != "eth0" ip saddr 10.200.0.0/16 ip daddr 10.70.0.0 meta mark != 0x0000ff00 snat prefix to 10.61.0.0/16 comment "acc43ae0-add0-48b2-83f8-63d8b14af742"
    }
}

This rule selects traffic that needs remapping (coming from the local remote PodCIDR) and destined to 10.70.0.0, and applies SNAT.

Note

In the conducted tests, it works even without this rule, thanks to conntrack storing the 5-tuple: the returning packet is matched by the existing conntrack entry, so the return remapping is not strictly needed. However, I'm not confident in relying on conntrack alone, without an explicit rule, so it was preferable to add it.

Why is there a protection against Geneve traffic (meta mark != 0x0000ff00)?

Consider the NodePort case, referring to the example in the picture above. Traffic 10.70.0.0 → 10.61.X.X is generated on a node of cluster A, where 10.61.X.X is how A sees the (remapped) target pod on cluster B. As already described, A immediately performs DNAT, and the packet reaches B as 10.70.0.0 → 10.200.X.X. At this point marking is not involved: it simply reaches the correct pod.

To simplify (ignoring an intermediate NAT step that doesn't affect this part of the flow), the pod's reply re-enters gateway B through a Geneve interface, as 10.200.X.X → 10.70.0.0. Since this traffic comes from Geneve, it does not match the conntrack-mark-to-meta-mark rule, so its mark remains 0xFF00.

This is exactly the traffic that must not be remapped on B: it must be forwarded as-is to gateway A, which is responsible for remapping it back to the originating node. The problem is that the remap-podcidr rule matches on oifname != eth0, saddr 10.200.0.0/16 and daddr 10.70.0.0, all of which are satisfied by this reply packet, even though it is incoming traffic coming from Geneve. Without the mark exclusion, the rule would remap the packet prematurely on B, before it has a chance to reach A.

This premature remapping would also be incorrect: it would apply B's remapping scheme for A's PodCIDR (10.61.0.0/16 in the example above), while the correct remapping to use is A's own scheme for B's PodCIDR, and the two are not guaranteed to match, even though they happen to coincide in this example.

MSS Clamping Interface Expansion

The mss-clamping rule, applied in the forward chain, matches on oifname: liqo-tunnel. Unlike the cases addressed in this PR, this rule cannot be replaced with a mark-based match and requires explicit interface enumeration. A set-based expansion over the actual WireGuard interfaces remains necessary here, unless an alternative expression is found.
(Implemented in 1a209f5)

Example of a rule for a scenario with 3 tunnels. Naming the set after the number of tunnels it contains is not ambiguous: by construction, interfaces are always named sequentially, so a set for N tunnels always contains liqo-tunnel, liqo-tunnel1, ..., liqo-tunnel(N-1). There is no case where, for example, tunnel-list-3 would contain arbitrary indices such as liqo-tunnel7, liqo-tunnel9, liqo-tunnel10:

table ip mss-clamping {
    set tunnel-list-3 {
        type ifname
        elements = { "liqo-tunnel",
                     "liqo-tunnel1",
                     "liqo-tunnel2" }
    }

    chain mss-clamping {
        type filter hook forward priority filter; policy accept;
        meta l4proto tcp oifname @tunnel-list-3 counter packets 1 bytes 1 tcp flags syn tcp option maxseg size set rt mtu comment "mss-clamping-out"
    }
}

The only problem can arise from the combination with ECMP. Suppose the MSS is set for flow X when the SYN packet passes through interface A. If ECMP later moves the flow to interface B, this can be problematic if MTU(A) != MTU(B). This scenario is true in general, but by construction, WireGuard interfaces are all set with the same MTU.


Tests

Tested in a local kind environment with two clusters by running iperf3 tests between two pods in distinct clusters, using a branch that combines all commits related to the multi-tunnel issue (#3225). Tests also cover service communication.

Unit tests were also added for the new firewall matches (Set, Mark), the MSS clamping expansion, the CIDR remapping rules and the mark-based route/firewall configurations.

@adamjensenbot

Copy link
Copy Markdown
Collaborator

Hi @MircoBarone. Thanks for your PR!

I am @adamjensenbot.
You can interact with me issuing a slash command in the first line of a comment.
Currently, I understand the following commands:

  • /rebase: Rebase this PR onto the master branch (You can add the option test=true to launch the tests
    when the rebase operation is completed)
  • /merge: Merge this PR into the master branch
  • /build Build Liqo components
  • /test Launch the E2E and Unit tests
  • /hold, /unhold Add/remove the hold label to prevent merging with /merge

Make sure this PR appears in the liqo changelog, adding one of the following labels:

  • feat: 🚀 New Feature
  • fix: 🐛 Bug Fix
  • refactor: 🧹 Code Refactoring
  • docs: 📝 Documentation
  • style: 💄 Code Style
  • perf: 🐎 Performance Improvement
  • test: ✅ Tests
  • chore: 🚚 Dependencies Management
  • build: 📦 Builds Management
  • ci: 👷 CI/CD
  • revert: ⏪ Reverts Previous Changes

@github-actions github-actions Bot added feat Adds a new feature to the codebase fix Fixes a bug in the codebase. labels Aug 27, 2026
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch from 8c2fb1a to 27b1f36 Compare September 4, 2026 19:33
@MircoBarone
MircoBarone marked this pull request as draft September 8, 2026 10:39
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch 4 times, most recently from f89a360 to fd315df Compare September 17, 2026 20:15
@MircoBarone MircoBarone changed the title Support multi-tunnel RouteConfiguration using traffic marking Support multi-tunnel RouteConfiguration adn FirewallConfiguration using traffic marking Sep 17, 2026
@MircoBarone MircoBarone changed the title Support multi-tunnel RouteConfiguration adn FirewallConfiguration using traffic marking Support multi-tunnel RouteConfiguration and FirewallConfiguration using traffic marking Sep 17, 2026
@MircoBarone MircoBarone changed the title Support multi-tunnel RouteConfiguration and FirewallConfiguration using traffic marking feat: support multi-tunnel RouteConfiguration and FirewallConfiguration using traffic marking Sep 17, 2026
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch from fd315df to 8337b66 Compare September 19, 2026 00:12
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch 5 times, most recently from cf5167d to 2897514 Compare September 24, 2026 20:01
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch 2 times, most recently from 6fb4492 to 27c99ed Compare September 24, 2026 21:18
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch from 27c99ed to ffad452 Compare September 24, 2026 21:26
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch from ffad452 to ac375fc Compare September 24, 2026 21:41
@MircoBarone
MircoBarone marked this pull request as ready for review September 28, 2026 22:14
@github-actions github-actions Bot added the test Adds or updates tests for the codebase label Sep 30, 2026
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch from 01f544e to 5c07119 Compare October 1, 2026 00:38
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch 3 times, most recently from 1dcaf21 to 0567e19 Compare October 5, 2026 23:28
The FirewallConfiguration gw-mss-clamping contains a rule matching liqo-tunnel
as outgoing interface. Unlike rules matching liqo-tunnel as incoming, which can
be handled with a traffic mark, there is no mark associated to traffic outgoing
a specific interface. The rule must therefore be expanded when multiple
WireGuard tunnels are present.

This expansion cannot be done upfront. gw-mss-clamping exists before any gateway
is created and is not tied to a specific one. Each gateway can have a different
number of tunnel interfaces. The idea is to treat liqo-tunnel as a placeholder
and delegate the expansion to the gateway itself, which expands the rule locally
based on the number of tunnel interfaces it finds.

The placeholder is replaced with an nftables set containing all real tunnel
interfaces present on that gateway. The resulting rule matches outgoing traffic
against the set, while the rest of the rule (proto, counters, MSS action)
remains unchanged.

Main changes:

* Introduced a new Match type called Set, expansion of Dev, containing a list
  of interfaces instead of a single one, and a Position in or out. Two new
  operations are associated to Set: In (corresponding to Eq for Dev) and Nin
  (corresponding to Neq).

* GetTunnelInterfaces reads net.Interfaces() and returns the sorted names of
  local interfaces matching the liqo-tunnel prefix. Called when rules must be
  created or cleaned.

* FromChainToRulesArray now takes the discovered tunnel list and passes each
  filter rule through expandOifFilterWithSet, which detects any Dev match
  targeting liqo-tunnel as outgoing interface and replaces it with a Set match
  in place, keeping the rest of the Match struct intact.

* ensureSetsForChain is called after rule expansion and creates the nftables
  set for each rule containing a Set match, before the rules are inserted.

* applyMatchSet is introduced to handle the new Set match type. Once the rules
  are expanded and the sets are in place, the gateway applies all rules
  automatically.
Traffic coming from a NodePort service and destined to offloaded pods in the
remote cluster (the endpoint of the service) has a source IP equal to the first
IP of the local ExternalCIDR (e.g. 10.70.0.0), which Liqo uses as the
unknown-source address. This traffic can enter the cluster on any node, since
a NodePort service is reachable on every node.

On the return path, packets enter from the WireGuard tunnel and receive
GwNodeMark, which also drives the SNAT remapping in postrouting. For traffic
destined to the unknown-source IP (10.70.0.0), the mark is then replaced with
the per-flow mark stored in conntrack, which identifies the correct Geneve
interface to reach the destination. This substitution overwrites GwNodeMark
before postrouting is reached, causing this class of traffic to miss the SNAT
rule and leave the source IP unremapped.

This commit introduces an additional SNAT rule that remaps packets whose source
is within the remote Pod or ExternalCIDR, ensuring correct remapping for
NodePort flows.
Add unit tests for the Set and Mark matches in pkg/firewall/utils,
both at match level and at filter rule level.

Code changes found while testing:
- applyMatchSet: reject ops other than in/nin and empty value lists,
  instead of silently treating them as "in" or emitting tunnel-list-0
- getMatchIPPositionOffset, getMatchPortPositionOffset,
  getMatchProtoValue, getMatchDevMetaKey: fix copy-pasted error
  messages that dereferenced m.Dev and caused a nil pointer panic
- add TunnelListSetName and use it in rule.go and applyMatchSet so the
  set name used at creation and in the Lookup cannot drift apart
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch from 0567e19 to 63a1091 Compare October 5, 2026 23:32
Created pkg/firewall/chain_test.go:
* Tested `expandOifFilterWithSet` (eq->in, neq->nin, no tunnels,
  no matching dev, multiple matches)
* Tested `FromChainToRulesArray` (tunnel threshold, rule order and
  selectivity)

Created pkg/firewall/rule_test.go:
* Tested `getMatchesFromRule` (filter, nat, route, nil rule)

Created pkg/liqo-controller-manager/networking/external-network/remapping/cidr_test.go:
* Tested `forgeCIDRFirewallConfigurationDNATRules` (Mark eq GwExtMark,
  no Dev match, PodCIDR and ExternalCIDR)
* Tested `forgeCIDRFirewallConfigurationSNATRules` (first rule: Mark eq
  GwNodeMark; second rule: Mark neq GwExtMark + unknown source IP)
* Tested that no rules are generated without remapping (identical
  pairs, empty lists, unknown CIDR type)
* Tested that only the first SNAT rule is generated when the local
  external CIDR is missing or invalid

Created pkg/liqo-controller-manager/networking/external-network/route/k8s_test.go:
* Tested `forgeMutateFirewallConfiguration` (gw-ext-mark and gw-node-mark:
  table, chain, wildcard Dev match, setmetamark value, labels, owner)
* Tested `forgeMutateRouteConfiguration` (one rule per CIDR with
  FwMark GwExtMark, no Iif, route via remote interface IP)
* Tested `GenerateRouteConfigurationName` and
  `GenerateFirewallConfigurationName` (names and non-collision)

Created pkg/liqo-controller-manager/networking/internal-network/route/route_k8s_test.go:
* Tested `forgeRoutePodUpdateFunction` (pod rule selected by GwNodeMark
  and not by Iif; host network pods get no pod rule)
* Tested `forgeFirewallConfigurationPreroutingChainRule` (Mark eq
  GwNodeMark instead of the tunnel Dev match)
* Tested `forgeRouteConfigurationExtCIDRRules` (FwMark GwNodeMark, no
  Iif, one rule per remote pod CIDR plus the final IP rule)

Code changes found while testing:
* forgeCIDRFirewallConfigurationSNATRules: skip the unknown-source-IP
  rule (used for NodePort traffic) when the local external CIDR is
  missing or invalid, instead of emitting a match with an empty
  destination IP
@MircoBarone
MircoBarone force-pushed the PR14-multitunnel-gw-node-trafficmarking branch from 63a1091 to c4057a2 Compare October 6, 2026 18:15

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

feat Adds a new feature to the codebase fix Fixes a bug in the codebase. size/XXL test Adds or updates tests for the codebase

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants