forked from cerc-io/stack-orchestrator
fix: ashburn relay playbooks and document DZ tunnel ACL root cause
Playbook fixes from testing: - ashburn-relay-biscayne: insert DNAT rules at position 1 before Docker's ADDRTYPE LOCAL rule (was being swallowed at position 3+) - ashburn-relay-mia-sw01: add inbound route for 137.239.194.65 via egress-vrf vrf1 (nexthop only, no interface — EOS silently drops cross-VRF routes that specify a tunnel interface) - ashburn-relay-was-sw01: replace PBR with static route, remove Loopback101 Bug doc (bug-ashburn-tunnel-port-filtering.md): root cause is the DoubleZero agent on mia-sw01 overwrites SEC-USER-500-IN ACL, dropping outbound gossip with src 137.239.194.65. The DZ agent controls Tunnel500's lifecycle. Fix requires a separate GRE tunnel using mia-sw01's free LAN IP (209.42.167.137) to bypass DZ infrastructure. Also adds all repo docs, scripts, inventory, and remaining playbooks. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This commit is contained in:
co-authored by
Claude Opus 4.6
parent
6841d5e3c3
commit
0b52fc99d7
@@ -0,0 +1,114 @@
|
||||
# Arista EOS Reference Notes
|
||||
|
||||
Collected from live switch CLI (`?` help) and Arista documentation search
|
||||
results. Switch platform: 7280CR3A, EOS 4.34.0F.
|
||||
|
||||
## PBR (Policy-Based Routing)
|
||||
|
||||
EOS uses `policy-map type pbr` — NOT `traffic-policy` (which is a different
|
||||
feature for ASIC-level traffic policies, not available on all platforms/modes).
|
||||
|
||||
### Syntax
|
||||
|
||||
```
|
||||
! ACL to match traffic
|
||||
ip access-list <ACL-NAME>
|
||||
10 permit <proto> <src> <dst> [ports]
|
||||
|
||||
! Class-map referencing the ACL
|
||||
class-map type pbr match-any <CLASS-NAME>
|
||||
match ip access-group <ACL-NAME>
|
||||
|
||||
! Policy-map with nexthop redirect
|
||||
policy-map type pbr <POLICY-NAME>
|
||||
class <CLASS-NAME>
|
||||
set nexthop <A.B.C.D> ! direct nexthop IP
|
||||
set nexthop recursive <A.B.C.D> ! recursive resolution
|
||||
! set nexthop-group <NAME> ! nexthop group
|
||||
! set ttl <value> ! TTL override
|
||||
|
||||
! Apply on interface
|
||||
interface <INTF>
|
||||
service-policy type pbr input <POLICY-NAME>
|
||||
```
|
||||
|
||||
### PBR `set` options (from CLI `?`)
|
||||
|
||||
```
|
||||
set ?
|
||||
nexthop Next hop IP address for forwarding
|
||||
nexthop-group next hop group name
|
||||
ttl TTL effective with nexthop/nexthop-group
|
||||
```
|
||||
|
||||
```
|
||||
set nexthop ?
|
||||
A.B.C.D next hop IP address
|
||||
A:B:C:D:E:F:G:H next hop IPv6 address
|
||||
recursive Enable Recursive Next hop resolution
|
||||
```
|
||||
|
||||
**No VRF qualifier on `set nexthop`.** The nexthop must be reachable in the
|
||||
VRF where the policy is applied. For cross-VRF PBR, use a static inter-VRF
|
||||
route to make the nexthop reachable (see below).
|
||||
|
||||
## Static Inter-VRF Routes
|
||||
|
||||
Source: [EOS 4.34.0F - Static Inter-VRF Route](https://www.arista.com/en/um-eos/eos-static-inter-vrf-route)
|
||||
|
||||
Allows configuring a static route in one VRF with a nexthop evaluated in a
|
||||
different VRF. Uses the `egress-vrf` keyword.
|
||||
|
||||
### Syntax
|
||||
|
||||
```
|
||||
ip route vrf <ingress-vrf> <prefix>/<mask> egress-vrf <egress-vrf> <nexthop-ip>
|
||||
ip route vrf <ingress-vrf> <prefix>/<mask> egress-vrf <egress-vrf> <interface>
|
||||
```
|
||||
|
||||
### Examples (from Arista docs)
|
||||
|
||||
```
|
||||
! Route in vrf1 with nexthop resolved in default VRF
|
||||
ip route vrf vrf1 1.0.1.0/24 egress-vrf default 1.0.0.2
|
||||
|
||||
! show ip route vrf vrf1 output:
|
||||
! S 1.0.1.0/24 [1/0] via 1.0.0.2, Vlan2180 (egress VRF default)
|
||||
```
|
||||
|
||||
### Key points
|
||||
|
||||
- For bidirectional traffic, static inter-VRF routes must be configured in
|
||||
both VRFs.
|
||||
- ECMP next-hop sets across same or heterogeneous egress VRFs are supported.
|
||||
- The `show ip route vrf` output displays the egress VRF name when it differs
|
||||
from the source VRF.
|
||||
|
||||
## Inter-VRF Local Route Leaking
|
||||
|
||||
Source: [EOS 4.35.1F - Inter-VRF Local Route Leaking](https://www.arista.com/en/um-eos/eos-inter-vrf-local-route-leaking)
|
||||
|
||||
An alternative to static inter-VRF routes that leaks routes dynamically from
|
||||
one VRF (source) to another VRF (destination) on the same router.
|
||||
|
||||
## Config Sessions
|
||||
|
||||
```
|
||||
configure session <name> ! enter named session
|
||||
show session-config diffs ! MUST be run from inside the session
|
||||
commit timer HH:MM:SS ! commit with auto-revert timer
|
||||
abort ! discard session
|
||||
```
|
||||
|
||||
From enable mode:
|
||||
```
|
||||
configure session <name> commit ! finalize a pending session
|
||||
```
|
||||
|
||||
## Checkpoints and Rollback
|
||||
|
||||
```
|
||||
configure checkpoint save <name>
|
||||
rollback running-config checkpoint <name>
|
||||
write memory
|
||||
```
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,181 @@
|
||||
<!-- Source: https://www.arista.com/um-eos/eos-ingress-and-egress-per-port-for-ipv4-and-ipv6-counters -->
|
||||
<!-- Scraped: 2026-03-06T20:50:41.080Z -->
|
||||
|
||||
# Ingress and Egress Per-Port for IPv4 and IPv6 Counters
|
||||
|
||||
|
||||
This feature supports per-interface ingress and egress packet and byte counters for IPv4
|
||||
and IPv6.
|
||||
|
||||
|
||||
This section describes Ingress and Egress per-port for IPv4 and IPv6 counters, including
|
||||
configuration instructions and command descriptions.
|
||||
|
||||
|
||||
Topics covered by this chapter include:
|
||||
|
||||
|
||||
- Configuration
|
||||
|
||||
- Show commands
|
||||
|
||||
- Dedicated ARP Entry for TX IPv4 and IPv6 Counters
|
||||
|
||||
- Considerations
|
||||
|
||||
|
||||
## Configuration
|
||||
|
||||
|
||||
IPv4 and IPv6 ingress counters (count **bridged and routed**
|
||||
traffic, supported only on front-panel ports) can be enabled and disabled using the
|
||||
**hardware counter feature ip in**
|
||||
command:
|
||||
|
||||
|
||||
```
|
||||
`**[no] hardware counter feature ip in**`
|
||||
```
|
||||
|
||||
|
||||
For IPv4 and IPv6 ingress and egress counters that include only
|
||||
**routed** traffic (supported on Layer3 interfaces such as
|
||||
routed ports and L3 subinterfaces only), use the following commands:
|
||||
|
||||
|
||||
Note: The DCS-7300X, DCS-7250X, DCS-7050X, and DCS-7060X platforms
|
||||
do not require configuration for IPv4 and IPv6 packet counters for only routed
|
||||
traffic. They are collected by default. Other platforms (DCS-7280SR, DCS-7280CR, and
|
||||
DCS-7500-R) need the feature enabled.
|
||||
|
||||
|
||||
```
|
||||
`**[no] hardware counter feature ip in layer3**`
|
||||
```
|
||||
|
||||
|
||||
```
|
||||
`**[no] hardware counter feature ip out layer3**`
|
||||
```
|
||||
|
||||
|
||||
### hardware counter feature ip
|
||||
|
||||
|
||||
Use the **hardware counter feature ip** command to enable ingress
|
||||
and egress counters at Layer 3. The **no** and **default** forms of the command
|
||||
disables the feature. The feature is enabled by default.
|
||||
|
||||
|
||||
**Command Mode**
|
||||
|
||||
|
||||
Configuration mode
|
||||
|
||||
|
||||
**Command Syntax**
|
||||
|
||||
|
||||
**hardware counter feature ip in|out layer3**
|
||||
|
||||
|
||||
**no hardware counter feature ip in|out layer3**
|
||||
|
||||
|
||||
**default hardware counter feature in|out layer3**
|
||||
|
||||
|
||||
**Example**
|
||||
|
||||
|
||||
This example enables ingress and egress ip counters for Layer 3.
|
||||
```
|
||||
`**switch(config)# hardware counter feature in layer3**`
|
||||
```
|
||||
|
||||
|
||||
```
|
||||
`**switch(config)# hardware counter feature out layer3**`
|
||||
```
|
||||
|
||||
|
||||
## Show commands
|
||||
|
||||
|
||||
Use the [**show interfaces counters ip**](/um-eos/eos-ethernet-ports#xzx_RbdvgrfI6B) command to
|
||||
display IPv4, IPv6 packets, and octets.
|
||||
|
||||
|
||||
**Example**
|
||||
|
||||
|
||||
```
|
||||
`switch# **show interfaces counters ip**
|
||||
Interface IPv4InOctets IPv4InPkts IPv6InOctets IPv6InPkts
|
||||
Et1/1 0 0 0 0
|
||||
Et1/2 0 0 0 0
|
||||
Et1/3 0 0 0 0
|
||||
Et1/4 0 0 0 0
|
||||
...
|
||||
Interface IPv4OutOctets IPv4OutPkts IPv6OutOctets IPv6OutPkts
|
||||
Et1/1 0 0 0 0
|
||||
Et1/2 0 0 0 0
|
||||
Et1/3 0 0 0 0
|
||||
Et1/4 0 0 0 0
|
||||
...`
|
||||
```
|
||||
|
||||
|
||||
You can also query the output from the **show interfaces counters
|
||||
ip** command through snmp via the ARISTA-IP-MIB.
|
||||
|
||||
|
||||
To clear the IPv4 or IPv6 counters, use the [**clear
|
||||
counters**](/um-eos/eos-ethernet-ports#topic_dnd_1nm_vnb) command.
|
||||
|
||||
|
||||
**Example**
|
||||
```
|
||||
`switch# **clear counters**`
|
||||
```
|
||||
|
||||
|
||||
## Dedicated ARP Entry for TX IPv4 and IPv6 Counters
|
||||
|
||||
|
||||
IPv4/IPv6 egress Layer 3 (**hardware counter feature ip out layer3**)
|
||||
counting on DCS-7280SR, DCS-7280CR, and DCS-7500-R platforms work based on ARP entry of
|
||||
the next hop. By default, IPv4's next-hop and IPv6's next-hop resolve to the same MAC
|
||||
address and interface that shared the ARP entry.
|
||||
|
||||
|
||||
To differentiate the counters between IPv4 and IPv6, disable
|
||||
**arp** entry sharing with the following command:
|
||||
|
||||
|
||||
```
|
||||
`**ip hardware fib next-hop arp dedicated**`
|
||||
```
|
||||
|
||||
|
||||
|
||||
|
||||
Note: This command is required for IPv4 and IPv6 egress counters
|
||||
to operate on the DCS-7280SR, DCS-7280CR, and DCS-7500-R platforms.
|
||||
|
||||
|
||||
|
||||
|
||||
## Considerations
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
- Packet sizes greater than 9236 bytes are not counted by per-port IPv4 and IPv6 counters.
|
||||
|
||||
- Only the DCS-7260X3, DCS-7368, DCS-7300, DCS-7050SX3, DCS-7050CX3, DCS-7280SR,
|
||||
DCS-7280CR and DCS-7500-R platforms support the **hardware counter feature ip in** command.
|
||||
|
||||
- Only the DCS-7280SR, DCS-7280CR and DCS-7500-R platforms support the **hardware counter feature ip [in|out] layer3** command.
|
||||
@@ -0,0 +1,305 @@
|
||||
<!-- Source: https://www.arista.com/en/um-eos/eos-inter-vrf-local-route-leaking -->
|
||||
<!-- Scraped: 2026-03-06T20:43:28.363Z -->
|
||||
|
||||
# Inter-VRF Local Route Leaking
|
||||
|
||||
|
||||
Inter-VRF local route leaking allows the leaking of routes from one VRF (the source VRF) to
|
||||
another VRF (the destination VRF) on the same router.
|
||||
Inter-VRF routes can exist in any VRF (including the
|
||||
default VRF) on the system. Routes can be leaked using the
|
||||
following methods:
|
||||
|
||||
- Inter-VRF Local Route Leaking using BGP
|
||||
VPN
|
||||
|
||||
- Inter-VRF Local Route Leaking using VRF-leak
|
||||
Agent
|
||||
|
||||
|
||||
## Inter-VRF Local Route Leaking using BGP VPN
|
||||
|
||||
|
||||
Inter-VRF local route leaking allows the user to export and import routes from one VRF to another
|
||||
on the same device. This is implemented by exporting routes from a VRF to the local VPN table
|
||||
using the route target extended community list and importing the same route target extended
|
||||
community lists from the local VPN table into the target VRF. VRF route leaking is supported
|
||||
on VPN-IPv4, VPN-IPv6, and EVPN types.
|
||||
|
||||
|
||||
Figure 1. Inter-VRF Local Route Leaking using Local VPN Table
|
||||
|
||||
|
||||
### Accessing Shared Resources Across VPNs
|
||||
|
||||
|
||||
To access shared resources across VPNs, all the routes from the shared services VRF must be
|
||||
leaked into each of the VPN VRFs, and customer routes must be leaked into the shared
|
||||
services VRF for return traffic. Accessing shared resources allows the route target of the
|
||||
shared services VRF to be exported into all customer VRFs, and allows the shared services
|
||||
VRF to import route targets from customers A and B. The following figure shows how to
|
||||
provide customers, corresponding to multiple VPN domains, access to services like DHCP
|
||||
available in the shared VRF.
|
||||
|
||||
|
||||
Route leaking across the VRFs is supported
|
||||
on VPN-IPv4, VPN-IPv6, and EVPN.
|
||||
|
||||
|
||||
Figure 2. Accessing Shared Resources Across VPNs
|
||||
|
||||
|
||||
### Configuring Inter-VRF Local Route Leaking
|
||||
|
||||
|
||||
Inter-VRF local route leaking is configured using VPN-IPv4, VPN-IPv6, and EVPN. Prefixes can be
|
||||
exported and imported using any of the configured VPN types. Ensure that the same VPN
|
||||
type that is exported is used while importing.
|
||||
|
||||
|
||||
Leaking unicast IPv4 or IPv6 prefixes is supported and achieved by exporting prefixes locally to
|
||||
the VPN table and importing locally from the VPN table into the target VRF on the same
|
||||
device as shown in the figure titled **Inter-VRF Local Route Leaking using Local VPN
|
||||
Table** using the **route-target** command.
|
||||
|
||||
|
||||
Exporting or importing the routes to or from the EVPN table is accomplished with the following
|
||||
two methods:
|
||||
|
||||
- Using VXLAN for encapsulation
|
||||
|
||||
- Using MPLS for encapsulation
|
||||
|
||||
|
||||
#### Using VXLAN for Encapsulation
|
||||
|
||||
|
||||
To use VXLAN encapsulation type, make sure that VRF to VNI mapping is present and the interface
|
||||
status for the VXLAN interface is up. This is the default encapsulation type for
|
||||
EVPN.
|
||||
|
||||
|
||||
**Example**
|
||||
|
||||
|
||||
The configuration for VXLAN encapsulation type is as
|
||||
follows:
|
||||
```
|
||||
`switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **address-family evpn**
|
||||
switch(config-router-bgp-af)# **neighbor default encapsulation VXLAN next-hop-self source-interface Loopback0**
|
||||
switch(config)# **hardware tcam**
|
||||
switch(config-hw-tcam)# **system profile VXLAN-routing**
|
||||
switch(config-hw-tcam)# **interface VXLAN1**
|
||||
switch(config-hw-tcam-if-Vx1)# **VXLAN source-interface Loopback0**
|
||||
switch(config-hw-tcam-if-Vx1)# **VXLAN udp-port 4789**
|
||||
switch(config-hw-tcam-if-Vx1)# **VXLAN vrf vrf-blue vni 20001**
|
||||
switch(config-hw-tcam-if-Vx1)# **VXLAN vrf vrf-red vni 10001**`
|
||||
```
|
||||
|
||||
|
||||
#### Using MPLS for Encapsulation
|
||||
|
||||
|
||||
To use MPLS encapsulation type to export
|
||||
to the EVPN table, MPLS needs to be enabled globally on the device and
|
||||
the encapsulation method needs to be changed from default type, that
|
||||
is VXLAN to MPLS under the EVPN address-family sub-mode.
|
||||
|
||||
|
||||
**Example**
|
||||
```
|
||||
`switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **address-family evpn**
|
||||
switch(config-router-bgp-af)# **neighbor default encapsulation mpls next-hop-self source-interface Loopback0**`
|
||||
```
|
||||
|
||||
|
||||
### Route-Distinguisher
|
||||
|
||||
|
||||
Route-Distinguisher (RD) uniquely identifies routes from a particular VRF.
|
||||
Route-Distinguisher is configured for every VRF from which routes are exported from or
|
||||
imported into.
|
||||
|
||||
|
||||
The following commands are used to configure Route-Distinguisher for a VRF.
|
||||
|
||||
|
||||
```
|
||||
`switch(config-router-bgp)# **vrf vrf-services**
|
||||
switch(config-router-bgp-vrf-vrf-services)# **rd 1.0.0.1:1**
|
||||
|
||||
switch(config-router-bgp)# **vrf vrf-blue**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **rd 2.0.0.1:2**`
|
||||
```
|
||||
|
||||
|
||||
### Exporting Routes from a VRF
|
||||
|
||||
|
||||
Use the **route-target export** command to export routes from a VRF to the
|
||||
local VPN or EVPN table using the route target
|
||||
extended community list.
|
||||
|
||||
|
||||
**Examples**
|
||||
|
||||
- These commands export routes from
|
||||
**vrf-red** to the local VPN
|
||||
table.
|
||||
```
|
||||
`switch(config)# **service routing protocols model multi-agent**
|
||||
switch(config)# **mpls ip**
|
||||
switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-red**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **rd 1:1**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export vpn-ipv4 10:10**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export vpn-ipv6 10:20**`
|
||||
```
|
||||
|
||||
- These commands export routes from
|
||||
**vrf-red** to the EVPN
|
||||
table.
|
||||
```
|
||||
`switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-red**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **rd 1:1**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export evpn 10:1**`
|
||||
```
|
||||
|
||||
|
||||
### Importing Routes into a VRF
|
||||
|
||||
|
||||
Use the **route-target import** command to import the exported routes from
|
||||
the local VPN or EVPN table to the target VRF
|
||||
using the route target extended community
|
||||
list.
|
||||
|
||||
|
||||
**Examples**
|
||||
|
||||
- These commands import routes from the VPN
|
||||
table to
|
||||
**vrf-blue**.
|
||||
```
|
||||
`switch(config)# **service routing protocols model multi-agent**
|
||||
switch(config)# **mpls ip**
|
||||
switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-blue**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **rd 2:2**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import vpn-ipv4 10:10**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import vpn-ipv6 10:20**`
|
||||
```
|
||||
|
||||
- These commands import routes from the EVPN
|
||||
table to
|
||||
**vrf-blue**.
|
||||
```
|
||||
`switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-blue**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **rd 2:2**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import evpn 10:1**`
|
||||
```
|
||||
|
||||
|
||||
### Exporting and Importing Routes using Route
|
||||
Map
|
||||
|
||||
|
||||
To manage VRF route leaking, control the export and import prefixes with route-map export or
|
||||
import commands. The route map is effective only if the VRF or the VPN
|
||||
paths are already candidates for export or import. The route-target
|
||||
export or import commandmust be configured first. Setting BGP
|
||||
attributes using route maps is effective only on the export end.
|
||||
|
||||
|
||||
Note: Prefixes that are leaked are not re-exported to the VPN table from the target VRF.
|
||||
|
||||
**Examples**
|
||||
|
||||
- These commands export routes from
|
||||
**vrf-red** to the local VPN
|
||||
table.
|
||||
```
|
||||
`switch(config)# **service routing protocols model multi-agent**
|
||||
switch(config)# **mpls ip**
|
||||
switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-red**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **rd 1:1**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export vpn-ipv4 10:10**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export vpn-ipv6 10:20**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export vpn-ipv4 route-map EXPORT_V4_ROUTES_T0_VPN_TABLE**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export vpn-ipv6 route-map EXPORT_V6_ROUTES_T0_VPN_TABLE**`
|
||||
```
|
||||
|
||||
- These commands export routes to from
|
||||
**vrf-red** to the EVPN
|
||||
table.
|
||||
```
|
||||
`switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-red**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **rd 1:1**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export evpn 10:1**
|
||||
switch(config-router-bgp-vrf-vrf-red)# **route-target export evpn route-map EXPORT_ROUTES_T0_EVPN_TABLE**`
|
||||
```
|
||||
|
||||
- These commands import routes from the VPN table to
|
||||
**vrf-blue**.
|
||||
```
|
||||
`switch(config)# **service routing protocols model multi-agent**
|
||||
switch(config)# **mpls ip**
|
||||
switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-blue**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **rd 1:1**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import vpn-ipv4 10:10**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import vpn-ipv6 10:20**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import vpn-ipv4 route-map IMPORT_V4_ROUTES_VPN_TABLE**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import vpn-ipv6 route-map IMPORT_V6_ROUTES_VPN_TABLE**`
|
||||
```
|
||||
|
||||
- These commands import routes from the EVPN table to
|
||||
**vrf-blue**.
|
||||
```
|
||||
`switch(config)# **router bgp 65001**
|
||||
switch(config-router-bgp)# **vrf vrf-blue**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **rd 2:2**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import evpn 10:1**
|
||||
switch(config-router-bgp-vrf-vrf-blue)# **route-target import evpn route-map IMPORT_ROUTES_FROM_EVPN_TABLE**`
|
||||
```
|
||||
|
||||
|
||||
## Inter-VRF Local Route Leaking using VRF-leak
|
||||
Agent
|
||||
|
||||
|
||||
Inter-VRF local route leaking allows routes to leak from one VRF to another using a route
|
||||
map as a VRF-leak agent. VRFs are leaked based on the preferences assigned to each
|
||||
VRF.
|
||||
|
||||
|
||||
### Configuring Route Maps
|
||||
|
||||
|
||||
To leak routes from one VRF to another using a route map, use the [router general](/um-eos/eos-evpn-and-vcs-commands#xx1351777) command to enter Router-General
|
||||
Configuration Mode, then enter the VRF submode for the destination VRF, and use the
|
||||
[leak routes](/um-eos/eos-evpn-and-vcs-commands#reference_g2h_2z3_hwb) command to specify the source
|
||||
VRF and the route map to be used. Routes in the source VRF that match the policy in the
|
||||
route map will then be considered for leaking into the configuration-mode VRF. If two or
|
||||
more policies specify leaking the same prefix to the same destination VRF, the route
|
||||
with a higher (post-set-clause) distance and preference is chosen.
|
||||
|
||||
|
||||
**Example**
|
||||
|
||||
|
||||
These commands configure a route map to leak routes from **VRF1**
|
||||
to **VRF2** using route map
|
||||
**RM1**.
|
||||
```
|
||||
`switch(config)# **router general**
|
||||
switch(config-router-general)# **vrf VRF2**
|
||||
switch(config-router-general-vrf-VRF2)# **leak routes source-vrf VRF1 subscribe-policy RM1**
|
||||
switch(config-router-general-vrf-VRF2)#`
|
||||
```
|
||||
File diff suppressed because it is too large
Load Diff
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,82 @@
|
||||
<!-- Source: https://www.arista.com/en/um-eos/eos-static-inter-vrf-route -->
|
||||
<!-- Scraped: 2026-03-06T20:43:17.977Z -->
|
||||
|
||||
# Static Inter-VRF Route
|
||||
|
||||
|
||||
The Static Inter-VRF Route feature adds support for static inter-VRF routes. This enables the configuration of routes to destinations in one ingress VRF with an ability to specify a next-hop in a different egress VRF through a static configuration.
|
||||
|
||||
|
||||
You can configure static inter-VRF routes in default and non-default VRFs. A different
|
||||
egress VRF is achieved by “tagging” the **next-hop** or **forwarding
|
||||
via** with a reference to an egress VRF (different from the source
|
||||
VRF) in which that next-hop should be evaluated. Static inter-VRF routes
|
||||
with ECMP next-hop sets in the same egress VRF or heterogenous egress VRFs
|
||||
can be specified.
|
||||
|
||||
|
||||
The Static Inter-VRF Route feature is independent and complementary to other mechanisms that can be used to setup local inter-VRF routes. The other supported mechanisms in EOS and the broader use-cases they support are documented here:
|
||||
|
||||
- [Inter-VRF Local Route Leaking using BGP VPN](/um-eos/eos-inter-vrf-local-route-leaking#xx1348142)
|
||||
|
||||
- [Inter-VRF Local Route Leaking using VRF-leak Agent](/um-eos/eos-inter-vrf-local-route-leaking#xx1346287)
|
||||
|
||||
|
||||
## Configuration
|
||||
|
||||
|
||||
The configuration to setup static-Inter VRF routes in an ingress (source) VRF to forward IP traffic to a different egress (target) VRF can be done in the following modes:
|
||||
|
||||
- This command creates a static route in one ingress VRF that points to a next-hop
|
||||
in a different egress VRF.
|
||||
ip | ipv6
|
||||
route [vrf
|
||||
vrf-name
|
||||
destination-prefix [egress-vrf
|
||||
egress-next-hop-vrf-name]
|
||||
next-hop]
|
||||
|
||||
|
||||
## Show Commands
|
||||
|
||||
|
||||
Use the **show ip route vrf** to display the egress VRF name if it
|
||||
differs from the source VRF.
|
||||
|
||||
|
||||
**Example**
|
||||
```
|
||||
`switch# **show ip route vrf vrf1**
|
||||
|
||||
VRF: vrf1
|
||||
Codes: C - connected, S - static, K - kernel,
|
||||
O - OSPF, IA - OSPF inter area, E1 - OSPF external type 1,
|
||||
E2 - OSPF external type 2, N1 - OSPF NSSA external type 1,
|
||||
N2 - OSPF NSSA external type2, B - BGP, B I - iBGP, B E - eBGP,
|
||||
R - RIP, I L1 - IS-IS level 1, I L2 - IS-IS level 2,
|
||||
O3 - OSPFv3, A B - BGP Aggregate, A O - OSPF Summary,
|
||||
NG - Nexthop Group Static Route, V - VXLAN Control Service,
|
||||
DH - DHCP client installed default route, M - Martian,
|
||||
DP - Dynamic Policy Route, L - VRF Leaked
|
||||
|
||||
Gateway of last resort is not set
|
||||
|
||||
S 1.0.1.0/24 [1/0] via 1.0.0.2, Vlan2180 (egress VRF default)
|
||||
S 1.0.7.0/24 [1/0] via 1.0.6.2, Vlan2507 (egress VRF vrf3)`
|
||||
```
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
## Limitations
|
||||
|
||||
|
||||
|
||||
|
||||
|
||||
- For bidirectional traffic to work correctly between a pair of VRFs, static inter-VRF
|
||||
routes in both VRFs must be configured.
|
||||
|
||||
- Static Inter-VRF routing is supported only in multi-agent routing protocol mode.
|
||||
File diff suppressed because it is too large
Load Diff
@@ -0,0 +1,275 @@
|
||||
# Ashburn Validator Relay — Full Traffic Redirect
|
||||
|
||||
## Overview
|
||||
|
||||
All validator traffic (gossip, repair, TVU, TPU) enters and exits from
|
||||
`137.239.194.65` (laconic-was-sw01, Ashburn). Peers see the validator as an
|
||||
Ashburn node. This improves repair peer count and slot catchup rate by reducing
|
||||
RTT to the TeraSwitch/Pittsburgh cluster from ~30ms (direct Miami) to ~5ms
|
||||
(Ashburn).
|
||||
|
||||
Supersedes the previous TVU-only shred relay (see `tvu-shred-relay.md`).
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
OUTBOUND (validator → peers)
|
||||
agave-validator (kind pod, ports 8001, 9000-9025)
|
||||
↓ Docker bridge → host FORWARD chain
|
||||
biscayne host (186.233.184.235)
|
||||
↓ mangle PREROUTING: fwmark 100 on sport 8001,9000-9025 from 172.20.0.0/16
|
||||
↓ nat POSTROUTING: SNAT → src 137.239.194.65
|
||||
↓ policy route: fwmark 100 → table ashburn → via 169.254.7.6 dev doublezero0
|
||||
laconic-mia-sw01 (209.42.167.133, Miami)
|
||||
↓ traffic-policy VALIDATOR-OUTBOUND: src 137.239.194.65 → nexthop 172.16.1.188
|
||||
↓ backbone Et4/1 (25.4ms)
|
||||
laconic-was-sw01 Et4/1 (Ashburn)
|
||||
↓ default route via 64.92.84.80 out Et1/1
|
||||
Internet (peers see src 137.239.194.65)
|
||||
|
||||
INBOUND (peers → validator)
|
||||
Solana peers → 137.239.194.65:8001,9000-9025
|
||||
↓ internet routing to was-sw01
|
||||
laconic-was-sw01 Et1/1 (Ashburn)
|
||||
↓ traffic-policy VALIDATOR-RELAY: ASIC redirect, line rate
|
||||
↓ nexthop 172.16.1.189 via Et4/1 backbone (25.4ms)
|
||||
laconic-mia-sw01 Et4/1 (Miami)
|
||||
↓ L3 forward → biscayne via doublezero0 GRE or ISP routing
|
||||
biscayne (186.233.184.235)
|
||||
↓ nat PREROUTING: DNAT dst 137.239.194.65:* → 172.20.0.2:* (kind node)
|
||||
↓ Docker bridge → validator pod
|
||||
agave-validator
|
||||
```
|
||||
|
||||
RPC traffic (port 8899) is NOT relayed — clients connect directly to biscayne.
|
||||
|
||||
## Switch Config: laconic-was-sw01
|
||||
|
||||
SSH: `install@137.239.200.198`
|
||||
|
||||
### Pre-change
|
||||
|
||||
```
|
||||
configure checkpoint save pre-validator-relay
|
||||
```
|
||||
|
||||
Rollback: `rollback running-config checkpoint pre-validator-relay` then `write memory`.
|
||||
|
||||
### Config session with auto-revert
|
||||
|
||||
```
|
||||
configure session validator-relay
|
||||
|
||||
! Loopback for 137.239.194.65 (do NOT touch Loopback100 which has .64)
|
||||
interface Loopback101
|
||||
ip address 137.239.194.65/32
|
||||
|
||||
! ACL covering all validator ports
|
||||
ip access-list VALIDATOR-RELAY-ACL
|
||||
10 permit udp any any eq 8001
|
||||
20 permit udp any any range 9000 9025
|
||||
30 permit tcp any any eq 8001
|
||||
|
||||
! Traffic-policy: ASIC redirect to backbone (mia-sw01)
|
||||
traffic-policy VALIDATOR-RELAY
|
||||
match VALIDATOR-RELAY-ACL
|
||||
set nexthop 172.16.1.189
|
||||
|
||||
! Replace old SHRED-RELAY on Et1/1
|
||||
interface Ethernet1/1
|
||||
no traffic-policy input SHRED-RELAY
|
||||
traffic-policy input VALIDATOR-RELAY
|
||||
|
||||
! system-rule overriding-action redirect (already present from SHRED-RELAY)
|
||||
|
||||
show session-config diffs
|
||||
commit timer 00:05:00
|
||||
```
|
||||
|
||||
After verification: `configure session validator-relay commit` then `write memory`.
|
||||
|
||||
### Cleanup (after stable)
|
||||
|
||||
Old SHRED-RELAY policy and ACL can be removed once VALIDATOR-RELAY is confirmed:
|
||||
|
||||
```
|
||||
configure session cleanup-shred-relay
|
||||
no traffic-policy SHRED-RELAY
|
||||
no ip access-list SHRED-RELAY-ACL
|
||||
show session-config diffs
|
||||
commit
|
||||
write memory
|
||||
```
|
||||
|
||||
## Switch Config: laconic-mia-sw01
|
||||
|
||||
### Pre-flight checks
|
||||
|
||||
Before applying config, verify:
|
||||
|
||||
1. Which EOS interface terminates the doublezero0 GRE from biscayne
|
||||
(endpoint 209.42.167.133). Check with `show interfaces tunnel` or
|
||||
`show ip interface brief | include Tunnel`.
|
||||
|
||||
2. Whether `system-rule overriding-action redirect` is already configured.
|
||||
Check with `show running-config | include system-rule`.
|
||||
|
||||
3. Whether EOS traffic-policy works on tunnel interfaces. If not, apply on
|
||||
the physical interface where GRE packets arrive (likely Et<X> facing
|
||||
biscayne's ISP network or the DZ infrastructure).
|
||||
|
||||
### Config session
|
||||
|
||||
```
|
||||
configure checkpoint save pre-validator-outbound
|
||||
|
||||
configure session validator-outbound
|
||||
|
||||
! ACL matching outbound validator traffic (source = Ashburn IP)
|
||||
ip access-list VALIDATOR-OUTBOUND-ACL
|
||||
10 permit ip 137.239.194.65/32 any
|
||||
|
||||
! Redirect to was-sw01 via backbone
|
||||
traffic-policy VALIDATOR-OUTBOUND
|
||||
match VALIDATOR-OUTBOUND-ACL
|
||||
set nexthop 172.16.1.188
|
||||
|
||||
! Apply on the interface where biscayne GRE traffic arrives
|
||||
! Replace Tunnel<X> with the actual interface from pre-flight check #1
|
||||
interface Tunnel<X>
|
||||
traffic-policy input VALIDATOR-OUTBOUND
|
||||
|
||||
! Add system-rule if not already present (pre-flight check #2)
|
||||
system-rule overriding-action redirect
|
||||
|
||||
show session-config diffs
|
||||
commit timer 00:05:00
|
||||
```
|
||||
|
||||
After verification: commit + `write memory`.
|
||||
|
||||
## Host Config: biscayne
|
||||
|
||||
Automated via ansible playbook `playbooks/ashburn-validator-relay.yml`.
|
||||
|
||||
### Manual equivalent
|
||||
|
||||
```bash
|
||||
# 1. Accept packets destined for 137.239.194.65
|
||||
sudo ip addr add 137.239.194.65/32 dev lo
|
||||
|
||||
# 2. Inbound DNAT to kind node (172.20.0.2)
|
||||
sudo iptables -t nat -A PREROUTING -p udp -d 137.239.194.65 --dport 8001 \
|
||||
-j DNAT --to-destination 172.20.0.2:8001
|
||||
sudo iptables -t nat -A PREROUTING -p tcp -d 137.239.194.65 --dport 8001 \
|
||||
-j DNAT --to-destination 172.20.0.2:8001
|
||||
sudo iptables -t nat -A PREROUTING -p udp -d 137.239.194.65 --dport 9000:9025 \
|
||||
-j DNAT --to-destination 172.20.0.2
|
||||
|
||||
# 3. Outbound: mark validator traffic
|
||||
sudo iptables -t mangle -A PREROUTING -s 172.20.0.0/16 -p udp --sport 8001 \
|
||||
-j MARK --set-mark 100
|
||||
sudo iptables -t mangle -A PREROUTING -s 172.20.0.0/16 -p udp --sport 9000:9025 \
|
||||
-j MARK --set-mark 100
|
||||
sudo iptables -t mangle -A PREROUTING -s 172.20.0.0/16 -p tcp --sport 8001 \
|
||||
-j MARK --set-mark 100
|
||||
|
||||
# 4. Outbound: SNAT to Ashburn IP (INSERT before Docker MASQUERADE)
|
||||
sudo iptables -t nat -I POSTROUTING 1 -m mark --mark 100 \
|
||||
-j SNAT --to-source 137.239.194.65
|
||||
|
||||
# 5. Policy routing table
|
||||
echo "100 ashburn" | sudo tee -a /etc/iproute2/rt_tables
|
||||
sudo ip rule add fwmark 100 table ashburn
|
||||
sudo ip route add default via 169.254.7.6 dev doublezero0 table ashburn
|
||||
|
||||
# 6. Persist
|
||||
sudo netfilter-persistent save
|
||||
# ip rule + ip route persist via /etc/network/if-up.d/ashburn-routing
|
||||
```
|
||||
|
||||
### Docker NAT port preservation
|
||||
|
||||
**Must verify before going live:** Docker masquerade must preserve source ports
|
||||
for kind's hostNetwork pods. If Docker rewrites the source port, the mangle
|
||||
PREROUTING match on `--sport 8001,9000-9025` will miss traffic.
|
||||
|
||||
Test: `tcpdump -i br-cf46a62ab5b2 -nn 'udp src port 8001'` — if you see
|
||||
packets with sport 8001 from 172.20.0.2, port preservation works.
|
||||
|
||||
If Docker does NOT preserve ports, the mark must be set inside the kind node
|
||||
container (on the pod's veth) rather than on the host.
|
||||
|
||||
## Execution Order
|
||||
|
||||
1. **was-sw01**: checkpoint → config session with 5min auto-revert → verify counters → commit
|
||||
2. **biscayne**: add 137.239.194.65/32 to lo, add inbound DNAT rules
|
||||
3. **Verify inbound**: `ping 137.239.194.65` from external host, check DNAT counters
|
||||
4. **mia-sw01**: pre-flight checks → config session with 5min auto-revert → commit
|
||||
5. **biscayne**: add outbound fwmark + policy routing + SNAT rules
|
||||
6. **Test outbound**: from biscayne, send UDP from port 8001, verify src 137.239.194.65 on was-sw01
|
||||
7. **Verify**: traffic-policy counters on both switches, iptables hit counts on biscayne
|
||||
8. **Restart validator** if needed (gossip should auto-refresh, but restart ensures clean state)
|
||||
9. **was-sw01 + mia-sw01**: `write memory` to persist
|
||||
10. **Cleanup**: remove old SHRED-RELAY and 64.92.84.81:20000 DNAT after stable
|
||||
|
||||
## Verification
|
||||
|
||||
1. `show traffic-policy counters` on was-sw01 — VALIDATOR-RELAY-ACL matches
|
||||
2. `show traffic-policy counters` on mia-sw01 — VALIDATOR-OUTBOUND-ACL matches
|
||||
3. `sudo iptables -t nat -L -v -n` on biscayne — DNAT and SNAT hit counts
|
||||
4. `sudo iptables -t mangle -L -v -n` on biscayne — fwmark hit counts
|
||||
5. `ip rule show` on biscayne — fwmark 100 lookup ashburn
|
||||
6. Validator gossip ContactInfo shows 137.239.194.65 for ALL addresses (gossip, repair, TVU, TPU)
|
||||
7. Repair peer count increases (target: 20+ peers)
|
||||
8. Slot catchup rate improves from ~0.9 toward ~2.5 slots/sec
|
||||
9. `traceroute --sport=8001 <remote_peer>` from biscayne routes via doublezero0/was-sw01
|
||||
|
||||
## Rollback
|
||||
|
||||
### biscayne
|
||||
|
||||
```bash
|
||||
sudo ip addr del 137.239.194.65/32 dev lo
|
||||
sudo iptables -t nat -D PREROUTING -p udp -d 137.239.194.65 --dport 8001 -j DNAT --to-destination 172.20.0.2:8001
|
||||
sudo iptables -t nat -D PREROUTING -p tcp -d 137.239.194.65 --dport 8001 -j DNAT --to-destination 172.20.0.2:8001
|
||||
sudo iptables -t nat -D PREROUTING -p udp -d 137.239.194.65 --dport 9000:9025 -j DNAT --to-destination 172.20.0.2
|
||||
sudo iptables -t mangle -D PREROUTING -s 172.20.0.0/16 -p udp --sport 8001 -j MARK --set-mark 100
|
||||
sudo iptables -t mangle -D PREROUTING -s 172.20.0.0/16 -p udp --sport 9000:9025 -j MARK --set-mark 100
|
||||
sudo iptables -t mangle -D PREROUTING -s 172.20.0.0/16 -p tcp --sport 8001 -j MARK --set-mark 100
|
||||
sudo iptables -t nat -D POSTROUTING -m mark --mark 100 -j SNAT --to-source 137.239.194.65
|
||||
sudo ip rule del fwmark 100 table ashburn
|
||||
sudo ip route del default table ashburn
|
||||
sudo netfilter-persistent save
|
||||
```
|
||||
|
||||
### was-sw01
|
||||
|
||||
```
|
||||
rollback running-config checkpoint pre-validator-relay
|
||||
write memory
|
||||
```
|
||||
|
||||
### mia-sw01
|
||||
|
||||
```
|
||||
rollback running-config checkpoint pre-validator-outbound
|
||||
write memory
|
||||
```
|
||||
|
||||
## Key Details
|
||||
|
||||
| Item | Value |
|
||||
|------|-------|
|
||||
| Ashburn relay IP | `137.239.194.65` (Loopback101 on was-sw01) |
|
||||
| Ashburn LAN block | `137.239.194.64/29` on was-sw01 Et1/1 |
|
||||
| Biscayne IP | `186.233.184.235` |
|
||||
| Kind node IP | `172.20.0.2` (Docker bridge br-cf46a62ab5b2) |
|
||||
| Validator ports | 8001 (gossip), 9000-9025 (TVU/repair/TPU) |
|
||||
| Excluded ports | 8899 (RPC), 8900 (WebSocket) — direct to biscayne |
|
||||
| GRE tunnel | doublezero0: 169.254.7.7 ↔ 169.254.7.6, remote 209.42.167.133 |
|
||||
| Backbone | was-sw01 Et4/1 172.16.1.188/31 ↔ mia-sw01 Et4/1 172.16.1.189/31 |
|
||||
| Policy routing table | 100 ashburn |
|
||||
| Fwmark | 100 |
|
||||
| was-sw01 SSH | `install@137.239.200.198` |
|
||||
| EOS version | 4.34.0F |
|
||||
@@ -0,0 +1,416 @@
|
||||
# Blue-Green Upgrades for Biscayne
|
||||
|
||||
Zero-downtime upgrade procedures for the agave-stack deployment on biscayne.
|
||||
Uses ZFS clones for instant data duplication, Caddy health-check routing for
|
||||
traffic shifting, and k8s native sidecars for independent container upgrades.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Caddy ingress (biscayne.vaasl.io)
|
||||
├── upstream A: localhost:8899 ← health: /health
|
||||
└── upstream B: localhost:8897 ← health: /health
|
||||
│
|
||||
┌─────────────────┴──────────────────┐
|
||||
│ kind cluster │
|
||||
│ │
|
||||
│ Deployment A Deployment B │
|
||||
│ ┌─────────────┐ ┌─────────────┐ │
|
||||
│ │ agave :8899 │ │ agave :8897 │ │
|
||||
│ │ doublezerod │ │ doublezerod │ │
|
||||
│ └──────┬──────┘ └──────┬──────┘ │
|
||||
└─────────┼─────────────────┼─────────┘
|
||||
│ │
|
||||
ZFS dataset A ZFS clone B
|
||||
(original) (instant CoW copy)
|
||||
```
|
||||
|
||||
Both deployments run in the same kind cluster with `hostNetwork: true`.
|
||||
Caddy active health checks route traffic to whichever deployment has a
|
||||
healthy `/health` endpoint.
|
||||
|
||||
## Storage Layout
|
||||
|
||||
| Data | Path | Type | Survives restart? |
|
||||
|------|------|------|-------------------|
|
||||
| Ledger | `/srv/solana/ledger` | ZFS zvol (xfs) | Yes |
|
||||
| Snapshots | `/srv/solana/snapshots` | ZFS zvol (xfs) | Yes |
|
||||
| Accounts | `/srv/solana/ramdisk/accounts` | `/dev/ram0` (xfs) | Until host reboot |
|
||||
| Validator config | `/srv/deployments/agave/data/validator-config` | ZFS | Yes |
|
||||
| DZ config | `/srv/deployments/agave/data/doublezero-config` | ZFS | Yes |
|
||||
|
||||
The ZFS zvol `biscayne/DATA/volumes/solana` backs `/srv/solana` (ledger, snapshots).
|
||||
The ramdisk at `/dev/ram0` holds accounts — it's a block device, not tmpfs, so it
|
||||
survives process restarts but not host reboots.
|
||||
|
||||
---
|
||||
|
||||
## Procedure 1: DoubleZero Binary Upgrade (zero downtime, single pod)
|
||||
|
||||
The GRE tunnel (`doublezero0`) and BGP routes live in kernel space. They persist
|
||||
across doublezerod process restarts. Upgrading the DZ binary does not require
|
||||
tearing down the tunnel or restarting the validator.
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- doublezerod is defined as a k8s native sidecar (`spec.initContainers` with
|
||||
`restartPolicy: Always`). See [Required Changes](#required-changes) below.
|
||||
- k8s 1.29+ (biscayne runs 1.35.1)
|
||||
|
||||
### Steps
|
||||
|
||||
1. Build or pull the new doublezero container image.
|
||||
|
||||
2. Patch the pod's sidecar image:
|
||||
```bash
|
||||
kubectl -n <ns> patch pod <pod> --type='json' -p='[
|
||||
{"op": "replace", "path": "/spec/initContainers/0/image",
|
||||
"value": "laconicnetwork/doublezero:new-version"}
|
||||
]'
|
||||
```
|
||||
|
||||
3. Only the doublezerod container restarts. The agave container is unaffected.
|
||||
The GRE tunnel interface and BGP routes remain in the kernel throughout.
|
||||
|
||||
4. Verify:
|
||||
```bash
|
||||
kubectl -n <ns> exec <pod> -c doublezerod -- doublezero --version
|
||||
kubectl -n <ns> exec <pod> -c doublezerod -- doublezero status
|
||||
ip route | grep doublezero0 # routes still present
|
||||
```
|
||||
|
||||
### Rollback
|
||||
|
||||
Patch the image back to the previous version. Same process, same zero downtime.
|
||||
|
||||
---
|
||||
|
||||
## Procedure 2: Agave Version Upgrade (zero RPC downtime, blue-green)
|
||||
|
||||
Agave is the main container and must be restarted for a version change. To maintain
|
||||
zero RPC downtime, we run two deployments simultaneously and let Caddy shift traffic
|
||||
based on health checks.
|
||||
|
||||
### Prerequisites
|
||||
|
||||
- Caddy ingress configured with dual upstreams and active health checks
|
||||
- A parameterized spec.yml that accepts alternate ports and volume paths
|
||||
- ZFS snapshot/clone scripts
|
||||
|
||||
### Steps
|
||||
|
||||
#### Phase 1: Prepare (no downtime, no risk)
|
||||
|
||||
1. **ZFS snapshot** for rollback safety:
|
||||
```bash
|
||||
zfs snapshot -r biscayne/DATA@pre-upgrade-$(date +%Y%m%d)
|
||||
```
|
||||
|
||||
2. **ZFS clone** the validator volumes:
|
||||
```bash
|
||||
zfs clone biscayne/DATA/volumes/solana@pre-upgrade-$(date +%Y%m%d) \
|
||||
biscayne/DATA/volumes/solana-blue
|
||||
```
|
||||
This is instant (copy-on-write). No additional storage until writes diverge.
|
||||
|
||||
3. **Clone the ramdisk accounts** (not on ZFS):
|
||||
```bash
|
||||
mkdir -p /srv/solana-blue/ramdisk/accounts
|
||||
cp -a /srv/solana/ramdisk/accounts/* /srv/solana-blue/ramdisk/accounts/
|
||||
```
|
||||
This is the slow step — 460GB on ramdisk. Consider `rsync` with `--inplace`
|
||||
to minimize copy time, or investigate whether the ramdisk can move to a ZFS
|
||||
dataset for instant cloning in future deployments.
|
||||
|
||||
4. **Build or pull** the new agave container image.
|
||||
|
||||
#### Phase 2: Start blue deployment (no downtime)
|
||||
|
||||
5. **Create Deployment B** in the same kind cluster, pointing at cloned volumes,
|
||||
with RPC on port 8897:
|
||||
```bash
|
||||
# Apply the blue deployment manifest (parameterized spec)
|
||||
kubectl apply -f deployment/k8s-manifests/agave-blue.yaml
|
||||
```
|
||||
|
||||
6. **Deployment B catches up.** It starts from the snapshot point and replays.
|
||||
Monitor progress:
|
||||
```bash
|
||||
kubectl -n <ns> exec <blue-pod> -c agave-validator -- \
|
||||
solana -u http://127.0.0.1:8897 slot
|
||||
```
|
||||
|
||||
7. **Validate** the new version works:
|
||||
- RPC responds: `curl -sf http://localhost:8897/health`
|
||||
- Correct version: `kubectl -n <ns> exec <blue-pod> -c agave-validator -- agave-validator --version`
|
||||
- doublezerod connected (if applicable)
|
||||
|
||||
Take as long as needed. Deployment A is still serving all traffic.
|
||||
|
||||
#### Phase 3: Traffic shift (zero downtime)
|
||||
|
||||
8. **Caddy routes traffic to B.** Once B's `/health` returns 200, Caddy's active
|
||||
health check automatically starts routing to it. Alternatively, update the
|
||||
Caddy upstream config to prefer B.
|
||||
|
||||
9. **Verify** B is serving live traffic:
|
||||
```bash
|
||||
curl -sf https://biscayne.vaasl.io/health
|
||||
# Check Caddy access logs for requests hitting port 8897
|
||||
```
|
||||
|
||||
#### Phase 4: Cleanup
|
||||
|
||||
10. **Stop Deployment A:**
|
||||
```bash
|
||||
kubectl -n <ns> delete deployment agave-green
|
||||
```
|
||||
|
||||
11. **Reconfigure B to use standard port** (8899) if desired, or update Caddy
|
||||
to only route to 8897.
|
||||
|
||||
12. **Clean up ZFS clone** (or keep as rollback):
|
||||
```bash
|
||||
zfs destroy biscayne/DATA/volumes/solana-blue
|
||||
```
|
||||
|
||||
### Rollback
|
||||
|
||||
At any point before Phase 4:
|
||||
- Deployment A is untouched and still serving traffic (or can be restarted)
|
||||
- Delete Deployment B: `kubectl -n <ns> delete deployment agave-blue`
|
||||
- Destroy the ZFS clone: `zfs destroy biscayne/DATA/volumes/solana-blue`
|
||||
|
||||
After Phase 4 (A already stopped):
|
||||
- `zfs rollback` to restore original data
|
||||
- Redeploy A with old image
|
||||
|
||||
---
|
||||
|
||||
## Required Changes to agave-stack
|
||||
|
||||
### 1. Move doublezerod to native sidecar
|
||||
|
||||
In the pod spec generation (laconic-so or compose override), doublezerod must be
|
||||
defined as a native sidecar container instead of a regular container:
|
||||
|
||||
```yaml
|
||||
spec:
|
||||
initContainers:
|
||||
- name: doublezerod
|
||||
image: laconicnetwork/doublezero:local
|
||||
restartPolicy: Always # makes it a native sidecar
|
||||
securityContext:
|
||||
privileged: true
|
||||
capabilities:
|
||||
add: [NET_ADMIN]
|
||||
env:
|
||||
- name: DOUBLEZERO_RPC_ENDPOINT
|
||||
value: https://api.mainnet-beta.solana.com
|
||||
volumeMounts:
|
||||
- name: doublezero-config
|
||||
mountPath: /root/.config/doublezero
|
||||
containers:
|
||||
- name: agave-validator
|
||||
image: laconicnetwork/agave:local
|
||||
# ... existing config
|
||||
```
|
||||
|
||||
This change means:
|
||||
- doublezerod starts before agave and stays running
|
||||
- Patching the doublezerod image restarts only that container
|
||||
- agave can be restarted independently without affecting doublezerod
|
||||
|
||||
This requires a laconic-so change to support `initContainers` with `restartPolicy`
|
||||
in compose-to-k8s translation — or a post-deployment patch.
|
||||
|
||||
### 2. Caddy dual-upstream config
|
||||
|
||||
Add health-checked upstreams for both blue and green deployments:
|
||||
|
||||
```caddyfile
|
||||
biscayne.vaasl.io {
|
||||
reverse_proxy {
|
||||
to localhost:8899 localhost:8897
|
||||
|
||||
health_uri /health
|
||||
health_interval 5s
|
||||
health_timeout 3s
|
||||
|
||||
lb_policy first
|
||||
}
|
||||
}
|
||||
```
|
||||
|
||||
`lb_policy first` routes to the first healthy upstream. When only A is running,
|
||||
all traffic goes to :8899. When B comes up healthy, traffic shifts.
|
||||
|
||||
### 3. Parameterized deployment spec
|
||||
|
||||
Create a parameterized spec or kustomize overlay that accepts:
|
||||
- RPC port (8899 vs 8897)
|
||||
- Volume paths (original vs ZFS clone)
|
||||
- Deployment name suffix (green vs blue)
|
||||
|
||||
### 4. Delete DaemonSet workaround
|
||||
|
||||
Remove `deployment/k8s-manifests/doublezero-daemonset.yaml` from agave-stack.
|
||||
|
||||
### 5. Fix container DZ identity
|
||||
|
||||
Copy the registered identity into the container volume:
|
||||
```bash
|
||||
sudo cp /home/solana/.config/doublezero/id.json \
|
||||
/srv/deployments/agave/data/doublezero-config/id.json
|
||||
```
|
||||
|
||||
### 6. Disable host systemd doublezerod
|
||||
|
||||
After the container sidecar is working:
|
||||
```bash
|
||||
sudo systemctl stop doublezerod
|
||||
sudo systemctl disable doublezerod
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Implementation Order
|
||||
|
||||
This is a spec-driven, test-driven plan. Each step produces a testable artifact.
|
||||
|
||||
### Step 1: Fix existing DZ bugs (no code changes to laconic-so)
|
||||
|
||||
Fixes BUG-1 through BUG-5 from [doublezero-status.md](doublezero-status.md).
|
||||
|
||||
**Spec:** Container doublezerod shows correct identity, connects to laconic-mia-sw01,
|
||||
host systemd doublezerod is disabled.
|
||||
|
||||
**Test:**
|
||||
```bash
|
||||
kubectl -n <ns> exec <pod> -c doublezerod -- doublezero address
|
||||
# assert: 3Bw6v7EruQvTwoY79h2QjQCs2KBQFzSneBdYUbcXK1Tr
|
||||
|
||||
kubectl -n <ns> exec <pod> -c doublezerod -- doublezero status
|
||||
# assert: BGP Session Up, laconic-mia-sw01
|
||||
|
||||
systemctl is-active doublezerod
|
||||
# assert: inactive
|
||||
```
|
||||
|
||||
**Changes:**
|
||||
- Copy `id.json` to container volume
|
||||
- Update `DOUBLEZERO_RPC_ENDPOINT` in spec.yml
|
||||
- Deploy with hostNetwork-enabled stack-orchestrator
|
||||
- Stop and disable host doublezerod
|
||||
- Delete DaemonSet manifest from agave-stack
|
||||
|
||||
### Step 2: Native sidecar for doublezerod
|
||||
|
||||
**Spec:** doublezerod image can be patched without restarting the agave container.
|
||||
GRE tunnel and routes persist across doublezerod restart.
|
||||
|
||||
**Test:**
|
||||
```bash
|
||||
# Record current agave container start time
|
||||
BEFORE=$(kubectl -n <ns> get pod <pod> -o jsonpath='{.status.containerStatuses[?(@.name=="agave-validator")].state.running.startedAt}')
|
||||
|
||||
# Patch DZ image
|
||||
kubectl -n <ns> patch pod <pod> --type='json' -p='[
|
||||
{"op":"replace","path":"/spec/initContainers/0/image","value":"laconicnetwork/doublezero:test"}
|
||||
]'
|
||||
|
||||
# Wait for DZ container to restart
|
||||
sleep 10
|
||||
|
||||
# Verify agave was NOT restarted
|
||||
AFTER=$(kubectl -n <ns> get pod <pod> -o jsonpath='{.status.containerStatuses[?(@.name=="agave-validator")].state.running.startedAt}')
|
||||
[ "$BEFORE" = "$AFTER" ] # assert: same start time
|
||||
|
||||
# Verify tunnel survived
|
||||
ip route | grep doublezero0 # assert: routes present
|
||||
```
|
||||
|
||||
**Changes:**
|
||||
- laconic-so: support `initContainers` with `restartPolicy: Always` in
|
||||
compose-to-k8s translation (or: define doublezerod as native sidecar in
|
||||
compose via `x-kubernetes-init-container` extension or equivalent)
|
||||
- Alternatively: post-deploy kubectl patch to move doublezerod to initContainers
|
||||
|
||||
### Step 3: Caddy dual-upstream routing
|
||||
|
||||
**Spec:** Caddy routes RPC traffic to whichever backend is healthy. Adding a second
|
||||
healthy backend on :8897 causes traffic to shift without configuration changes.
|
||||
|
||||
**Test:**
|
||||
```bash
|
||||
# Start a test HTTP server on :8897 with /health
|
||||
python3 -c "
|
||||
from http.server import HTTPServer, BaseHTTPRequestHandler
|
||||
class H(BaseHTTPRequestHandler):
|
||||
def do_GET(self):
|
||||
self.send_response(200); self.end_headers(); self.wfile.write(b'ok')
|
||||
HTTPServer(('', 8897), H).serve_forever()
|
||||
" &
|
||||
|
||||
# Verify Caddy discovers it
|
||||
sleep 10
|
||||
curl -sf https://biscayne.vaasl.io/health
|
||||
# assert: 200
|
||||
|
||||
kill %1
|
||||
```
|
||||
|
||||
**Changes:**
|
||||
- Update Caddy ingress config with dual upstreams and health checks
|
||||
|
||||
### Step 4: ZFS clone and blue-green tooling
|
||||
|
||||
**Spec:** A script creates a ZFS clone, starts a blue deployment on alternate ports
|
||||
using the cloned data, and the deployment catches up and becomes healthy.
|
||||
|
||||
**Test:**
|
||||
```bash
|
||||
# Run the clone + deploy script
|
||||
./scripts/blue-green-prepare.sh --target-version v2.2.1
|
||||
|
||||
# assert: ZFS clone exists
|
||||
zfs list biscayne/DATA/volumes/solana-blue
|
||||
|
||||
# assert: blue deployment exists and is catching up
|
||||
kubectl -n <ns> get deployment agave-blue
|
||||
|
||||
# assert: blue RPC eventually becomes healthy
|
||||
timeout 600 bash -c 'until curl -sf http://localhost:8897/health; do sleep 5; done'
|
||||
```
|
||||
|
||||
**Changes:**
|
||||
- `scripts/blue-green-prepare.sh` — ZFS snapshot, clone, deploy B
|
||||
- `scripts/blue-green-promote.sh` — tear down A, optional port swap
|
||||
- `scripts/blue-green-rollback.sh` — destroy B, restore A
|
||||
- Parameterized deployment spec (kustomize overlay or env-driven)
|
||||
|
||||
### Step 5: End-to-end upgrade test
|
||||
|
||||
**Spec:** Full upgrade cycle completes with zero dropped RPC requests.
|
||||
|
||||
**Test:**
|
||||
```bash
|
||||
# Start continuous health probe in background
|
||||
while true; do
|
||||
curl -sf -o /dev/null -w "%{http_code} %{time_total}\n" \
|
||||
https://biscayne.vaasl.io/health || echo "FAIL $(date)"
|
||||
sleep 0.5
|
||||
done > /tmp/health-probe.log &
|
||||
|
||||
# Execute full blue-green upgrade
|
||||
./scripts/blue-green-prepare.sh --target-version v2.2.1
|
||||
# wait for blue to sync...
|
||||
./scripts/blue-green-promote.sh
|
||||
|
||||
# Stop probe
|
||||
kill %1
|
||||
|
||||
# assert: no FAIL lines in probe log
|
||||
grep -c FAIL /tmp/health-probe.log
|
||||
# assert: 0
|
||||
```
|
||||
@@ -0,0 +1,85 @@
|
||||
# Bug: Ashburn Relay — 137.239.194.65 Not Routable from Public Internet
|
||||
|
||||
## Summary
|
||||
|
||||
`--gossip-host 137.239.194.65` correctly advertises the Ashburn relay IP in
|
||||
ContactInfo for all sockets (gossip, TVU, repair, TPU). However, 137.239.194.65
|
||||
is a DoubleZero overlay IP (137.239.192.0/19, IS-IS only) that is NOT announced
|
||||
via BGP to the public internet. Public peers cannot route to it, so TVU shreds,
|
||||
repair requests, and TPU traffic never arrive at was-sw01.
|
||||
|
||||
## Evidence
|
||||
|
||||
- Gossip traffic arrives on `doublezero0` interface:
|
||||
```
|
||||
doublezero0 In IP 64.130.58.70.8001 > 137.239.194.65.8001: UDP, length 132
|
||||
```
|
||||
- Zero TVU/repair traffic arrives:
|
||||
```
|
||||
tcpdump -i doublezero0 'dst host 137.239.194.65 and udp and not port 8001'
|
||||
0 packets captured
|
||||
```
|
||||
- ContactInfo correctly advertises all sockets on 137.239.194.65:
|
||||
```json
|
||||
{
|
||||
"gossip": "137.239.194.65:8001",
|
||||
"tvu": "137.239.194.65:9000",
|
||||
"serveRepair": "137.239.194.65:9011",
|
||||
"tpu": "137.239.194.65:9002"
|
||||
}
|
||||
```
|
||||
- Outbound gossip from biscayne exits via `doublezero0` with source
|
||||
137.239.194.65 — SNAT and routing work correctly in the outbound direction.
|
||||
|
||||
## Root Cause
|
||||
|
||||
**137.239.194.0/24 is not routable from the public internet.** The prefix
|
||||
belongs to DoubleZero's overlay address space (137.239.192.0/19, Momentum
|
||||
Telecom, WHOIS OriginAS: empty). It is advertised only via IS-IS within the
|
||||
DoubleZero switch mesh. There is no eBGP session on was-sw01 to advertise it
|
||||
to the ISP — all BGP peers are iBGP AS 65342 (DoubleZero internal).
|
||||
|
||||
When the validator advertises `tvu: 137.239.194.65:9000` in ContactInfo,
|
||||
public internet peers attempt to send turbine shreds to that IP, but the
|
||||
packets have no route through the global BGP table to reach was-sw01. Only
|
||||
DoubleZero-connected peers could potentially reach it via the overlay.
|
||||
|
||||
The old shred relay pipeline worked because it used `--public-tvu-address
|
||||
64.92.84.81:20000` — was-sw01's Et1/1 ISP uplink IP, which IS publicly
|
||||
routable. The `--gossip-host 137.239.194.65` approach advertises a
|
||||
DoubleZero-only IP for ALL sockets, making TVU/repair/TPU unreachable from
|
||||
non-DoubleZero peers.
|
||||
|
||||
The original hypothesis (ACL/PBR port filtering) was wrong. The tunnel and
|
||||
switch routing work correctly — the problem is upstream: traffic never arrives
|
||||
at was-sw01 in the first place.
|
||||
|
||||
## Impact
|
||||
|
||||
The validator cannot receive turbine shreds or serve repair requests via the
|
||||
low-latency Ashburn path. It falls back to the Miami public IP (186.233.184.235)
|
||||
for all shred/repair traffic, negating the benefit of `--gossip-host`.
|
||||
|
||||
## Fix Options
|
||||
|
||||
1. **Use 64.92.84.81 (was-sw01 Et1/1) for ContactInfo sockets.** This is the
|
||||
publicly routable Ashburn IP. Requires `--gossip-host 64.92.84.81` (or
|
||||
equivalent `--bind-address` config) and DNAT/forwarding on was-sw01 to relay
|
||||
traffic through the backbone → mia-sw01 → Tunnel500 → biscayne. The old
|
||||
`--public-tvu-address` pipeline used this IP successfully.
|
||||
|
||||
2. **Get DoubleZero to announce 137.239.194.0/24 via eBGP to the ISP.** This
|
||||
would make the current `--gossip-host 137.239.194.65` setup work, but
|
||||
requires coordination with DoubleZero operations.
|
||||
|
||||
3. **Hybrid approach**: Use 64.92.84.81 for public-facing sockets (TVU, repair,
|
||||
TPU) and 137.239.194.65 for gossip (which works via DoubleZero overlay).
|
||||
Requires agave to support per-protocol address binding, which it does not
|
||||
(`--gossip-host` sets ALL sockets to the same IP).
|
||||
|
||||
## Previous Workaround
|
||||
|
||||
The old `--public-tvu-address` pipeline used socat + shred-unwrap.py to relay
|
||||
shreds from 64.92.84.81:20000 to the validator. That pipeline is not persistent
|
||||
across reboots and was superseded by the `--gossip-host` approach (which turned
|
||||
out to be broken for non-DoubleZero peers).
|
||||
@@ -0,0 +1,51 @@
|
||||
# Bug: laconic-so etcd cleanup wipes core kubernetes service
|
||||
|
||||
## Summary
|
||||
|
||||
`_clean_etcd_keeping_certs()` in laconic-stack-orchestrator 1.1.0 deletes the `kubernetes` service from etcd, breaking cluster networking on restart.
|
||||
|
||||
## Component
|
||||
|
||||
`stack_orchestrator/deploy/k8s/helpers.py` — `_clean_etcd_keeping_certs()`
|
||||
|
||||
## Reproduction
|
||||
|
||||
1. Deploy with `laconic-so` to a k8s-kind target with persisted etcd (hostPath mount in kind-config.yml)
|
||||
2. `laconic-so deployment --dir <dir> stop` (destroys cluster)
|
||||
3. `laconic-so deployment --dir <dir> start` (recreates cluster with cleaned etcd)
|
||||
|
||||
## Symptoms
|
||||
|
||||
- `kindnet` pods enter CrashLoopBackOff with: `panic: unable to load in-cluster configuration, KUBERNETES_SERVICE_HOST and KUBERNETES_SERVICE_PORT must be defined`
|
||||
- `kubectl get svc kubernetes -n default` returns `NotFound`
|
||||
- coredns, caddy, local-path-provisioner stuck in Pending (no CNI without kindnet)
|
||||
- No pods can be scheduled
|
||||
|
||||
## Root Cause
|
||||
|
||||
`_clean_etcd_keeping_certs()` uses a whitelist that only preserves `/registry/secrets/caddy-system` keys. All other etcd keys are deleted, including `/registry/services/specs/default/kubernetes` — the core `kubernetes` ClusterIP service that kube-apiserver auto-creates.
|
||||
|
||||
When the kind cluster starts with the cleaned etcd, kube-apiserver sees the existing etcd data and does not re-create the `kubernetes` service. kindnet depends on the `KUBERNETES_SERVICE_HOST` environment variable which is injected by the kubelet from this service — without it, kindnet panics.
|
||||
|
||||
## Fix Options
|
||||
|
||||
1. **Expand the whitelist** to include `/registry/services/specs/default/kubernetes` and other core cluster resources
|
||||
2. **Fully wipe etcd** instead of selective cleanup — let the cluster bootstrap fresh (simpler, but loses Caddy TLS certs)
|
||||
3. **Don't persist etcd at all** — ephemeral etcd means clean state every restart (recommended for kind deployments)
|
||||
|
||||
## Workaround
|
||||
|
||||
Fully delete the kind cluster before `start`:
|
||||
|
||||
```bash
|
||||
kind delete cluster --name <cluster-name>
|
||||
laconic-so deployment --dir <dir> start
|
||||
```
|
||||
|
||||
This forces fresh etcd bootstrap. Downside: all other services deployed to the cluster (DaemonSets, other namespaces) are destroyed.
|
||||
|
||||
## Impact
|
||||
|
||||
- Affects any k8s-kind deployment with persisted etcd
|
||||
- Cluster is unrecoverable without full destroy+recreate
|
||||
- All non-laconic-so-managed workloads in the cluster are lost
|
||||
@@ -0,0 +1,75 @@
|
||||
# Bug: laconic-so crashes on re-deploy when caddy ingress already exists
|
||||
|
||||
## Summary
|
||||
|
||||
`laconic-so deployment start` crashes with `FailToCreateError` when the kind cluster already has caddy ingress resources installed. The deployer uses `create_from_yaml()` which fails on `AlreadyExists` conflicts instead of applying idempotently. This prevents the application deployment from ever being reached — the crash happens before any app manifests are applied.
|
||||
|
||||
## Component
|
||||
|
||||
`stack_orchestrator/deploy/k8s/deploy_k8s.py:366` — `up()` method
|
||||
`stack_orchestrator/deploy/k8s/helpers.py:369` — `install_ingress_for_kind()`
|
||||
|
||||
## Reproduction
|
||||
|
||||
1. `kind delete cluster --name laconic-70ce4c4b47e23b85`
|
||||
2. `laconic-so deployment --dir /srv/deployments/agave start` — creates cluster, loads images, installs caddy ingress, but times out or is interrupted before app deployment completes
|
||||
3. `laconic-so deployment --dir /srv/deployments/agave start` — crashes immediately after image loading
|
||||
|
||||
## Symptoms
|
||||
|
||||
- Traceback ending in:
|
||||
```
|
||||
kubernetes.utils.create_from_yaml.FailToCreateError:
|
||||
Error from server (Conflict): namespaces "caddy-system" already exists
|
||||
Error from server (Conflict): serviceaccounts "caddy-ingress-controller" already exists
|
||||
Error from server (Conflict): clusterroles.rbac.authorization.k8s.io "caddy-ingress-controller" already exists
|
||||
...
|
||||
```
|
||||
- Namespace `laconic-laconic-70ce4c4b47e23b85` exists but is empty — no pods, no deployments, no events
|
||||
- Cluster is healthy, images are loaded, but no app manifests are applied
|
||||
|
||||
## Root Cause
|
||||
|
||||
`install_ingress_for_kind()` calls `kubernetes.utils.create_from_yaml()` which uses `POST` (create) semantics. If the resources already exist (from a previous partial run), every resource returns `409 Conflict` and `create_from_yaml` raises `FailToCreateError`, aborting the entire `up()` method before the app deployment step.
|
||||
|
||||
The first `laconic-so start` after a fresh `kind delete` works because:
|
||||
1. Image loading into the kind node takes 5-10 minutes (images are ~10GB+)
|
||||
2. Caddy ingress is installed successfully
|
||||
3. App deployment begins
|
||||
|
||||
But if that first run is interrupted (timeout, Ctrl-C, ansible timeout), the second run finds caddy already installed and crashes.
|
||||
|
||||
## Fix Options
|
||||
|
||||
1. **Use server-side apply** instead of `create_from_yaml()` — `kubectl apply` is idempotent
|
||||
2. **Check if ingress exists before installing** — skip `install_ingress_for_kind()` if caddy-system namespace exists
|
||||
3. **Catch `AlreadyExists` and continue** — treat 409 as success for infrastructure resources
|
||||
|
||||
## Workaround
|
||||
|
||||
Delete the caddy ingress resources before re-running:
|
||||
|
||||
```bash
|
||||
kubectl delete namespace caddy-system
|
||||
kubectl delete clusterrole caddy-ingress-controller
|
||||
kubectl delete clusterrolebinding caddy-ingress-controller
|
||||
kubectl delete ingressclass caddy
|
||||
laconic-so deployment --dir /srv/deployments/agave start
|
||||
```
|
||||
|
||||
Or nuke the entire cluster and start fresh:
|
||||
|
||||
```bash
|
||||
kind delete cluster --name laconic-70ce4c4b47e23b85
|
||||
laconic-so deployment --dir /srv/deployments/agave start
|
||||
```
|
||||
|
||||
## Interaction with ansible timeout
|
||||
|
||||
The `biscayne-redeploy.yml` playbook sets a 600s timeout on the `laconic-so deployment start` task. Image loading alone can exceed this on a fresh cluster (images must be re-loaded into the new kind node). When ansible kills the process at 600s, the caddy ingress is already installed but the app is not — putting the cluster into the broken state described above. Subsequent playbook runs hit this bug on every attempt.
|
||||
|
||||
## Impact
|
||||
|
||||
- Blocks all re-deploys on biscayne without manual cleanup
|
||||
- The playbook cannot recover automatically — every retry hits the same conflict
|
||||
- Discovered 2026-03-05 during full wipe redeploy of biscayne validator
|
||||
@@ -0,0 +1,121 @@
|
||||
# DoubleZero Multicast Access Requests
|
||||
|
||||
## Status (2026-03-06)
|
||||
|
||||
DZ multicast is **still in testnet** (client v0.2.2). Multicast groups are defined
|
||||
on the DZ ledger with on-chain access control (publishers/subscribers). The testnet
|
||||
allocates addresses from 233.84.178.0/24 (AS21682). Not yet available for production
|
||||
Solana shred delivery.
|
||||
|
||||
## Biscayne Connection Details
|
||||
|
||||
Provide these details when requesting subscriber access:
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Client IP | 186.233.184.235 |
|
||||
| Validator identity | 4WeLUxfQghbhsLEuwaAzjZiHg2VBw87vqHc4iZrGvKPr |
|
||||
| DZ identity | 3Bw6v7EruQvTwoY79h2QjQCs2KBQFzSneBdYUbcXK1Tr |
|
||||
| DZ device | laconic-mia-sw01 |
|
||||
| Contributor / tenant | laconic |
|
||||
|
||||
## Jito ShredStream
|
||||
|
||||
**Not a DZ multicast group.** ShredStream is Jito's own shred delivery service,
|
||||
independent of DoubleZero multicast. It provides low-latency shreds from leaders
|
||||
on the Solana network via a proxy client that connects to the Jito Block Engine.
|
||||
|
||||
| Property | Value |
|
||||
|----------|-------|
|
||||
| What it does | Delivers shreds from Jito-connected leaders with low latency. Provides a redundant shred path for servers in remote locations. |
|
||||
| How it works | `shredstream-proxy` authenticates to a Jito Block Engine via keypair, receives shreds, forwards them to configured UDP destinations (e.g. validator TVU port). |
|
||||
| Cost | **Unknown.** Docs don't list pricing. Was previously "complimentary" for searchers (2024). May require approval. |
|
||||
| Requirements | Approved Solana pubkey (form submission), auth keypair, firewall open on UDP 20000, TVU port of your node. |
|
||||
| Regions | Amsterdam, Dublin, Frankfurt, London, New York, Salt Lake City, Singapore, Tokyo. Max 2 regions selectable. |
|
||||
| Limitations | No NAT support. Bridge networking incompatible with multicast mode. |
|
||||
| Repo | https://github.com/jito-labs/shredstream-proxy |
|
||||
| Docs | https://docs.jito.wtf/lowlatencytxnfeed/ |
|
||||
| Status for biscayne | **Not yet requested.** Need to submit pubkey for approval. |
|
||||
|
||||
ShredStream is relevant to our shred completeness problem — it provides an additional
|
||||
shred source beyond turbine and the Ashburn relay. It would run as a sidecar process
|
||||
forwarding shreds to the validator's TVU port.
|
||||
|
||||
## DZ Multicast Groups
|
||||
|
||||
DZ multicast uses PIM (Protocol Independent Multicast) and MSDP (Multicast Source
|
||||
Discovery Protocol). Group owners define allowed publishers and subscribers on the
|
||||
DZ ledger. Switch ASICs handle packet replication — no CPU overhead.
|
||||
|
||||
### bebop
|
||||
|
||||
Listed in earlier notes as a multicast shred distribution group. **No public
|
||||
documentation found.** Cannot confirm this exists as a DZ multicast group.
|
||||
|
||||
- **Owner:** Unknown
|
||||
- **Status:** Unverified — may not exist as described
|
||||
|
||||
### turbine (future)
|
||||
|
||||
Solana's native shred propagation via DZ multicast. Jito has expressed interest
|
||||
in leveraging multicast for shred delivery. Not yet available for production use.
|
||||
|
||||
- **Owner:** Solana Foundation / Anza (native turbine), Jito (shredstream)
|
||||
- **Status:** Testnet only (DZ client v0.2.2)
|
||||
|
||||
## bloXroute OFR (Optimized Feed Relay)
|
||||
|
||||
Commercial shred delivery service. Runs a gateway docker container on your node that
|
||||
connects to bloXroute's BDN (Blockchain Distribution Network) to receive shreds
|
||||
faster than default turbine (~30-50ms improvement, beats turbine ~98% of the time).
|
||||
|
||||
| Property | Value |
|
||||
|----------|-------|
|
||||
| What it does | Delivers shreds via bloXroute's BDN with optimized relay topologies. Not just a different turbine path — uses their own distribution network. |
|
||||
| How it works | Docker gateway container on your node, communicates with bloXroute OFR relay over UDP 18888. Forwards shreds to your validator. |
|
||||
| Cost | **$300/mo** (Professional, 1500 tx/day), **$1,250/mo** (Enterprise, unlimited tx). OFR gateway without local node requires Enterprise Elite ($5,000+/mo). |
|
||||
| Requirements | Docker, UDP port 18888 open, bloXroute subscription. |
|
||||
| Open source | Gateway at https://github.com/bloXroute-Labs/solana-gateway |
|
||||
| Docs | https://docs.bloxroute.com/solana/optimized-feed-relay |
|
||||
| Status for biscayne | **Not yet evaluated.** Monthly cost may not be justified. |
|
||||
|
||||
bloXroute's value proposition: they operate nodes at multiple turbine tree positions
|
||||
across their network, aggregate shreds, and redistribute via their BDN. This is the
|
||||
"multiple identities collecting different shreds" approach — but operated by bloXroute,
|
||||
not by us.
|
||||
|
||||
## How These Services Get More Shreds
|
||||
|
||||
Turbine tree position is determined by validator identity (pubkey). A single validator
|
||||
gets shreds from one position in the tree per slot. Services like Jito ShredStream
|
||||
and bloXroute OFR operate many nodes with different identities across the turbine
|
||||
tree, aggregate the shreds they each receive, and redistribute the combined set to
|
||||
subscribers. This is why they can deliver shreds the subscriber's own turbine position
|
||||
would never see.
|
||||
|
||||
**An open-source equivalent would require running multiple lightweight validator
|
||||
identities (non-voting, minimal stake) at different locations, each collecting shreds
|
||||
from their unique turbine tree position, and forwarding them to the main validator.**
|
||||
No known open-source project implements this pattern.
|
||||
|
||||
## Sources
|
||||
|
||||
- [Jito ShredStream docs](https://docs.jito.wtf/lowlatencytxnfeed/)
|
||||
- [shredstream-proxy repo](https://github.com/jito-labs/shredstream-proxy)
|
||||
- [bloXroute OFR docs](https://docs.bloxroute.com/solana/optimized-feed-relay)
|
||||
- [bloXroute pricing](https://bloxroute.com/pricing/)
|
||||
- [bloXroute OFR intro](https://bloxroute.com/pulse/introducing-ofrs-faster-shreds-better-performance-on-solana/)
|
||||
- [DZ multicast announcement](https://doublezero.xyz/journal/doublezero-introduces-multicast-support-smarter-faster-data-delivery-for-distributed-systems)
|
||||
|
||||
## Request Template
|
||||
|
||||
When contacting a group owner, use something like:
|
||||
|
||||
> We'd like to subscribe to your DoubleZero multicast group for our Solana
|
||||
> validator. Our details:
|
||||
>
|
||||
> - Validator: 4WeLUxfQghbhsLEuwaAzjZiHg2VBw87vqHc4iZrGvKPr
|
||||
> - DZ identity: 3Bw6v7EruQvTwoY79h2QjQCs2KBQFzSneBdYUbcXK1Tr
|
||||
> - Client IP: 186.233.184.235
|
||||
> - Device: laconic-mia-sw01
|
||||
> - Tenant: laconic
|
||||
@@ -0,0 +1,121 @@
|
||||
# DoubleZero Current State and Bug Fixes
|
||||
|
||||
## Biscayne Connection Details
|
||||
|
||||
| Field | Value |
|
||||
|-------|-------|
|
||||
| Host | biscayne.vaasl.io (186.233.184.235) |
|
||||
| DZ identity | `3Bw6v7EruQvTwoY79h2QjQCs2KBQFzSneBdYUbcXK1Tr` |
|
||||
| Validator identity | `4WeLUxfQghbhsLEuwaAzjZiHg2VBw87vqHc4iZrGvKPr` |
|
||||
| Nearest device | laconic-mia-sw01 (0.3ms) |
|
||||
| DZ version (host) | 0.8.10 |
|
||||
| DZ version (container) | 0.8.11 |
|
||||
| k8s version | 1.35.1 (kind) |
|
||||
|
||||
## Current State (2026-03-03)
|
||||
|
||||
The host systemd `doublezerod` is connected and working. The container sidecar
|
||||
doublezerod is broken. Both are running simultaneously.
|
||||
|
||||
| Instance | Identity | Status |
|
||||
|----------|----------|--------|
|
||||
| Host systemd | `3Bw6v7...` (correct) | BGP Session Up, IBRL to laconic-mia-sw01 |
|
||||
| Container sidecar | `Cw9qun...` (wrong) | Disconnected, error loop |
|
||||
| DaemonSet manifest | N/A | Never applied, dead code |
|
||||
|
||||
### Access pass
|
||||
|
||||
The access pass for 186.233.184.235 is registered and connected:
|
||||
|
||||
```
|
||||
type: prepaid
|
||||
payer: 3Bw6v7EruQvTwoY79h2QjQCs2KBQFzSneBdYUbcXK1Tr
|
||||
status: connected
|
||||
owner: DZfLKFDgLShjY34WqXdVVzHUvVtrYXb7UtdrALnGa8jw
|
||||
```
|
||||
|
||||
## Bugs
|
||||
|
||||
### BUG-1: Container doublezerod has wrong identity
|
||||
|
||||
The entrypoint script (`entrypoint.sh`) auto-generates a new `id.json` if one isn't
|
||||
found. The volume at `/srv/deployments/agave/data/doublezero-config/` was empty at
|
||||
first boot, so it generated `Cw9qun...` instead of using the registered identity.
|
||||
|
||||
**Root cause:** The real `id.json` lives at `/home/solana/.config/doublezero/id.json`
|
||||
(created by the host-level DZ install). The container volume is a separate path that
|
||||
was never seeded.
|
||||
|
||||
**Fix:**
|
||||
```bash
|
||||
sudo cp /home/solana/.config/doublezero/id.json \
|
||||
/srv/deployments/agave/data/doublezero-config/id.json
|
||||
```
|
||||
|
||||
### BUG-2: Container doublezerod can't resolve DZ passport program
|
||||
|
||||
`DOUBLEZERO_RPC_ENDPOINT` in `spec.yml` is `http://127.0.0.1:8899` — the local
|
||||
validator. But the local validator hasn't replayed enough slots to have the DZ
|
||||
passport program accounts (`ser2VaTMAcYTaauMrTSfSrxBaUDq7BLNs2xfUugTAGv`).
|
||||
doublezerod calls `GetProgramAccounts` every 30 seconds and gets empty results.
|
||||
|
||||
**Fix in `deployment/spec.yml`:**
|
||||
```yaml
|
||||
# Use public RPC for DZ bootstrapping until local validator is caught up
|
||||
DOUBLEZERO_RPC_ENDPOINT: https://api.mainnet-beta.solana.com
|
||||
```
|
||||
|
||||
Switch back to `http://127.0.0.1:8899` once the local validator is synced.
|
||||
|
||||
### BUG-3: Container doublezerod lacks hostNetwork
|
||||
|
||||
laconic-so was not translating `network_mode: host` from compose files to
|
||||
`hostNetwork: true` in generated k8s pod specs. Without host network access, the
|
||||
container can't create GRE tunnels (IP proto 47) or run BGP (tcp/179 on
|
||||
169.254.0.0/16).
|
||||
|
||||
**Fix:** Deploy with stack-orchestrator branch `fix/k8s-port-mappings-hostnetwork-v2`
|
||||
(commit `fb69cc58`, 2026-03-03) which adds automatic hostNetwork detection.
|
||||
|
||||
### BUG-4: DaemonSet workaround is dead code
|
||||
|
||||
`deployment/k8s-manifests/doublezero-daemonset.yaml` was a workaround for BUG-3.
|
||||
Now that laconic-so supports hostNetwork natively, it should be deleted.
|
||||
|
||||
**Fix:** Remove `deployment/k8s-manifests/doublezero-daemonset.yaml` from agave-stack.
|
||||
|
||||
### BUG-5: Two doublezerod instances running simultaneously
|
||||
|
||||
The host systemd `doublezerod` and the container sidecar are both running. Once the
|
||||
container is fixed (BUG-1 through BUG-3), the host service must be disabled to avoid
|
||||
two processes fighting over the GRE tunnel.
|
||||
|
||||
**Fix:**
|
||||
```bash
|
||||
sudo systemctl stop doublezerod
|
||||
sudo systemctl disable doublezerod
|
||||
```
|
||||
|
||||
## Diagnostic Commands
|
||||
|
||||
Always use `sudo -u solana` for host-level DZ commands — the identity is under
|
||||
`/home/solana/.config/doublezero/`.
|
||||
|
||||
```bash
|
||||
# Host
|
||||
sudo -u solana doublezero address # expect 3Bw6v7...
|
||||
sudo -u solana doublezero status # tunnel state
|
||||
sudo -u solana doublezero latency # device reachability
|
||||
sudo -u solana doublezero access-pass list | grep 186.233.184 # access pass
|
||||
sudo -u solana doublezero balance # credits
|
||||
ip route | grep doublezero0 # BGP routes
|
||||
|
||||
# Container (from kind node)
|
||||
kubectl -n <ns> exec <pod> -c doublezerod -- doublezero address
|
||||
kubectl -n <ns> exec <pod> -c doublezerod -- doublezero status
|
||||
kubectl -n <ns> exec <pod> -c doublezerod -- doublezero --version
|
||||
|
||||
# Logs
|
||||
kubectl -n <ns> logs <pod> -c doublezerod --tail=30
|
||||
sudo journalctl -u doublezerod -f # host systemd logs
|
||||
```
|
||||
@@ -0,0 +1,65 @@
|
||||
# Feature: Use local registry for kind image loading
|
||||
|
||||
## Summary
|
||||
|
||||
`laconic-so deployment start` uses `kind load docker-image` to copy container images from the host Docker daemon into the kind node's containerd. This serializes the full image (`docker save`), pipes it through `docker exec`, and deserializes it (`ctr image import`). For biscayne's ~837MB agave image plus the doublezero image, this takes 5-10 minutes on every cluster recreate — copying between two container runtimes on the same machine.
|
||||
|
||||
## Current behavior
|
||||
|
||||
```
|
||||
docker build → host Docker daemon (image stored once)
|
||||
kind load docker-image → docker save | docker exec kind-node ctr import (full copy)
|
||||
```
|
||||
|
||||
This happens in `stack_orchestrator/deploy/k8s/deploy_k8s.py` every time `laconic-so deployment start` runs and the image isn't already present in the kind node.
|
||||
|
||||
## Proposed behavior
|
||||
|
||||
Run a persistent local registry (`registry:2`) on the host. `laconic-so` pushes images there after build. Kind's containerd is configured to pull from it.
|
||||
|
||||
```
|
||||
docker build → docker tag localhost:5001/image → docker push localhost:5001/image
|
||||
kind node containerd → pulls from localhost:5001 (fast, no serialization)
|
||||
```
|
||||
|
||||
The registry container persists across kind cluster deletions. Images are always available without reloading.
|
||||
|
||||
## Implementation
|
||||
|
||||
1. **Registry container**: `docker run -d --restart=always -p 5001:5000 --name kind-registry registry:2`
|
||||
|
||||
2. **Kind config** — add registry mirror to `containerdConfigPatches` in kind-config.yml:
|
||||
```yaml
|
||||
containerdConfigPatches:
|
||||
- |-
|
||||
[plugins."io.containerd.grpc.v1.cri".registry.mirrors."localhost:5001"]
|
||||
endpoint = ["http://kind-registry:5000"]
|
||||
```
|
||||
|
||||
3. **Connect registry to kind network**: `docker network connect kind kind-registry`
|
||||
|
||||
4. **laconic-so change** — in `deploy_k8s.py`, replace `kind load docker-image` with:
|
||||
```python
|
||||
# Tag and push to local registry instead of kind load
|
||||
docker tag image:local localhost:5001/image:local
|
||||
docker push localhost:5001/image:local
|
||||
```
|
||||
|
||||
5. **Compose files** — image references change from `laconicnetwork/agave:local` to `localhost:5001/laconicnetwork/agave:local`
|
||||
|
||||
Kind documents this pattern: https://kind.sigs.k8s.io/docs/user/local-registry/
|
||||
|
||||
## Impact
|
||||
|
||||
- Eliminates 5-10 minute image loading step on every cluster recreate
|
||||
- Registry persists across `kind delete cluster` — no re-push needed unless the image itself changes
|
||||
- `docker push` to a local registry is near-instant (shared filesystem, layer dedup)
|
||||
- Unblocks faster iteration on redeploy cycles
|
||||
|
||||
## Scope
|
||||
|
||||
This is a `stack-orchestrator` change, specifically in `deploy_k8s.py`. The kind-config.yml also needs the registry mirror config, which `laconic-so` generates from `spec.yml`.
|
||||
|
||||
## Discovered
|
||||
|
||||
2026-03-05 — during biscayne full wipe redeploy, `laconic-so start` spent most of its runtime on `kind load docker-image`, causing ansible timeouts and cascading failures (caddy ingress conflict bug).
|
||||
@@ -0,0 +1,78 @@
|
||||
# Known Issues
|
||||
|
||||
## BUG-6: Validator logging not configured, only stdout available
|
||||
|
||||
**Observed:** 2026-03-03
|
||||
|
||||
The validator only logs to stdout. kubectl logs retains ~2 minutes of history
|
||||
at current log volume before the buffer fills. When diagnosing a replay stall,
|
||||
the startup logs (snapshot load, initial replay, error conditions) were gone.
|
||||
|
||||
**Impact:** Cannot determine why the validator replay stage stalled — the
|
||||
startup logs that would show the root cause are not available.
|
||||
|
||||
**Fix:** Configure the `--log` flag in the validator start script to write to
|
||||
a persistent volume, so logs survive container restarts and aren't limited
|
||||
to the kubectl buffer.
|
||||
|
||||
## BUG-7: Metrics endpoint unreachable from validator pod
|
||||
|
||||
**Observed:** 2026-03-03
|
||||
|
||||
```
|
||||
WARN solana_metrics::metrics submit error: error sending request for url
|
||||
(http://localhost:8086/write?db=agave_metrics&u=admin&p=admin&precision=n)
|
||||
```
|
||||
|
||||
The validator is configured with `SOLANA_METRICS_CONFIG` pointing to
|
||||
`http://172.20.0.1:8086` (the kind docker bridge gateway), but the logs show
|
||||
it trying `localhost:8086`. The InfluxDB container (`solana-monitoring-influxdb-1`)
|
||||
is running on the host, but the validator can't reach it.
|
||||
|
||||
**Impact:** No metrics collection. Cannot use Grafana dashboards to diagnose
|
||||
performance issues or track sync progress over time.
|
||||
|
||||
## BUG-8: sysctl values not visible inside kind container
|
||||
|
||||
**Observed:** 2026-03-03
|
||||
|
||||
```
|
||||
ERROR solana_core::system_monitor_service Failed to query value for net.core.rmem_max: no such sysctl
|
||||
WARN solana_core::system_monitor_service net.core.rmem_max: recommended=134217728, current=-1 too small
|
||||
```
|
||||
|
||||
The host has correct sysctl values (`net.core.rmem_max = 134217728`), but
|
||||
`/proc/sys/net/core/` does not exist inside the kind node container. The
|
||||
validator reads `-1` and reports the buffer as too small.
|
||||
|
||||
The network buffers themselves may still be effective (they're set on the
|
||||
host network namespace which the pod shares via `hostNetwork: true`), but
|
||||
this is unverified. If the buffers are not effective, it could limit shred
|
||||
ingestion throughput and contribute to slow repair.
|
||||
|
||||
**Fix options:**
|
||||
- Set sysctls on the kind node container at creation time
|
||||
(`kind` supports `kubeadmConfigPatches` and sysctl configuration)
|
||||
- Verify empirically whether the host sysctls apply to hostNetwork pods
|
||||
by checking actual socket buffer sizes from inside the pod
|
||||
|
||||
## Validator replay stall (under investigation)
|
||||
|
||||
**Observed:** 2026-03-03
|
||||
|
||||
The validator root has been stuck at slot 403,892,310 for 55+ minutes.
|
||||
The gap to the cluster tip is ~120,000 slots and growing.
|
||||
|
||||
**Observed symptoms:**
|
||||
- Zero `Frozen` banks in log history — replay stage is not processing slots
|
||||
- All incoming slots show `bank_status: Unprocessed`
|
||||
- Repair only requests tip slots and two specific old slots (403,892,310,
|
||||
403,909,228) — not the ~120k slot gap
|
||||
- Repair peer count is 3-12 per cycle (vs 1,000+ gossip peers)
|
||||
- Startup logs have rotated out (BUG-6), so initialization context is lost
|
||||
|
||||
**Unknown:**
|
||||
- What snapshot the validator loaded at boot
|
||||
- Whether replay ever started or was blocked from the beginning
|
||||
- Whether the sysctl issue (BUG-8) is limiting repair throughput
|
||||
- Whether the missing metrics (BUG-7) would show what's happening internally
|
||||
@@ -0,0 +1,191 @@
|
||||
# Shred Collector Relay
|
||||
|
||||
## Problem
|
||||
|
||||
Turbine assigns each validator a single position in the shred distribution tree
|
||||
per slot, determined by its pubkey. A validator in Miami with one identity receives
|
||||
shreds from one set of tree neighbors — typically ~60-70% of shreds for any given
|
||||
slot. The remaining 30-40% must come from the repair protocol, which is too slow
|
||||
to keep pace with chain production (see analysis below).
|
||||
|
||||
Commercial services (Jito ShredStream, bloXroute OFR) solve this by running many
|
||||
nodes with different identities across the turbine tree, aggregating shreds, and
|
||||
redistributing the combined set to subscribers. This works but costs $300-5,000/mo
|
||||
and adds a dependency on a third party.
|
||||
|
||||
## Concept
|
||||
|
||||
Run lightweight **shred collector** nodes at multiple geographic locations on
|
||||
the Laconic network (Ashburn, Dallas, etc.). Each collector has its own keypair,
|
||||
joins gossip with a unique identity, receives turbine shreds from its unique tree
|
||||
position, and forwards raw shred packets to the main validator in Miami. The main
|
||||
validator inserts these shreds into its blockstore alongside its own turbine shreds,
|
||||
increasing completeness toward 100% without relying on repair.
|
||||
|
||||
```
|
||||
Turbine Tree
|
||||
/ | \
|
||||
/ | \
|
||||
collector-ash collector-dfw biscayne (main validator)
|
||||
(Ashburn) (Dallas) (Miami)
|
||||
identity A identity B identity C
|
||||
~60% shreds ~60% shreds ~60% shreds
|
||||
\ | /
|
||||
\ | /
|
||||
→ UDP forward via DZ backbone →
|
||||
|
|
||||
biscayne blockstore
|
||||
~95%+ shreds (union of A∪B∪C)
|
||||
```
|
||||
|
||||
Each collector sees a different ~60% slice of the turbine tree. The union of
|
||||
three independent positions yields ~94% coverage (1 - 0.4³ = 0.936). Four
|
||||
collectors yield ~97%. The main validator fills the remaining few percent via
|
||||
repair, which is fast when only 3-6% of shreds are missing.
|
||||
|
||||
## Why This Works
|
||||
|
||||
The math from biscayne's recovery (2026-03-06):
|
||||
|
||||
| Metric | Value |
|
||||
|--------|-------|
|
||||
| Compute-bound replay (complete blocks) | 5.2 slots/sec |
|
||||
| Repair-bound replay (incomplete blocks) | 0.5 slots/sec |
|
||||
| Chain production rate | 2.5 slots/sec |
|
||||
| Turbine + relay delivery per identity | ~60-70% |
|
||||
| Repair bandwidth | ~600 shreds/sec (estimated) |
|
||||
| Repair needed to converge at 60% delivery | 5x current bandwidth |
|
||||
| Repair needed to converge at 95% delivery | Easily sufficient |
|
||||
|
||||
At 60% shred delivery, repair must fill 40% per slot — too slow to converge.
|
||||
At 95% delivery (3 collectors), repair fills 5% per slot — well within capacity.
|
||||
The validator replays at near compute-bound speed (5+ slots/sec) and converges.
|
||||
|
||||
## Infrastructure
|
||||
|
||||
Laconic already has DZ-connected switches at multiple sites:
|
||||
|
||||
| Site | Device | Latency to Miami | Backbone |
|
||||
|------|--------|-------------------|----------|
|
||||
| Miami | laconic-mia-sw01 | 0.24ms | local |
|
||||
| Ashburn | laconic-was-sw01 | ~29ms | Et4/1 25.4ms |
|
||||
| Dallas | laconic-dfw-sw01 | ~30ms | TBD |
|
||||
|
||||
The DZ backbone carries traffic between sites at line rate. Shred packets are
|
||||
~1280 bytes each. At ~3,000 shreds/slot and 2.5 slots/sec, each collector
|
||||
forwards ~7,500 packets/sec (~10 MB/s) — trivial bandwidth for the backbone.
|
||||
|
||||
## Collector Architecture
|
||||
|
||||
The collector does NOT need to be a full validator. It needs to:
|
||||
|
||||
1. **Join gossip** — advertise a ContactInfo with its own pubkey and a TVU
|
||||
address (the site's IP)
|
||||
2. **Receive turbine shreds** — UDP packets on the advertised TVU port
|
||||
3. **Forward shreds** — retransmit raw UDP packets to biscayne's TVU port
|
||||
|
||||
It does NOT need to: replay transactions, maintain accounts state, store a
|
||||
ledger, load a snapshot, vote, or run RPC.
|
||||
|
||||
### Option A: Firedancer Minimal Build
|
||||
|
||||
Firedancer (Apache 2, C) has a tile-based architecture where each function
|
||||
(net, gossip, shred, bank, store, etc.) runs as an independent Linux process.
|
||||
A minimal build using only the networking + gossip + shred tiles would:
|
||||
|
||||
- Join gossip and advertise a TVU address
|
||||
- Receive turbine shreds via the shred tile
|
||||
- Forward shreds to a configured destination instead of to bank/store
|
||||
|
||||
This requires modifying the shred tile to add a UDP forwarder output instead
|
||||
of (or in addition to) the normal bank handoff. The rest of the tile pipeline
|
||||
(bank, pack, poh, store) is simply not started.
|
||||
|
||||
**Estimated effort:** Moderate. Firedancer's tile architecture is designed for
|
||||
this kind of composition. The main work is adding a forwarder sink to the shred
|
||||
tile and testing gossip participation without the full validator stack.
|
||||
|
||||
**Source:** https://github.com/firedancer-io/firedancer
|
||||
|
||||
### Option B: Agave Non-Voting Minimal
|
||||
|
||||
Run `agave-validator --no-voting` with `--limit-ledger-size 0` and minimal
|
||||
config. Agave still requires a snapshot to start and runs the full process, but
|
||||
with no voting and minimal ledger it would be lighter than a full node.
|
||||
|
||||
**Downside:** Agave is monolithic — you can't easily disable replay/accounts.
|
||||
It still loads a snapshot, builds the accounts index, and runs replay. This
|
||||
defeats the purpose of a lightweight collector.
|
||||
|
||||
### Option C: Custom Gossip + TVU Receiver
|
||||
|
||||
Write a minimal Rust binary using agave's `solana-gossip` and `solana-streamer`
|
||||
crates to:
|
||||
1. Bootstrap into gossip via entrypoints
|
||||
2. Advertise ContactInfo with TVU socket
|
||||
3. Receive shred packets on TVU
|
||||
4. Forward them via UDP
|
||||
|
||||
**Estimated effort:** Significant. Gossip protocol participation is complex
|
||||
(CRDS protocol, pull/push protocol, protocol versioning). Using the agave
|
||||
crates directly is possible but poorly documented for standalone use.
|
||||
|
||||
### Option D: Run Collectors on Biscayne
|
||||
|
||||
Run the collector processes on biscayne itself, each advertising a TVU address
|
||||
at a remote site. The switches at each site forward inbound TVU traffic to
|
||||
biscayne via the DZ backbone using traffic-policy redirects (same pattern as
|
||||
`ashburn-validator-relay.md`).
|
||||
|
||||
**Advantage:** No compute needed at remote sites. Just switch config + loopback
|
||||
IPs. All collector processes run in Miami.
|
||||
|
||||
**Risk:** Gossip advertises IP + port. If the collector runs on biscayne but
|
||||
advertises an Ashburn IP, gossip protocol interactions (pull requests, pings)
|
||||
arrive at the Ashburn IP and must be forwarded back to biscayne. This adds
|
||||
~58ms RTT to gossip protocol messages, which may cause timeouts or peer
|
||||
quality degradation. Needs testing.
|
||||
|
||||
## Recommendation
|
||||
|
||||
Option A (Firedancer minimal build) is the correct long-term approach. It
|
||||
produces a single binary that does exactly one thing: collect shreds from a
|
||||
unique turbine tree position and forward them. It runs on minimal hardware
|
||||
(a small VM or container at each site, or on biscayne with remote TVU
|
||||
addresses).
|
||||
|
||||
Option D (collectors on biscayne with switch forwarding) is the fastest to
|
||||
test since it needs no new software — just switch config and multiple
|
||||
agave-validator instances with `--no-voting`. The question is whether agave
|
||||
can start without a snapshot if we only care about gossip + TVU.
|
||||
|
||||
## Deployment Topology
|
||||
|
||||
```
|
||||
biscayne (186.233.184.235)
|
||||
├── agave-validator (main, identity C, TVU 186.233.184.235:9000)
|
||||
├── collector-ash (identity A, TVU 137.239.194.65:9000)
|
||||
│ └── shreds forwarded via was-sw01 traffic-policy
|
||||
├── collector-dfw (identity B, TVU <dfw-ip>:9000)
|
||||
│ └── shreds forwarded via dfw-sw01 traffic-policy
|
||||
└── blockstore receives union of A∪B∪C shreds
|
||||
|
||||
was-sw01 (Ashburn)
|
||||
└── Loopback: 137.239.194.65
|
||||
└── traffic-policy: UDP dst 137.239.194.65:9000 → nexthop mia-sw01
|
||||
|
||||
dfw-sw01 (Dallas)
|
||||
└── Loopback: <assigned IP>
|
||||
└── traffic-policy: UDP dst <assigned IP>:9000 → nexthop mia-sw01
|
||||
```
|
||||
|
||||
## Open Questions
|
||||
|
||||
1. Can agave-validator start in gossip-only mode without a snapshot?
|
||||
2. Does Firedancer's shred tile work standalone without bank/replay?
|
||||
3. What is the gossip protocol timeout for remote TVU addresses (Option D)?
|
||||
4. How does the turbine tree handle multiple identities from the same IP
|
||||
(if running all collectors on biscayne)?
|
||||
5. Do we need stake on collector identities to be placed in the turbine tree,
|
||||
or do unstaked nodes still participate?
|
||||
6. What IP block is available on dfw-sw01 for a collector loopback?
|
||||
@@ -0,0 +1,161 @@
|
||||
# TVU Shred Relay — Data-Plane Redirect
|
||||
|
||||
## Overview
|
||||
|
||||
Biscayne's agave validator advertises `64.92.84.81:20000` (laconic-was-sw01 Et1/1) as its TVU
|
||||
address. Turbine shreds arrive as normal UDP to the switch's front-panel IP. The 7280CR3A ASIC
|
||||
handles front-panel traffic without punting to Linux userspace — it sees a local interface IP
|
||||
with no service and drops at the hardware level.
|
||||
|
||||
### Previous approach (monitor + socat)
|
||||
|
||||
EOS monitor session mirrored matched packets to CPU (mirror0 interface). socat read from mirror0
|
||||
and relayed to biscayne. shred-unwrap.py on biscayne stripped encapsulation headers.
|
||||
|
||||
Fragile: socat ran as a foreground process, died on disconnect.
|
||||
|
||||
### New approach (traffic-policy redirect)
|
||||
|
||||
EOS `traffic-policy` with `set nexthop` and `system-rule overriding-action redirect` overrides
|
||||
the ASIC's "local IP, handle myself" decision. The ASIC forwards matched packets to the
|
||||
specified next-hop at line rate. Pure data plane, no CPU involvement, persists in startup-config.
|
||||
|
||||
Available since EOS 4.28.0F on R3 platforms. Confirmed on 4.34.0F.
|
||||
|
||||
## Architecture
|
||||
|
||||
```
|
||||
Turbine peers (hundreds of validators)
|
||||
|
|
||||
v UDP shreds to 64.92.84.81:20000
|
||||
laconic-was-sw01 Et1/1 (Ashburn)
|
||||
| ASIC matches traffic-policy SHRED-RELAY
|
||||
| Redirects to nexthop 172.16.1.189 (data plane, line rate)
|
||||
v Et4/1 backbone (25.4ms)
|
||||
laconic-mia-sw01 Et4/1 (Miami)
|
||||
| forwards via default route (same metro)
|
||||
v 0.13ms
|
||||
biscayne (186.233.184.235, Miami)
|
||||
| iptables DNAT: dst 64.92.84.81:20000 -> 127.0.0.1:9000
|
||||
v
|
||||
agave-validator TVU port (localhost:9000)
|
||||
```
|
||||
|
||||
## Production Config: laconic-was-sw01
|
||||
|
||||
### Pre-change safety
|
||||
|
||||
```
|
||||
configure checkpoint save pre-shred-relay
|
||||
```
|
||||
|
||||
Rollback: `rollback running-config checkpoint pre-shred-relay` then `write memory`.
|
||||
|
||||
### Config session with auto-revert
|
||||
|
||||
```
|
||||
configure session shred-relay
|
||||
|
||||
! ACL for traffic-policy match
|
||||
ip access-list SHRED-RELAY-ACL
|
||||
10 permit udp any any eq 20000
|
||||
|
||||
! Traffic policy: redirect matched packets to backbone next-hop
|
||||
traffic-policy SHRED-RELAY
|
||||
match SHRED-RELAY-ACL
|
||||
set nexthop 172.16.1.189
|
||||
|
||||
! Override ASIC punt-to-CPU for redirected traffic
|
||||
system-rule overriding-action redirect
|
||||
|
||||
! Apply to Et1/1 ingress
|
||||
interface Ethernet1/1
|
||||
traffic-policy input SHRED-RELAY
|
||||
|
||||
! Remove old monitor session and its ACL
|
||||
no monitor session 1
|
||||
no ip access-list SHRED-RELAY
|
||||
|
||||
! Review before committing
|
||||
show session-config diffs
|
||||
|
||||
! Commit with 5-minute auto-revert safety net
|
||||
commit timer 00:05:00
|
||||
```
|
||||
|
||||
After verification: `configure session shred-relay commit` then `write memory`.
|
||||
|
||||
### Linux cleanup on was-sw01
|
||||
|
||||
```bash
|
||||
# Kill socat relay (PID 27743)
|
||||
kill 27743
|
||||
# Remove Linux kernel route
|
||||
ip route del 186.233.184.235/32
|
||||
```
|
||||
|
||||
The EOS static route `ip route 186.233.184.235/32 172.16.1.189` stays (general reachability).
|
||||
|
||||
## Production Config: biscayne
|
||||
|
||||
### iptables DNAT
|
||||
|
||||
Traffic-policy sends normal L3-forwarded UDP packets (no mirror encapsulation). Packets arrive
|
||||
with dst `64.92.84.81:20000` containing clean shred payloads directly in the UDP body.
|
||||
|
||||
```bash
|
||||
sudo iptables -t nat -A PREROUTING -p udp -d 64.92.84.81 --dport 20000 \
|
||||
-j DNAT --to-destination 127.0.0.1:9000
|
||||
|
||||
# Persist across reboot
|
||||
sudo apt install -y iptables-persistent
|
||||
sudo netfilter-persistent save
|
||||
```
|
||||
|
||||
### Cleanup
|
||||
|
||||
```bash
|
||||
# Kill shred-unwrap.py (PID 2497694)
|
||||
kill 2497694
|
||||
rm /tmp/shred-unwrap.py
|
||||
```
|
||||
|
||||
## Verification
|
||||
|
||||
1. `show traffic-policy interface Ethernet1/1` — policy applied
|
||||
2. `show traffic-policy counters` — packets matching and redirected
|
||||
3. `sudo iptables -t nat -L PREROUTING -v -n` — DNAT rule with packet counts
|
||||
4. Validator logs: slot replay rate should maintain ~3.3 slots/sec
|
||||
5. `ss -unp | grep 9000` — validator receiving on TVU port
|
||||
|
||||
## What was removed
|
||||
|
||||
| Component | Host |
|
||||
|-----------|------|
|
||||
| monitor session 1 | was-sw01 |
|
||||
| SHRED-RELAY ACL (old) | was-sw01 |
|
||||
| socat relay process | was-sw01 |
|
||||
| Linux kernel static route | was-sw01 |
|
||||
| shred-unwrap.py | biscayne |
|
||||
|
||||
## What was added
|
||||
|
||||
| Component | Host | Persistent? |
|
||||
|-----------|------|-------------|
|
||||
| traffic-policy SHRED-RELAY | was-sw01 | Yes (startup-config) |
|
||||
| SHRED-RELAY-ACL | was-sw01 | Yes (startup-config) |
|
||||
| system-rule overriding-action redirect | was-sw01 | Yes (startup-config) |
|
||||
| iptables DNAT rule | biscayne | Yes (iptables-persistent) |
|
||||
|
||||
## Key Details
|
||||
|
||||
| Item | Value |
|
||||
|------|-------|
|
||||
| Biscayne validator identity | `4WeLUxfQghbhsLEuwaAzjZiHg2VBw87vqHc4iZrGvKPr` |
|
||||
| Biscayne IP | `186.233.184.235` |
|
||||
| laconic-was-sw01 public IP | `64.92.84.81` (Et1/1) |
|
||||
| laconic-was-sw01 backbone IP | `172.16.1.188` (Et4/1) |
|
||||
| laconic-was-sw01 SSH | `install@137.239.200.198` |
|
||||
| laconic-mia-sw01 backbone IP | `172.16.1.189` (Et4/1) |
|
||||
| Backbone RTT (WAS-MIA) | 25.4ms |
|
||||
| EOS version | 4.34.0F |
|
||||
Reference in New Issue
Block a user