Includes: - package/ GL-XE3000 kit-busrouter (opkg) - scripts/provision.sh (GL) and provision-synology.sh (Synology) - syno-balance/ — new WAN balancer replacing aiwanbal (SmartWAN adapter) - kit-connect/ — unified connectivity SPK (Tailscale + reverse SSH) - docs/deployment/synology-rt2600ac-checklist.md — 62-point checklist - docs/provisioning/device-identity.md — fleet identity spec - docs/pilot/checklist.md — field pilot validation - x4078_20260721.dss — reference config backup Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
9.5 KiB
Handoff — WG mgmt incident + resume Task 2.1 (2026-07-01)
Audience: local Claude Code running on a machine with LAN access to the bench router.
Why this exists: the cloud session hit the router only over the WireGuard mgmt tunnel, which is
currently DOWN. The remaining work must be driven from the LAN side. This doc + the two scripts in
scripts/diag/ contain everything needed to (1) restore WG mgmt, (2) make the tunnel robust across
either WAN, then (3) resume Task 2.1.
1. Current project state
- Phase 0 (0.1 SSH, 0.2 recon, 0.3 package skeleton) — DONE, committed, plan boxes checked.
- Phase 1 (scoring/decision pure-shell libs + bats) — DONE, committed.
- Task 2.0 (install
mwan3 2.8.15-2from GL's own feed; disable GLkmwan) — DONE + independently verified. Commit4f33ca3. Recon doc §4. Default route + WG mgmt survived that change. - BLOCKER (this doc): WG mgmt tunnel to the bench router (
10.88.0.2) went dead the moment ATT fiber was plugged intoeth0. On-device Task 2.1 work is blocked until mgmt is restored.
Plan of record: docs/superpowers/plans/2026-06-30-pioneer-bus-router-balancer.md
Recon / ground truth: docs/recon/2026-06-30-xe3000-environment.md (see §5 for the fiber pre-check).
2. Environment & access
| Thing | Value |
|---|---|
| Router | GL.iNet GL-XE3000, OpenWrt 21.02-SNAPSHOT, busybox/ash, kernel 5.4 |
| Router LAN | 192.168.8.1 — ssh root@192.168.8.1, password kitPLANE1!! (dropbear) |
| Router WG mgmt IP | 10.88.0.2 (iface wgclient1, route table 1001) |
| Uplink 1 (cellular) | rmnet_mhi0 — T-Mobile CGNAT, default-route metric 40 |
| Uplink 2 (fiber, = "Starlink" stand-in) | eth0 / UCI iface wan — ATT fiber, metric 10 (now the preferred default). Public IP seen by hub: 99.163.124.171 |
| Fleet hub | DO droplet kit-fleet-hub, public 167.172.237.162, WG 10.88.0.1, wg iface wg0. Cloud session reaches it via mcp__lab-do__do_ssh_command droplet="kit-fleet-hub". |
| NOT the hub | 162.243.83.36 is a different lab server — ignore it for this work. |
Reachability rule: while WG is down, the LAN is the only way in. The hub has NO route to
192.168.8.1 (router LAN is private); and no interactive shell survives over the broken tunnel.
3. The incident — evidence gathered so far
Symptom: fiber plugged into eth0 → fiber (metric 10) became the default route → the WG underlay
roamed onto fiber (hub now sees the peer at 99.163.124.171:44881, was T-Mobile CGNAT 100.x).
Handshakes succeed both directions and are always "fresh", but no data flows — ICMP/SSH over
the tunnel get zero replies in either direction.
Verified CLEAN on the hub (167.172.237.162, checked directly via DO API):
wg showpeer allowed-ips =10.88.0.2/32✓ip route get 10.88.0.2→dev wg0 src 10.88.0.1✓;net.ipv4.ip_forward = 1✓- Firewall:
nft list rulesetempty; iptables INPUT/FORWARD/OUTPUT policy ACCEPT, 0 packets ✓ - Endpoint roam succeeded (hub has the fiber endpoint, fresh handshake).
- → The three "usual hub culprits" (allowed-ips, firewall, return route) are all ruled out.
Verified on the router (LAN side):
rp_filter= 0 on all interfaces → NOT the cause (an earlier hypothesis; discard it).- Routing correct:
ip route get 10.88.0.1 from 10.88.0.2→dev wgclient1 table 1001✓ wgclient1MTU = 1420 (WG default).- Source-bound ping (
ping -I 10.88.0.2 10.88.0.1) fails, butwgtransfer incremented only ~32 bytes — suspiciously small for 3 ICMP packets (~430B expected), so it's unclear the data packets actually hit the wire.
Measurement caveat: hub-side and router-side counters were sampled at different moments, so cross-side counter comparisons are unreliable. The §B capture in the diagnose script fixes this by capturing on the router itself while it generates traffic.
Ruled out: rp_filter (=0), hub firewall (none), hub return route (correct), hub allowed-ips (correct), MTU-for-tiny-pings (a 144B encapsulated ping ≈ the 148B handshake that succeeds — MTU only kills large packets like SSH data, not pings).
Leading hypotheses, in priority order:
persistent-keepalivemissing on the router'swgclient1peer → behind ATT fiber NAT the return mapping expires between handshakes (T-Mobile CGNAT likely masked this). Not yet confirmed — we never captured the router'swg show. Getwg showfirst.- Router-side cryptokey routing / peer AllowedIPs doesn't include
10.88.0.1→ inbound replies silently dropped by WG. Also revealed bywg show. - MTU/PMTU black-hole for large packets — unlikely to block pings but WILL break SSH data; worth
pre-empting with
wgclient1 mtu 1280+ MSS clamp.
4. DO THIS FIRST — restore mgmt (LAN, root@192.168.8.1)
Step 1 — diagnose (read-only, safe)
sh scripts/diag/wg-mgmt-diagnose.sh # or paste it; it's plain ash
The single most important output is wg show (§A) — we have never seen the router's peer config.
Look for:
persistent keepalive: (none)→ hypothesis #1 is live.allowed ips:on the hub peer — must include10.88.0.1(else hypothesis #2).- §B capture: does
[eth0]show inbound167.172.237.162.51820 →packets during the ping?- inbound present → replies arrive on fiber → decrypt/inner issue (chase AllowedIPs/keys).
- outbound only → router's packets aren't leaving, or hub isn't returning → coordinate a hub capture.
Step 2 — apply the fix (mutating; wg-mgmt-fix.sh is gated, read its header)
Default fix order (safe, reversible, and each independently supports "works over either WAN"):
sh scripts/diag/wg-mgmt-fix.sh keepalive # set persistent_keepalive=25 (try FIRST)
# re-test: ping -c3 -I 10.88.0.2 10.88.0.1 → expect replies
sh scripts/diag/wg-mgmt-fix.sh mtu # set wgclient1 mtu=1280 (if keepalive alone didn't fix)
sh scripts/diag/wg-mgmt-fix.sh rpfilter # persist loose rp_filter (already 0 at runtime; makes it durable)
Each is uci set … ; uci commit ; ifup wgclient1 — the tunnel bounces (~5s) but nothing else is touched.
Step 3 — verify restored
ping -c3 -W2 -I 10.88.0.2 10.88.0.1 # replies from the hub = tunnel data path healthy
When that replies, mgmt over WG is back. (The cloud session can then re-confirm from the hub with
ssh -i /root/.ssh/busfleet root@10.88.0.2 'ubus call system board' — but you can also just keep
driving from the LAN.)
5. Robustness goal — WG mgmt must survive on EITHER WAN
Bake these in (they're the durable fix, not just incident cleanup), then fold into the Phase 8 provisioning profile:
persistent_keepalive=25— holds the NAT mapping open on whichever WAN is active.wgclient1 mtu=1280— fits under cellular and ATT-fiber/PPPoE path MTUs; kills PMTU black-holes.- rp_filter loose (
=2) persisted via/etc/sysctl.conf; if GL re-tightens on WAN flap, add a/etc/hotplug.d/iface/hook that re-asserts it (the fix script writes one underrpfilter). - The mgmt return rule
from 10.88.0.2 lookup 1001is already default-route-independent — leave it. - After the fix, prove failover:
ifdown wan(fiber) → mgmt should ride cellular;ifup wan→ rides fiber again. Do this only once mgmt is confirmed healthy, and see the guardrail below.
6. THEN resume Task 2.1 (per the plan)
Follow superpowers:subagent-driven-development (implement → spec review → code-quality review →
commit → check the plan box), same as Tasks 0.3 / 2.0.
Task 2.1 files & specifics (from the plan):
package/kit-busrouter/files/usr/lib/busrouter/wan-mwan.sh- mwan3 config: members cellular (
modem_0001/rmnet_mhi0) + fiber-as-WAN2 (wan/eth0); policybalanced;track_ip1.1.1.1 & 8.8.8.8;reliability 1;down 3;up 3. mwan_set_weight <member> <weight>=uci set mwan3.<member>.weight+mwan3 restart, guarded byweight_changed(already implemented inlib-decide.sh) so we don't flush conntrack needlessly.- Rewrite mwan3's stock config — it ships referencing
wan/wwan/tethering, none of which match our live ifaces (noted in Task 2.0).
Failover drill — SAFE direction (critical):
- Fiber currently carries the WG mgmt underlay, so it is SAFE to bounce CELLULAR (
rmnet_mhi0) for the failover test, and NEVER downeth0/fiber while it carries mgmt. - Always re-check first:
ip route get 167.172.237.162tells you which uplink the WG underlay is on right now; down the OTHER one. (After you set persistent_keepalive + prove either-WAN failover, the mgmt-bearing uplink can shift — re-check every time.)
7. Guardrails (learned this session)
- No dead-man watchdog on this firmware:
setsidandnohupare both absent, andstart-stop-daemon -x /bin/shself-matches the session shell. So you cannot arm a reliable auto-rollback over SSH. Mitigation: do risky routing changes from the LAN side (you're physically present) so a mistake doesn't lock you out; verify mgmt after every mutation. - No
base64on the router — transfer files by decoding on the hub/LAN box and piping cleartext intocat >on the router, or just paste scripts. - busybox
pinghas no-M do— can't do a clean DF/PMTU probe with it; infer MTU issues from large-vs-small behavior ortracepathif present. - Verify mgmt access (a WG-sourced ping or SSH) after every routing/firewall/mwan3 change.