← Research

ENGINEERING NOTE · OCTOBER 2026 · 6 MIN READ

Designing resilient network recovery

Connections fail in ordinary, everyday ways. A resilient system notices, responds sensibly and tells the truth about what it is doing.

Failure is normal

Phones move between Wi-Fi and mobile data, pass through areas of weak signal, join networks with captive sign-in pages and sit behind congested links. None of this is exceptional. A system that treats a dropped connection as a rare event will behave badly in daily use.

Recovery should be designed in from the start, not added after the first complaint.

A simple loop: detect, diagnose, act, verify

Principles that keep recovery safe

The hardest decision: fail open or fail closed

When a protection layer cannot work, should traffic be blocked or allowed? Blocking keeps people private but can leave them offline. Allowing keeps them connected but unprotected. There is no universal answer. The right choice depends on context, and it should be made deliberately and communicated clearly, not left to chance.

Keeping sessions alive across changes

Many connections break when a device changes network because they are tied to the addresses at each end. Newer transport designs, such as QUIC, can carry a session across a change of network, which is a good example of recovery handled at the protocol level rather than patched on top.

KEY TAKEAWAYS

  • Treat failure as normal and design recovery in advance.
  • Use a loop of detect, diagnose, act and verify.
  • Back off, avoid flapping and limit effort.
  • Decide fail-open or fail-closed on purpose, and tell people what is happening.

Have a question or a correction?