|
Home IO Control
ESPHome add-on for IO-Homecontrol devices
|
Status: Accepted · Recorded: 2026-08
Every blocking wait in this project — six of them, across ExchangeEngine and PairingEngine — is waiting for one of exactly two things, and which one determines where a reply can arrive:
Before this refactor, that distinction was expressed six different ways, in six separately evolved wait loops, and two of the six drivers had independently rediscovered it and encoded it as a chip-specific number instead of a loop-specific policy:
Both comments describe the same protocol fact — a unicast wait should never hop — worked around at the driver layer because the loop itself offered no way to say it. The workaround also leaked into the one loop that genuinely needs to rotate: collect_broadcast_responses() reused the same per-chip accessor, so on LR1121 the broadcast roll-call inherited "never hop" from a constant that existed to serve an unrelated unicast loop — about three dwells across a 2000 ms window instead of the many short ones a rotating listen needs.
Channel policy is decided by what kind of reply a listen is waiting for, not by which chip is listening. Three named policies, each used by exactly the loops whose reply shape it matches:
| Policy | Behaviour | Used by |
|---|---|---|
| HOLD_REQUEST_CHANNEL | Never retunes, never dwells; one wait_for_packet() call spans the whole remaining window. | wait_for_key_challenge_(), wait_for_key_confirm_() (pairing); wait_for_first_response_(), wait_for_final_response_() (every command exchange) |
| ROTATE_ALL_CHANNELS | CH1→CH2→CH3→CH1, starting on the request channel. | collect_broadcast_responses() for the roll-call's low-power pass |
| ROTATE_SKIPPING_REQUEST | The two channels that are not the request channel, leaving it before the first dwell. | wait_for_discovery_response_() (pairing discovery); collect_broadcast_responses() for the roll-call's always-alive pass |
All three share one primitive, ExchangeEngine::listen(), parameterised by a ListenSpec. The policy is chosen once, at the call site, by the caller stating which kind of reply it is waiting for — never by asking the driver. The roll-call is two frames with two reply populations, so it states two policies, one per pass. Whether a rotating listen hops away after an ignored frame, or stays put after a reception, is a separate ListenSpec setting of each loop: discovery hops on, because its window is full of unrelated traffic, while the roll-call stays where responders already are.
A consequence worth stating plainly: a holding listen needs no dwell at all. Slicing exists so a rotating loop gets a chance to hop between channels and so a long silent wait keeps feeding the watchdog; wait_for_packet() already feeds the watchdog internally while it blocks, so a HOLD_REQUEST_CHANNEL listen has nothing to gain from slicing and every expired slice on the soft-PHY chips costs a full RX re-arm (standby → clear IRQ → buffer base → packet params → RX). This is what let the two chip-specific "exchange dwell" constants disappear entirely rather than being folded into one shared value — there was never a real number to share, because a holding wait does not need one.
What survives as the one remaining chip-specific wait-timing knob is RadioDriver::hop_dwell_ms() (renamed from discovery_hop_slice_ms()): the dwell a rotating listen spends per channel before moving on. That is a chip question — how long must this radio sit on a channel after retuning before it can hear anything at all — and it now answers it for every rotating listen, discovery and both roll-call passes alike, through the same user-facing tuning field each driver already exposed.
The per-chip spread is large (5 ms on SX1276, 200 ms on SX1262 and LR1121) because retuning costs wildly different amounts on the two families, which is visible in the drivers themselves. The SX1276 changes channel with three register writes and never leaves RX, so a hop is essentially free. The SX1262 and LR1121 share SoftPhyDriverBase, whose change_frequency() must go standby → set frequency → clear stale IRQ/DIO latches → re-enter RX; hopping every few milliseconds there would spend most of the window re-arming the receiver instead of listening with it.
A dwell has to clear two independent floors, and only one of them is a chip property. The move above is what exposed this, and it is the most reusable thing in this ADR:
hop_dwell_ms() answers the first floor only. SX1262 and LR1121 clear the second one incidentally, because 200 ms is far longer than any frame. SX1276's 5 ms does not clear it at all — the fast dwell was only ever safe because pairing discovery paired it with a preamble/sync guard that extends a dwell in progress rather than cutting a reception off at the channel boundary. The roll-call inherited the fast dwell without inheriting the guard, and hardware measurement was unambiguous: scan_paired_devices found-rate over 25 scans collapsed from an 83.3% baseline to 12%. Giving collect_broadcast_responses() the same guard restored it to 25/25 scans finding every device, better than the pre-refactor baseline.
The rule this leaves behind: a rotating listen may dwell for less than one frame's air time only if it lingers on a detected preamble or sync word. Fast hopping and preamble gating are one design, not two independent choices — which is also how the implementations this 5 ms value was modelled on do it.