Summary
Device models three teardowns and implements one. There is no way for a producer to say "I am dying, not shutting down", and no way to tear down without announcing $state=disconnected. Every caller that needs either has to reach around Device.stop() to the concrete client, and that reach-around is where the bugs are.
| Teardown |
Correct $state |
SDK today |
| Graceful shutdown |
disconnected |
Device.stop() |
| Ungraceful death (crash, power loss) |
lost |
the will, which only fires on an unclean disconnect |
| Deliberate death (fatal error handler, or a simulator acting the part) |
lost |
nothing |
Device.stop() publishes disconnected unconditionally:
# homie.py:1651
def stop(self, *, flush_timeout: float = 1.0, stop_timeout: float = 2.0) -> None:
# homie.py:1684-1690
if mqttc.is_connected():
root._state = DeviceState.DISCONNECTED
...
There is no announce= or equivalent on the teardown path, and DeviceState.LOST is never published anywhere in homie.py except inside the will() descriptor at line 2255. So a producer that knows it is failing has exactly two options: announce disconnected, which is a lie, or reach past the SDK.
Why this is not a simulator-only concern
It reads like one, because a simulator is the obvious consumer. But consider a producer whose supervisor is about to kill it after a fatal error, or one whose hardware backing the tree has gone away. It is dying. disconnected tells consumers this was orderly and expected; lost tells them something broke. A consumer rendering $state into availability treats those very differently, and the SDK's own doc/consuming-a-homie-tree.md is emphatic that consumers MUST react to every transition.
Today such a producer cannot tell the truth, and Device.stop() will actively misreport on its behalf.
What the gap has already cost
ebus-panel-sim wanted the third row for stop(graceful=False). With no API for it, it reached around Device.stop() to mqttc.stop(). Four separate bugs are downstream of that one reach-around, all found and fixed in the last two days:
- It did not type-check.
Device.mqttc is MqttDeviceTransport, which deliberately omits stop because that is owned-only. main was red until panel-sim#7.
- It would have torn down a caller-injected client. Ownership decides, not type: an injected client can itself be an
MqttClient driven by asyncio_driver, so an isinstance narrowing is not sufficient.
- The will was suppressed, so the tree said
ready forever. MqttClient.stop() performs a clean DISCONNECT by design, which discards the will. Verified against a real broker: a consumer joining after an ungraceful teardown read the whole retained tree as ready, indefinitely. Fixed in panel-sim 0.3.1 by publishing the will's own payload.
- That publish left the Device's
_state on ready, so a later refresh_tree() republished ready over the lost. Fixed in panel-sim 0.3.3 by routing through Device.set_state(DeviceState.LOST).
Each fix was correct locally. Collectively they are a downstream package reimplementing teardown because the upstream one had no way to express what it meant. panel-sim#17, currently open, is carrying still more of this: an owned/injected split inside the reach-around, and an unresolved question about flushing on a caller-driven event loop.
Suggested shape
Not prescribing an API, but the two things that would remove the need for any of the above:
A way to announce death. Something like Device.declare_lost(): set _state to LOST, publish it retained through whichever transport the tree has, flush on the owned path. It is Device.stop()'s existing disconnected block with a different state and no close.
A way to tear down without announcing, for the caller who wants to publish something else first (or nothing at all). Either a parameter on stop() or a separate owned-only close.
Two details worth carrying over, since both cost real time to discover downstream:
- The state move and the publish must happen together. Publishing a state the
Device does not hold means any later re-announce silently undoes it.
- On an injected transport the message is queued on the caller's loop and
publish_and_flush is not available (it is owned-only, off MqttDeviceTransport). Blocking there would stall the loop that has to drain it, so the caller obligation needs stating rather than solving.
Why file it now
ebus-panel-sim has a working, tested local implementation today, so nothing is blocked. But it intends to adopt the SDK API and delete its own once this lands: the local version exists only because there was nothing to call, and every one of the four bugs above came from maintaining it. That gives this a committed consumer rather than a hypothetical one.
The evidence that this is a wall rather than a preference is that one downstream took four attempts to get over it, and a second simulator (ebus-utility-meter-sim) is being scaffolded now that would otherwise start the same climb.
Summary
Devicemodels three teardowns and implements one. There is no way for a producer to say "I am dying, not shutting down", and no way to tear down without announcing$state=disconnected. Every caller that needs either has to reach aroundDevice.stop()to the concrete client, and that reach-around is where the bugs are.$statedisconnectedDevice.stop()lostlostDevice.stop()publishesdisconnectedunconditionally:There is no
announce=or equivalent on the teardown path, andDeviceState.LOSTis never published anywhere inhomie.pyexcept inside thewill()descriptor at line 2255. So a producer that knows it is failing has exactly two options: announcedisconnected, which is a lie, or reach past the SDK.Why this is not a simulator-only concern
It reads like one, because a simulator is the obvious consumer. But consider a producer whose supervisor is about to kill it after a fatal error, or one whose hardware backing the tree has gone away. It is dying.
disconnectedtells consumers this was orderly and expected;losttells them something broke. A consumer rendering$stateinto availability treats those very differently, and the SDK's owndoc/consuming-a-homie-tree.mdis emphatic that consumers MUST react to every transition.Today such a producer cannot tell the truth, and
Device.stop()will actively misreport on its behalf.What the gap has already cost
ebus-panel-simwanted the third row forstop(graceful=False). With no API for it, it reached aroundDevice.stop()tomqttc.stop(). Four separate bugs are downstream of that one reach-around, all found and fixed in the last two days:Device.mqttcisMqttDeviceTransport, which deliberately omitsstopbecause that is owned-only.mainwas red until panel-sim#7.MqttClientdriven byasyncio_driver, so anisinstancenarrowing is not sufficient.readyforever.MqttClient.stop()performs a clean DISCONNECT by design, which discards the will. Verified against a real broker: a consumer joining after an ungraceful teardown read the whole retained tree asready, indefinitely. Fixed in panel-sim 0.3.1 by publishing the will's own payload._stateonready, so a laterrefresh_tree()republishedreadyover thelost. Fixed in panel-sim 0.3.3 by routing throughDevice.set_state(DeviceState.LOST).Each fix was correct locally. Collectively they are a downstream package reimplementing teardown because the upstream one had no way to express what it meant. panel-sim#17, currently open, is carrying still more of this: an owned/injected split inside the reach-around, and an unresolved question about flushing on a caller-driven event loop.
Suggested shape
Not prescribing an API, but the two things that would remove the need for any of the above:
A way to announce death. Something like
Device.declare_lost(): set_statetoLOST, publish it retained through whichever transport the tree has, flush on the owned path. It isDevice.stop()'s existingdisconnectedblock with a different state and no close.A way to tear down without announcing, for the caller who wants to publish something else first (or nothing at all). Either a parameter on
stop()or a separate owned-only close.Two details worth carrying over, since both cost real time to discover downstream:
Devicedoes not hold means any later re-announce silently undoes it.publish_and_flushis not available (it is owned-only, offMqttDeviceTransport). Blocking there would stall the loop that has to drain it, so the caller obligation needs stating rather than solving.Why file it now
ebus-panel-simhas a working, tested local implementation today, so nothing is blocked. But it intends to adopt the SDK API and delete its own once this lands: the local version exists only because there was nothing to call, and every one of the four bugs above came from maintaining it. That gives this a committed consumer rather than a hypothetical one.The evidence that this is a wall rather than a preference is that one downstream took four attempts to get over it, and a second simulator (
ebus-utility-meter-sim) is being scaffolded now that would otherwise start the same climb.