Skip to content

No way to announce death: Device models three teardowns and implements one #46

Description

@dcj

Summary

Device models three teardowns and implements one. There is no way for a producer to say "I am dying, not shutting down", and no way to tear down without announcing $state=disconnected. Every caller that needs either has to reach around Device.stop() to the concrete client, and that reach-around is where the bugs are.

Teardown Correct $state SDK today
Graceful shutdown disconnected Device.stop()
Ungraceful death (crash, power loss) lost the will, which only fires on an unclean disconnect
Deliberate death (fatal error handler, or a simulator acting the part) lost nothing

Device.stop() publishes disconnected unconditionally:

# homie.py:1651
def stop(self, *, flush_timeout: float = 1.0, stop_timeout: float = 2.0) -> None:

# homie.py:1684-1690
if mqttc.is_connected():
    root._state = DeviceState.DISCONNECTED
    ...

There is no announce= or equivalent on the teardown path, and DeviceState.LOST is never published anywhere in homie.py except inside the will() descriptor at line 2255. So a producer that knows it is failing has exactly two options: announce disconnected, which is a lie, or reach past the SDK.

Why this is not a simulator-only concern

It reads like one, because a simulator is the obvious consumer. But consider a producer whose supervisor is about to kill it after a fatal error, or one whose hardware backing the tree has gone away. It is dying. disconnected tells consumers this was orderly and expected; lost tells them something broke. A consumer rendering $state into availability treats those very differently, and the SDK's own doc/consuming-a-homie-tree.md is emphatic that consumers MUST react to every transition.

Today such a producer cannot tell the truth, and Device.stop() will actively misreport on its behalf.

What the gap has already cost

ebus-panel-sim wanted the third row for stop(graceful=False). With no API for it, it reached around Device.stop() to mqttc.stop(). Four separate bugs are downstream of that one reach-around, all found and fixed in the last two days:

  1. It did not type-check. Device.mqttc is MqttDeviceTransport, which deliberately omits stop because that is owned-only. main was red until panel-sim#7.
  2. It would have torn down a caller-injected client. Ownership decides, not type: an injected client can itself be an MqttClient driven by asyncio_driver, so an isinstance narrowing is not sufficient.
  3. The will was suppressed, so the tree said ready forever. MqttClient.stop() performs a clean DISCONNECT by design, which discards the will. Verified against a real broker: a consumer joining after an ungraceful teardown read the whole retained tree as ready, indefinitely. Fixed in panel-sim 0.3.1 by publishing the will's own payload.
  4. That publish left the Device's _state on ready, so a later refresh_tree() republished ready over the lost. Fixed in panel-sim 0.3.3 by routing through Device.set_state(DeviceState.LOST).

Each fix was correct locally. Collectively they are a downstream package reimplementing teardown because the upstream one had no way to express what it meant. panel-sim#17, currently open, is carrying still more of this: an owned/injected split inside the reach-around, and an unresolved question about flushing on a caller-driven event loop.

Suggested shape

Not prescribing an API, but the two things that would remove the need for any of the above:

A way to announce death. Something like Device.declare_lost(): set _state to LOST, publish it retained through whichever transport the tree has, flush on the owned path. It is Device.stop()'s existing disconnected block with a different state and no close.

A way to tear down without announcing, for the caller who wants to publish something else first (or nothing at all). Either a parameter on stop() or a separate owned-only close.

Two details worth carrying over, since both cost real time to discover downstream:

  • The state move and the publish must happen together. Publishing a state the Device does not hold means any later re-announce silently undoes it.
  • On an injected transport the message is queued on the caller's loop and publish_and_flush is not available (it is owned-only, off MqttDeviceTransport). Blocking there would stall the loop that has to drain it, so the caller obligation needs stating rather than solving.

Why file it now

ebus-panel-sim has a working, tested local implementation today, so nothing is blocked. But it intends to adopt the SDK API and delete its own once this lands: the local version exists only because there was nothing to call, and every one of the four bugs above came from maintaining it. That gives this a committed consumer rather than a hypothetical one.

The evidence that this is a wall rather than a preference is that one downstream took four attempts to get over it, and a second simulator (ebus-utility-meter-sim) is being scaffolded now that would otherwise start the same climb.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions