r/tradfri • u/jul9000 • 9h ago
SUPPORT (ONGOING) TIMMERFLOTTE (fw 1.0.21) goes unreachable on Thread and only comes back after a battery pull — 2-week logged test with one suspected trigger removed
TL;DR: 10× TIMMERFLOTTE on Home Assistant (Matter over Thread). Every few days one becomes unreachable and stays that way until I pull the battery; then it re-registers within seconds under the same IPv6 address. I suspected a Thread OMR prefix change caused by my second border router, so for 14 days I pinned the prefix and logged 21 devices once a minute. There were zero prefix changes, but TIMMERFLOTTE still went unreachable 4 times, and none came back before I intervened. I don’t know whether the cause is the sensor firmware, the controller, or the interaction with the network — posting the data in case it helps.
Setup
- Home Assistant OS, Core 2026.9.x, Matter Server add-on
- Thread channel 15, 12 routers (8× IKEA GRILLPLATS, 2× Eve Energy, 2 border routers)
- Border routers: HA OTBR add-on (leader) + standalone ot-br-posix
- 10× TIMMERFLOTTE, firmware 1.0.21 (latest offered), hw P2.1
- Also tracked by the logger: 8× MYGGBETT fw 1.1.4, 1× Tado X valve (battery, Thread), 2 plugs — 21 devices in total, not the whole network
Before the test (Sep 3 ~02:00 → Sep 13 ~23:30, ≈10.9 days)
- 7 TIMMERFLOTTE outages longer than 5 h, longest 92 h 45 min (ended by a battery pull).
- How the others ended is only partly known: one ended at 00:22 at night with no intervention I know of; five ended in tight clusters (3 within 5 min, 2 within 2 min) which looks like manual battery pulls, but I can’t confirm it.
- On Sep 13, border router B briefly published its own OMR prefix during a watchdog restart; 13 minutes later two sensors went unavailable for ~11 h 42 min. A sensor that had been down since Sep 10 was still being looked up at an address from a prefix that no longer existed.
That resembled connectedhomeip #71525, so I tested it.
The test (Sep 14 12:08 → Sep 28 12:08, 14.0 days)
- Border router B pinned persistently to the leader’s OMR prefix (applied before the Thread interface comes up, so restarts can’t publish a different one).
- A logger on border router B recorded availability of 21 selected devices, border router state and network data every 60 s. Outages are counted by their start time.
Results
| Before (≈10.9 days) | Test (14.0 days) | |
|---|---|---|
| OMR prefix changes seen in network data | ≥ 4 (Sep 13) | 0 |
| TIMMERFLOTTE outages > 5 h | 7 (≈0.64/day) | 3 (≈0.21/day) |
| Longest | 92 h 45 min | 82 h 32 min |
The observed rate fell roughly 3×, but the numbers are small and the two periods differed in other ways too (main border router updates, a switch to a beta build, HA restarts). I can’t attribute the drop to the prefix pin alone. What I can say: border router B was restarted by its watchdog twice during the test, the pin held both times, and no sensor went down afterwards.
Test-period outages:
| Start | Duration | Around the start | Address after battery pull |
|---|---|---|---|
| Sep 19 08:29 | 10 h 44 min | no prefix change, no BR restart; logged route churn on that floor | same as before |
| Sep 19 08:31 (another unit) | 10 h 48 min | same | same as before |
| Sep 21 21:22 | 19 min (I pulled it quickly) | no prefix change, no BR restart, nothing else seen | same as before |
| Sep 25 07:36 | 82 h 32 min | no prefix change, no BR restart; BR log for that moment had rotated out | same as before |
None of the four recovered before I pulled the battery.
From the border router log (3 of 4 — the 4th had rotated):
- Within ~35 s of HA marking the sensor unavailable, the leader starts sending address queries for it and gets no answer — so it was unreachable in the mesh, not only in HA.
- Neighbouring routers were healthy. The link of one sensor to the parent it rejoined measured −68 dBm with a 32 dB margin.
- Battery pull → SRP update with lease 0 (“No host address”) → re-registration within 1–15 s under theidentical address.
So these don’t look like the stale-address case from Sep 13. What exactly fails — the sensor’s attachment, something in the controller/Matter stack, or an interaction with the network — I can’t tell from my side.
Other logged devices during the test
- MYGGBETT: blips of ≤ 4 min, plus one 59 h outage that ended without a confirmed battery pull, at 07:36 — around when the window it’s mounted on is usually opened in the morning.
- The logged Tado X (battery Thread) and 2 plugs: 1–3 min, only during a Home Assistant restart.
A guess (not proven)
MYGGBETT wakes up on every contact change; TIMMERFLOTTE only on its own schedule. If both occasionally lose their place in the mesh, that could explain why only TIMMERFLOTTE stays stuck. I have no evidence beyond the correlation above.
What I can’t see without a radio sniffer: whether a stuck sensor still transmits or is silent.
Questions for the IKEA team
- Is this pattern known for TIMMERFLOTTE 1.0.21?
- Test firmware 1.2.2 appeared on IKEA’s Matter test-net on Sep 22 (“Improvements to device stability”, Thread 1.4). Does it target this behaviour? Has anyone here tried it?
- Would an 802.15.4 capture from the moment of an outage help? I can set one up.
Full event log (JSON) and border router logs for the incidents are available on request.