RoonServer hangs twice daily on Innuos Zenith Mk.3 with TIDAL (ref#K9517L)

What app are you having the slowness issue with?

· Roon

What kind of performance/speed issue are you experiencing?

· Other

Please try to reboot your Roon Server

· Yes, rebooting helps, but the issue returns after some time

Please try to reboot your networking gear (Router/Switches/etc.)

· No, the issue is still the same even after a reboot

Is there any change in behavior if you try to navigate to Roon Settings -> Library and set both Background and On-Demand Audio Analysis to Throttled or Off?

· No, the issue is still the same

Does the issue happen on multiple Roon Remotes (controllers) or just one?

· I only have one Roon remote to test with

Please try to restart your Roon Remote (controller) app

· No, the issue is still the same even after a restart

What is the operating system of your Roon Remote (controller)?

· Mac

Reinstall Windows/MacOS Roon Remote App

· No, I am still having the issue even after reinstalling

Router Domain Name System (DNS) change

· I don't know how to do this

What is the operating system of your Roon Server host machine?

· Prebuilt Branded Roon Server (SonicTransporter, Innous, Merging...)

Timestamp of issue occurrences

· Two occurrences, both captured server-side. The timestamps below are exactly as they appear in RoonServer_log.txt on the Core, so they should line up directly with the logs when you read them. Occurrence 1 — 30 July 2026, 18:32:15 Last log line before the silence, mid-playback: 18:32:14 [] [HighQuality 35.6x, 16/44 MQA TIDAL FLAC => 16/44] [100% buf] [PLAYING @ 1:46/3:03] 18:32:15 [stats] ... 399 Handles, 88 Threads Nothing is written after that line. In this occurrence port 9330 also stopped accepting TCP connections entirely. Occurrence 2 — 31 July 2026, 01:17:31 Freeze began while the Core was LOADING a track: 01:17:14 [] [HighQuality, 16/44 TIDAL FLAC => 16/44] [LOADING @ 0:00] 01:17:14 [easyhttp] GET api.tidal.com/v1/tracks//playbackinfopostpaywall ...audioquality=HI_RES_LOSSLESS... returned 200 in 24 ms 01:17:16 [stats] ... 698 Handles, 73 Threads 01:17:31 [stats] ... 698 Handles, 90 Threads Nothing after that. Thread count rose 73 -> 90 in fifteen seconds, and those two [stats] lines are 15 s apart instead of the usual ten-minute cadence. In this occurrence port 9330 kept serving cached images normally (a real 166 KB cover in 42 ms, an hour into the freeze) while playback and the Remoting ports were dead. I only noticed occurrence 2 at about 02:10, when music would not start; the freeze itself had begun at 01:17:31 with nobody touching the system. Recovery both times: a full restart of the appliance. After the restart that followed occurrence 1, the Core ran for exactly 5 h 42 min before occurrence 2. One caveat on clocks, so you are not misled: the appliance's timezone is set to Europe/London, while I am in Europe/Paris (CEST). If your reading of the logs needs my local wall-clock time, add one hour to the values above. The relative intervals and the log-internal sequence are of course unaffected.

Describe the issue

Hello,

RoonServer hangs on my Core roughly twice a day. I have captured server-side logs
for two consecutive occurrences and the signature is consistent, so I think this
is a deadlock rather than a crash.

SETUP
Core RoonServer build 1671, platform linuxx64
running on an Innuos Zenith Mk.3 appliance (no shell access)
Library 100% TIDAL streaming, zero local files
Remotes Roon on macOS
Endpoint Meridian Explorer² USB DAC attached to the Innuos
Network wired path, 3-9 ms to the Core, 0% packet loss measured
continuously (44 consecutive probes) around the incidents

WHAT HAPPENS
RoonServer_log.txt simply stops. No exception, no stack trace, no restart, no
further line at any log level. The process stays up and keeps its listening
sockets bound, but playback never starts and remotes can no longer load anything.

OCCURRENCE 1 — 2026-07-30 18:32:15 (mid-playback)
Last lines:
18:32:14 Trace: [] [HighQuality 35.6x, 16/44 MQA TIDAL FLAC => 16/44]
[100% buf] [PLAYING @ 1:46/3:03]
18:32:15 Info: [stats] 20888mb Virtual; 2691mb Physical = 1411mb GC-committed
(1047mb Managed-live = 74% of committed) + 1280mb Native;
399 Handles, 88 Threads
Nothing after that. [stats] normally appears every ~10 minutes (3124 entries in
that file alone), so the process went silent for over an hour.
In this occurrence the image server on port 9330 stopped accepting TCP
connections entirely — connections were dropped, never refused.

OCCURRENCE 2 — 2026-07-31 01:17:31 (while LOADING a track)
Last lines:
01:16:46 Trace: [] [HighQuality, 16/44 TIDAL FLAC => 16/44] [STOPPED @ 0:00]
01:17:14 Trace: [] [HighQuality, 16/44 TIDAL FLAC => 16/44] [LOADING @ 0:00]
01:17:14 Debug: [easyhttp] GET https://api.tidal.com/v1/tracks//playbackinfo
postpaywall?...audioquality=HI_RES_LOSSLESS... returned 200
01:17:14 Trace: [tidal/http] ... => Success
01:17:16 Info: [stats] ... 698 Handles, 73 Threads
01:17:31 Info: [stats] ... 698 Handles, 90 Threads
Nothing after that.

Two things stand out here:
- Thread count rose from 73 to 90 in fifteen seconds, and two [stats] lines
landed 15 s apart instead of the usual 10 minutes — something was spinning up
threads right before the freeze.
- Unlike occurrence 1, the image server on 9330 kept working perfectly
afterwards: it served a real 166 KB cover in 42 ms an hour into the freeze,
while the Remoting ports and all playback were dead. So parts of the process
are alive while the core loops are stuck.

TRIAGE STEPS ALREADY COVERED
- Rebooting the Core fixes it, and it comes back: the restart at 19:35 bought
exactly 5h42 of uptime before the next freeze at 01:17.
- Rebooting the networking gear changes nothing: the router was restarted at
23:55 that same night and the Core still froze at 01:17.
- Background and On-Demand Audio Analysis set to Off. Worth noting that my
library is 100% TIDAL streaming with zero local files, so audio analysis has
nothing to process here in the first place.
- Restarting the remote app changes nothing, as expected.
- DNS: I cannot change the router's DNS servers (the gateway is a managed
firewall and I do not hold its admin credentials), but I measured DNS
instead of guessing: 30 out of 30 lookups succeeded against the gateway
resolver, 28 ms average, 32 ms worst. More to the point, at the exact moment
of freeze 2 the Core resolved and reached api.tidal.com in 24 ms and got a
200. DNS was healthy while the Core was hanging.
- I have NOT reinstalled the remote app, and I want to be transparent about
why rather than claim a step I did not take: the fault is captured
server-side. RoonServer's own log stops mid-line on the Core while the
process stays up. Occurrence 2 began while the Core was loading a track. No
change to a controller on another machine can act on that, so a reinstall
would not test anything. I will of course do it if you still want it ruled
out formally.

WHAT IT IS NOT
- Not TIDAL: every TIDAL API call at the moment of the freeze returned 200 in
24-29 ms, including the playbackinfo request.
- Not the network: 0% loss, 3-9 ms to the Core, and the appliance's own web
interface answered in 11 ms during both freezes.
- Not resource exhaustion: 399 and 698 handles respectively, memory flat for
hours beforehand. No "Too many open files", nothing in the log at all.
- Not the hardware: the appliance itself stayed fully responsive throughout.

FREQUENCY
20 Jul 22:00 → 22 Jul 12:00 outage
22 Jul 17:00 → 29 Jul 12:00 seven clean days, ~1900 successful image requests,
zero failures
29 Jul 13:00 → 29 Jul 16:00 outage
30 Jul 18:32 freeze (occurrence 1)
31 Jul 01:17 freeze (occurrence 2), ~5h40 after a restart

RECOVERY
Only a full restart of the appliance. Restarting the Roon remote does nothing,
as expected.

PRIOR REPORTS THAT LOOK RELATED
I searched your community before writing, and three existing threads share parts
of this signature:
- "Playback stops and continues sometimes later, mostly on MQA TIDAL tracks"
community.roonlabs.com/t/.../85727
- "A memory leak in RoonAppliance" / "Since build 952 a real big memory leak
(Linux)" — threads climbing, server becoming unresponsive with no error
community.roonlabs.com/t/.../236489 and /202771
- "Roon Server crashes and restarts with no indication of trouble in the log
files" — nothing at app level, fault only visible at system level
community.roonlabs.com/t/.../174763
I mention them in case they help you correlate; my case differs in that the
process does not crash or restart — it stays up and silent indefinitely.

QUESTIONS
1. Does this signature match a known deadlock in build 1671 on linuxx64?
2. Both freezes occurred during TIDAL track transitions, and occurrence 2 was
requesting HI_RES_LOSSLESS. Given the MQA-related playback threads above, is
switching TIDAL streaming quality away from MQA to hi-res FLAC a sensible
thing for me to try as a diagnostic? If so I will run it and report back.
3. I have the complete log archives from immediately after each freeze. Where
should I upload them?

Happy to run any diagnostic you need — the fault reproduces roughly twice a day.

Thank you,

Describe your network setup

ISP: XEFI (Lyon, France), fibre.

Gateway / firewall: a Sophos UTM 9 appliance. It is managed by a third-party IT
integrator and I do not hold its admin credentials, so I cannot inspect or change
its configuration myself. It also acts as the LAN's DHCP server (24-hour leases)
and DNS resolver. Worth flagging: this model reached end of life on 30 June 2026.

LAN: a single flat /24 subnet, no VLANs.

Wi-Fi: a TP-Link access point. My Roon remote (macOS) connects over Wi-Fi at
5 GHz on a 160 MHz channel, -59 dBm, 816 Mbps negotiated link rate.

Roon Server host: Innuos Zenith Mk.3 in Roon Core mode.
[TO CONFIRM: the Innuos is connected by Ethernet / Wi-Fi — please correct]

Audio endpoints:
- Meridian Explorer² USB DAC attached to the Innuos
- LUMIN U1 MINI as a network endpoint
- occasional AirPlay devices

Measurements taken around both freezes, so you do not have to take "the network
is fine" on faith:
- 0% packet loss over 44 consecutive probes to the gateway, to the Core, and
to the public internet
- ~75 Mbps download throughput
- DNS: 30 of 30 lookups succeeded, 28 ms average, 32 ms worst
- at the exact moment of the second freeze, the Core reached api.tidal.com in
24 ms and received a 200

Note on jitter: round-trip times from the Mac to every LAN device (gateway, Core,
LUMIN) show the same profile, ~4-5 ms minimum with occasional 85-100 ms spikes.
Since it is identical to all three targets, that variance comes from my Mac's
own Wi-Fi link, not from the Core.

One relevant network event, for completeness: the router went down entirely one
night and was restarted at 23:55. The Core still froze at 01:17 afterwards, so
the freeze is not downstream of that outage.

Hi @jc1,

Thank you for the thorough report. We’ve reviewed this case and have a few follow-up questions. It’s likely that this is something you can resolve, or at least alleviate, for this setup in the meantime.

The “log goes silent” symptom is log rotation, not a hang. Roon rolls the file over into .01, .02, etc., and the end of a file looks exactly like the process dying. In your 01:17 case, logging appears to have actually continued, then the Roon Server gets a normal restart at 01:19:44. Nothing froze, unless you’re referring to the GUI itself (these would be Roon logs, rather than the Innuos Roon Server logs).

The real event is a crash at 09:40:58, where the server ran out of worker threads. Right before it, your own extension (com.jc.roon-display) got stuck in a loop: it re-registers using one saved token, and each new registration kicks off the previous one, so it never settles. It did that ~12,000 times in about two and a half minutes, and that’s what exhausted the server.

Two of your other sessions ran 3 to 8 hours with the extension not connected, and both stayed perfectly healthy. So we’re fairly confident it’s the extension.

We have a brief test we’d like you to try that will equip us with more precise logging from here. Stop the extension, restart the Roon Server, and leave it off for 12 hours (needs to be that long, since it ran 8 hours before crashing last time). Then tell us two things separately:

  • Did it crash or restart on its own? (We expect no.)
  • Did music still cut out? (We expect it might, see below.)

Those are two different problems. The music dropouts look like your LUMIN U1 MINI running over AirPlay instead of Roon Ready. It drops and stutters in the logs even when the extension isn’t involved. Roon already recognizes your U1 MINI as a RoonReady device, so you can switch it in Settings → Audio. Do that after the 12-hour test so we only change one thing at a time.

A couple of quick answers: no need to mess with MQA/TIDAL quality, and no need to reinstall the remote, since you were right that the fault is server-side.

Once you’ve got the 12-hour result, we’re happy to help sort out the extension’s re-registration logic.

Thank you and we’ll watch for your reply.

Hi @connor,

Thanks — and you’re right about the log: the appliance is on Europe/London while
I’m in Europe/Paris, so what I read as an hour of silence was a one-hour offset.
The restart at 01:19:44 was mine. My mistake, noted.

I counted the registrations in the Core logs myself and your finding holds — peak
4707 in a single minute. I also found the defect in my extension: core_unpaired()
and the websocket’s onclose() each schedule a reconnect on a fixed 10s timer, with
no tracking of pending timers, so pending reconnects double every 10 seconds. A
fix is written but deliberately not deployed until after the test window.

Test started: extension stopped and Roon Server restarted at 20:50 local, 31 July.
The LUMIN stays on AirPlay for now, so only one variable changes.

I’ll come back with your two answers separately after 08:50 local tomorrow.

:folded_hands:

Hi Connor,

Thank you — that was a much sharper read of the logs than mine. Two corrections
I’ve taken on board: the “log goes silent” was rotation, not a hang, and I also
hadn’t accounted for the appliance running on Europe/London while I’m on CEST,
so my timestamps were an hour off. The 01:19:44 restart was mine, not a crash.

The 12-hour test is done. Extension stopped, RoonServer restarted, extension
left off. Your two questions, separately:

1. Did it crash or restart on its own? No.

I pulled the server logs off the appliance directly. RoonServer_log.txt contains
exactly one “Starting RoonServer v2.70 (build 1671)” line, at 19:49:23 server
time — that’s my restart — and logs continuously through to 07:54:47 server time
when I pulled them. Twelve hours, no second startup line, no gap longer than 15
minutes anywhere in the file, zero “out of worker threads”, zero unhandled
exceptions, zero OOM. I also probed ports 9330/9331/9332 every two minutes
throughout: 211 consecutive healthy checks, zero failures.

One disclosure, because it affects how much you should trust that window.
The extension does appear in the log once, at 00:52:27 server time: six
registrations from 192.168.3.55 within a single second, then nothing. That was
me. While writing tests for the reconnect fix, my first test harness used a
module loader that bypassed the mocking layer, so it briefly connected to the
live Core instead of a stub. I caught it, verified the appliance was unaffected,
and hard-wired the tests to 127.0.0.1:9 so it can’t happen again. For scale
against what you diagnosed: that was ~4,700 registrations per minute; this was
six, in one second, once. No measurable effect in the logs. But you should know
the window isn’t perfectly clean rather than find it yourself.

2. Did music still cut out? Not during the window — but it did afterwards, and
we now know why. It isn’t AirPlay.

During the 12 hours: zero dropouts, zero underruns, zero errors in
RAATServer_log.txt. Only one playback session though, so that was a weak
negative at best.

Then we switched the U1 MINI to RoonReady as you suggested. Signal path is now
Lossless, 24/192 TIDAL FLAC straight through RAAT, and it sounds better — thank
you for that.

And a track still dropped. Korn, “It’s On!”, today at 15:07:25 server time.
That gave us the real answer, and it changes the picture:

15:05:18  poor connection kbps:5076.0 (min:5253.0)
15:05:28  poor connection kbps:4066.0 (min:5253.0)
... 14 consecutive warnings, all ~4,100 kbps ...
15:07:22  poor connection kbps:4074.0 (min:5253.0)
15:07:22  [prebuffer] sleeping in read -- this isn't good   (x14 in 3s)
15:07:25  Too many dropouts (>3s dropped out in the last 30s). Killing stream
15:07:25  advance didn't change the track. returning short read

So the source stream was being delivered ~22% below what the track needs, for two
solid minutes, until Roon gave up and skipped. Nothing to do with the endpoint or
the transport.

Two things make this more interesting than plain congestion:

First, the rate is pinned, not erratic — 4066, 4006, 4232, 4160, 4064, 4314,
4148, 4051, 4125, 4110, 4076, 4074. Congestion oscillates; that profile looks
like a ceiling.

Second, the same starvation hit a 16/44 stream the night before, on AirPlay:
poor connection kbps:1067.0 (min:1148.0). Eight times less demanding and still
short. So it isn’t “hi-res is too much for the line”.

The line itself is not the constraint — measured from the same network minutes
later: 92 Mb/s to one host, 40 Mb/s to another. But 2.3 Mb/s to a third, at the
same moment. Throughput here is strongly path-dependent, and the house firewall
is an EOL Sophos UTM 9 that stopped routing entirely two nights ago. We’re taking
that up with whoever administers it.

I’m not asking you to debug our network — but if Roon’s own poor connection
accounting is the right thing to trust here, that’s useful to know, and if
there’s a buffering setting that would ride out a 20% shortfall rather than
killing the stream, I’d take the pointer.

On the extension

I found the same loop you did, from the other end. Two independent reconnect
paths — core_unpaired and the websocket onclose — each armed a setTimeout without
tracking the pending one, so every disconnect roughly doubled the queued
reconnects. That matches the ~12,000 registrations you saw.

The fix is written and committed, held back until now so it couldn’t contaminate
this test. Five changes: a single tracked timer; a connect guard that reschedules
rather than returning empty-handed; exponential backoff capped at 5 minutes
instead of a fixed 10 seconds; a watchdog so the daemon can never sit neither
paired, nor connecting, nor scheduled; and a per-attempt epoch so a late callback
from an abandoned attempt can’t release the lock on one still in flight.

It’s covered by 21 tests, each verified by mutation — disabling any one mechanism
makes tests fail, so they aren’t decorative. Worth saying plainly: the repo’s 9
pre-existing tests never imported the daemon at all, which is exactly why a
defect this loud shipped with a green suite.

Deployed, and the U1 MINI is now on RoonReady — both done, one at a time as you
asked. Freeze: solved. Sound quality: improved, and we hadn’t even realised it
was degraded. Dropouts: understood, and pointing at our own network rather than
at Roon.

Thanks again for the precise diagnosis — the extension call was exactly right.

Hello @jc1

This is about as clean a piece of diagnostic work as we see come through here, thank you for running it exactly as asked and reporting back with this level of detail, including flagging your own test harness’s stray six-registration blip rather than letting it slide. That kind of transparency makes the result easy to trust.

On the extension: glad the fix is deployed and holding. That matches what we saw on our end.

On your buffering question: we checked, and there isn’t a setting that would help here. The 3-seconds-of-dropout-in-30-seconds threshold that ends playback is fixed, not adjustable, and the buffer settings Roon does expose (Resync Delay, hardware buffer size, WASAPI/ALSA buffer size) all apply to the local output device, not to the network stream coming in from TIDAL. So there’s no dial on our side that would ride out a sustained throughput shortfall like the one you measured.

Given the pinned-ceiling profile you found and the EOL Sophos UTM 9 dropping out entirely two nights ago, that really does point to the firewall/routing path rather than anything in Roon or on the endpoint side. Once that’s sorted with your network administrator, we’d expect the dropouts to clear along with it.

Marking this resolved on our end: RoonServer freeze fixed (extension re-registration loop), U1 MINI on RoonReady, and remaining dropouts traced to your network path. Please reopen if anything resurfaces.