ZimaOS crashing every 3 or so days

Ok. So I just replaced my ZimaOS SSD back, and forced upgrade to 1.6.2. Installed ZimaBrain. And… now what :slight_smile: I’ve been looking and clicking around, and I can ask it questions about disk health and such as per your how-to. But you say you want two reports, when I download brainsession I get a few-bytes file. So… What do you want me to ask it?

Sidenote, even when subscribed / tracking enabled on this thread, I don’t get any notification of it anymore. Not in spam either. Are there issues with that at the moment?

Thanks for testing it, and sorry, my earlier explanation was unclear.

For now, please use ZimaBrain to create a baseline before changing the kernel parameters. Ask these questions one at a time:

What needs attention?
Are there any failed services?
What kernel and boot parameters are active?
Are there any hardware, memory, storage or Intel GPU warnings?

After asking all four questions, please export the entire session and attach it to this thread.

Then, after making Jerry’s cmdline.txt changes, ask the same four questions again and export the full session once more. That will give us a clean before-and-after comparison.

The small brainsession file you downloaded earlier does not sound correct, so please also mention its file size when you attach it.

Please don’t apologise :slight_smile: Attached two logs, one with the default cmdline.txt and one for the modified one where I removed intel_iommu=on and vfio_iommu_type1.allow_unsafe_interrupts=1 completely. Output is almost identical except for a systemd-networkd-wait-online.service error.

If I should have exported another file tell me.

zimabrain-session_modified_cmdline.txt_redacted-20260713-112354.md.txt (5.8 KB)

zimabrain-session_default_cmdline.txt_redacted-20260713-112354.md.txt (6.1 KB)

Thank you, these are the correct files and you do not need to export anything else at this stage.

The reports are mostly identical, as expected, but there is one confirmed difference:

With the default cmdline.txt, ZimaBrain detected:

systemd-networkd-wait-online.service

as failed.

After removing:

intel_iommu=on
vfio_iommu_type1.allow_unsafe_interrupts=1

that failed-service finding was no longer present.

This does not yet prove those parameters caused the system crashes, but it gives us a clean baseline and confirms that the modified boot configuration changed at least one part of the system state.

Please continue running with the modified configuration and let us know how long it remains stable. If it crashes again, reboot, ask ZimaBrain the same questions and export the complete session again so we can compare it with these two reports.

Your exports also exposed a few areas where ZimaBrain’s question routing needs improvement, particularly the failed-services and active kernel-parameter questions. We will include those improvements in the next release. Thank you for helping us test it.

Just a heads-up, 3 days, 20 hours so far, without issues. That’s nothing spectacular, I have had uptimes up to 6 days. But usually a few minutes after posting all is well it crashes :wink:

I’ve removed

intel_iommu=on
vfio_iommu_type1.allow_unsafe_interrupts=1

from the config.

1 Like

That is a very useful update, thank you.

3 days and 20 hours without issues is not final proof yet, but it is a good sign, especially compared with the previous behaviour.

The important point is that the system has been running with these removed:

intel_iommu=on
vfio_iommu_type1.allow_unsafe_interrupts=1

So at the moment it looks like one of those kernel parameters may have been involved, or at least made the system less stable on this hardware / VM combination.

I would keep it running exactly as it is for now and avoid changing anything else, so the test stays clean.

If it reaches more than your previous 6 day uptime, that becomes much stronger evidence that removing those parameters helped.

If it crashes again, please export another ZimaBrain report after reboot, and mention clearly that this was with both IOMMU/VFIO parameters removed.

Thanks again for keeping track of the uptime. This is the kind of slow testing that actually helps narrow down the cause.

@Rataplan626, the new ZimaBrain CE Docker version is now available. This update adds persistent local memory, system-health history, memory/swap checks, failed-service analysis, SMART/NVMe trends, container history and ZimaOS update comparison.

This should be much more useful for your recurring hard-lock issue because you can create a baseline, then after the next crash and reboot ask the same or reworded questions and ZimaBrain will compare the new evidence with the previous scan.

Please test:

Give me a complete system-health assessment.
What changed since my previous scan?
What changed after my latest ZimaOS update?
Which services are failed or degraded?
Is memory pressure high?

Then export the redacted session. A complete hard lock may still leave no final journal entry, but the new memory layer should help identify changes or worsening signals leading up to repeated crashes.

Quote: If it crashes again, please export another ZimaBrain report after reboot, and mention clearly that this was with both IOMMU/VFIO parameters removed.

It just crashed again with both options removed. I’ve not yet rebooted it. Should I ask the same questions to ZimaBrain? Or should I mention to zimabrain I have those options disabled?

Best give exact instructions that aid you the most :slightly_smiling_face:

I changed Zimabrain tag to ‘latest’, let it update and now it doesn’t start anymore..

For the crashes, should I update to 1.7.0 beta? Or keep 1.6.2 for now.

Stay on ZimaOS 1.6.2 for now.

To get ZimaBrain CE working again, change the image tag from latest to the 1.6.0-beta tag.

Here’s zimabrains report, with Gelbuildings’ questions.

Side question, I now asked it all questions one by one. Could I have copied the whole set of questions all at once?

zimabrain-session-redacted-20260720-162859.md.txt (19.2 KB)

Thank you, this report is useful and confirms an important point: the system crashed again with both IOMMU/VFIO parameters removed, so removing them did not resolve the hard lock.

The report after reboot shows no failed services, normal memory pressure, a normal captured temperature of 47°C, no SMART/NVMe failure markers, and all five containers running.

One item worth watching is zimaos-local-st, which was doing approximately 1.25 MB/s of disk activity across all three recorded scans. That proves persistent background activity, but it does not yet prove it caused the crash.

I also found a ZimaBrain routing gap. Your question about active kernel and boot parameters reached the install/boot guidance layer and did not display the actual kernel command line. Therefore, the report did not independently verify that the two parameters were absent. I will correct that.

For now, please run these two commands after the reboot:

cat /proc/cmdline
journalctl -b -1 -k --no-pager | grep -Ei 'i915|drm|gpu|hang|reset|iommu|vfio|nvme|pcie|aer|watchdog|lockup|oom|thermal|fatal|error' | tail -300 > /DATA/previous-boot-kernel-filtered.txt

Please paste the first output and attach /DATA/previous-boot-kernel-filtered.txt.

Regarding the questions, please continue asking them one at a time. Pasting the entire set together would currently be treated as one combined question and could route to the wrong diagnostic layer. A proper batch-question option would need to be added separately.

Thanks for pushing through!

root@RataNAS:/root ➜ # cat /proc/cmdline
BOOT_IMAGE=(hd1,gpt2)/bzImage root=PARTUUID=8d3d53e3-6d49-4c38-8349-aff6859e82fd rootwait net.naming-scheme=v250 systemd.machine_id=4073bb4f60c14bbeb7016546f3b3098c fsck.repair=yes console=tty1 quiet splash loglevel=3 systemd.show_status=1 rd.udev.log_level=3 net.ifnames=0 biosdevname=0 thunderbolt.host_reset=false rauc.slot=A

previous-boot-kernel-filtered.txt (4.6 KB)

Thanks, this confirms the test was valid. Your /proc/cmdline shows that both intel_iommu=on and vfio_iommu_type1.allow_unsafe_interrupts=1 were removed, but the system still crashed.

The DMAR/IOMMU and VFIO messages are normal kernel subsystem initialization. They do not mean those removed parameters are still active.

The filtered log does not contain a confirmed cause. I cannot see a kernel panic, OOM event, thermal shutdown, NVMe error, GPU hang/reset, watchdog lockup or PCIe/AER failure. The Intel graphics driver initializes successfully, and Cannot find any crtc or sizes usually means no active display was detected.

The log only contains startup messages and does not reach the time of the freeze. Please run these two commands:

journalctl -b -1 -o short-precise --no-pager | tail -300 > /DATA/crashed-boot-last-300.txt
ls -la /sys/fs/pstore /var/log/journal 2>&1

The first command creates /DATA/crashed-boot-last-300.txt from the boot session that crashed immediately before the current boot. Please review it for personal information and attach it here.

The second command prints its result directly in the terminal. Please copy and paste that output into your reply. This will show whether anything was recorded immediately before the freeze and whether persistent crash logging is available.

root@RataNAS:/root ➜ # ls -la /sys/fs/pstore /var/log/journal 2>&1
/sys/fs/pstore:
total 0
drwxr-x— 2 root root 0 Jul 20 17:23 .
drwxr-xr-x 10 root root 0 Jul 20 17:23 ..

/var/log/journal:
total 20
drwxr-sr-x+ 3 root systemd-journal 4096 Apr 20 16:35 .
drwxr-xr-x 8 root root 4096 Apr 20 16:35 ..
drwxr-sr-x+ 2 root systemd-journal 4096 Jul 20 17:23 4073bb4f60c14bbeb7016546f3b3098c

crashed-boot-last-300.txt (71.7 KB)

Thank you. This gives us a much clearer result.

Persistent journal storage is enabled, so ZimaOS was saving logs across reboots. However, /sys/fs/pstore is empty, which means the kernel did not preserve a panic or crash record.

The crashed boot journal ends abruptly at 14:06:08 with an ordinary Samba network message. There is no normal shutdown, reboot sequence, kernel panic, OOM event, GPU reset, NVMe error, thermal event or watchdog lockup recorded before it stops.

I found several unrelated warnings:

  • Samba/AppArmor denied access to lock /etc/samba/smbpasswd.
  • A system hourly cron job exited with status 127.
  • One App Store compose file failed validation.
  • ZimaOS downloaded the 1.7.0-beta1 update bundle, but the log does not show it being installed.

None of these currently provides evidence of the complete system freeze.

Do you remember approximately what time the system became unresponsive? The final recorded message was at 14:06:08. If the freeze happened around that time, it confirms the machine locked so completely that it could not write the actual failure to disk.

Since the crash also occurred without the two IOMMU/VFIO parameters, we can now rule out removing those parameters as the solution. If the time matches, the next useful step will be remote kernel logging so another device can capture the final messages during the next lockup.

I suggest that, if the freeze occurred around 14:06, we configure Linux netconsole for the next test. This would allow your RataNAS to send live kernel messages across your local network to another computer, which may capture the final error when RataNAS locks before it can save anything locally. Do you have another Linux or ZimaOS device that can remain running, and what is its LAN IP address?

My monitoring system (zabbix, but it’s remote. My firewall runs a zabbix proxy) also reports 14.06.

I can fire up a vm with any distro that suits the job, or if it’s something basic that runs from freebsd I can use my firewall maybe. I’ll substitute ip’s myself🙂

Perfect. A small Debian or Ubuntu VM will be the cleanest option. I suggest keeping the firewall out of this test so we do not change anything on a critical network device.

Configure the VM with a bridged network adapter on the same LAN as RataNAS and give it a fixed or reserved IP address. Linux netconsole is designed specifically to send kernel messages over UDP when local disk logging fails.

[Linux kernel netconsole documentation]

On the VM, run:

sudo apt-get update && sudo apt-get install -y socat

Then start the receiver:

sudo sh -c 'socat -u UDP-RECV:6666,reuseaddr - | tee -a /var/log/ratanas-netconsole.log'

The second command will remain open with a blank terminal while it waits for messages. Leave it running and tell me the VM’s LAN IP address. We will then configure RataNAS temporarily and send a harmless test message before making anything persistent.

That will be Debian based then, hostname ZimaLogger, ip 192.168.6.252

[edit] just as update, it crashed again, with the 2 kernel options still removed (I haven’t changed anything) so that’s within a day. It’s unpredictable.

Perfect. Debian VM ZimaLogger at 192.168.6.252 on the same subnet is ideal.

The second crash, with both kernel options still removed, confirms that removing them did not solve the problem. The shorter interval also confirms how unpredictable the lockup is, which makes continuous remote kernel logging the correct next step.

Please keep RataNAS unchanged now. Once ZimaLogger is online and the socat receiver is running, run these two commands on RataNAS:

ip -br -4 addr
ping -c 1 192.168.6.252 >/dev/null && ip neigh show 192.168.6.252

Please paste both results. They will identify the correct RataNAS network interface and ZimaLogger MAC address. I will then provide the temporary netconsole command and a harmless test message. We will only make it persistent after confirming that ZimaLogger receives the test.