RTX PRO 4000 Blackwell Xid 8 / CUDA_ERROR_LAUNCH_TIMEOUT in 2D Classification and Ab-initio

Hi,

I am currently having a problem with CryoSPARC v5.0.6 on a workstation with two NVIDIA RTX PRO 4000 Blackwell GPUs.

The problem mainly happens during 2D Classification and Ab-initio Reconstruction. Other jobs, especially Non-uniform Refinement, seem to run without problems.

The CryoSPARC jobs fail with errors like:

CUDA_ERROR_LAUNCH_TIMEOUT

For example:

RuntimeError: cuStreamSynchronize(...):
CUDA ERROR: (CUDA_ERROR_LAUNCH_TIMEOUT)
the launch timed out and was terminated

I have also seen:

numba.cuda.cudadrv.driver.CudaAPIError:
[702] Call to cuMemcpyHtoDAsync results in CUDA_ERROR_LAUNCH_TIMEOUT

and:

[702] Call to cuMemGetInfo results in CUDA_ERROR_LAUNCH_TIMEOUT

At the same time, the Linux kernel reports an NVIDIA Xid 8 error:

NVRM: krcWatchdog_IMPL: RC watchdog: GPU is probably locked! Notify Timeout Seconds: 7
NVRM: GPU at PCI:0000:01:00
NVRM: Xid (PCI:0000:01:00): 8, pid=5937, name=python, channel 0x0000000e

What makes this a bit strange is that it happens on both GPUs, not only one card.

From the kernel logs I found several Xid 8 events:

GPU 0 / PCI 01:00.0:
2026-08-04  Xid 8
2026-08-04  Xid 8
2026-08-06  Xid 8

GPU 1 / PCI 61:00.0:
2026-08-10  Xid 8
2026-08-16  Xid 8

Our system:

Ubuntu 24.04.4 LTS
Kernel: 6.17.0-1032-oem
CryoSPARC: v5.0.6

2 × NVIDIA RTX PRO 4000 Blackwell
24 GB VRAM each

595.84 open


CRYOSPARC_NO_PAGELOCK=true
THP: always [madvise] never

I first thought this could be a driver problem, so I updated from NVIDIA 580.173.02 to 595.84.

Unfortunately, the same problem still happens:

580.173.02 -> Xid 8 / CUDA_ERROR_LAUNCH_TIMEOUT
595.84     -> Xid 8 / CUDA_ERROR_LAUNCH_TIMEOUT

I also repeated one of the problematic 2D Classification jobs after the driver update.

The job was running only on GPU 0:

GPU=[0]
MAIN PROCESS PID 5937

Just before the crash, the CryoSPARC log showed:

DIE: cuMemsetD32Async(...):
CUDA ERROR: (CUDA_ERROR_LAUNCH_TIMEOUT)
the launch timed out and was terminated

Received SIGSEGV

At exactly the same time, the kernel reported:

NVRM: RC watchdog: GPU is probably locked!
NVRM: Xid (PCI:0000:01:00): 8, pid=5937, name=python

So the Xid is coming from the same CryoSPARC Python process.

I also checked that the CRYOSPARC_NO_PAGELOCK=true setting is actually being used. The job log contains:

HOST ALLOCATION FUNCTION: using n.empty
(CRYOSPARC_NO_PAGELOCK==true)

I do not see obvious signs of a hardware problem. After the crash there were no PCIe replay errors, no thermal throttling and no memory remapping errors. The GPU was also not close to running out of VRAM. We are also running other GPU heavy applications on that workstation without any issues.

I would be happy if anyone has an idea how to fix it.

I can also provide the full CryoSPARC job log and NVIDIA bug report if that would help.

Thanks!

@OleUns Please can you post the output of the nvidia-smi command

@wtempel yes:

$ nvidia-smi
Mon Aug 17 21:17:53 2026       
+-----------------------------------------------------------------------------------------+
| NVIDIA-SMI 595.84                 Driver Version: 595.84         CUDA Version: 13.2     |
+-----------------------------------------+------------------------+----------------------+
| GPU  Name                 Persistence-M | Bus-Id          Disp.A | Volatile Uncorr. ECC |
| Fan  Temp   Perf          Pwr:Usage/Cap |           Memory-Usage | GPU-Util  Compute M. |
|                                         |                        |               MIG M. |
|=========================================+========================+======================|
|   0  NVIDIA RTX PRO 4000 Blac...    Off |   00000000:01:00.0  On |                  Off |
| 30%   41C    P1             29W /  145W |     524MiB /  24467MiB |      0%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+
|   1  NVIDIA RTX PRO 4000 Blac...    Off |   00000000:61:00.0 Off |                  Off |
| 30%   28C    P8              2W /  145W |      18MiB /  24467MiB |      0%      Default |
|                                         |                        |                  N/A |
+-----------------------------------------+------------------------+----------------------+

+-----------------------------------------------------------------------------------------+
| Processes:                                                                              |
|  GPU   GI   CI              PID   Type   Process name                        GPU Memory |
|        ID   ID                                                               Usage      |
|=========================================================================================|
|    0   N/A  N/A            2332      G   /usr/lib/xorg/Xorg                      107MiB |
|    0   N/A  N/A            2499    C+G   ...c/gnome-remote-desktop-daemon        225MiB |
|    0   N/A  N/A            2564      G   /usr/bin/gnome-shell                    120MiB |
|    0   N/A  N/A           80913      G   /usr/bin/nautilus                        14MiB |
|    1   N/A  N/A            2332      G   /usr/lib/xorg/Xorg                        4MiB |
+-----------------------------------------------------------------------------------------+

I had a lot of problems with driver 595. RTX PRO 6000s having the same issue (but in patch motion). Reverting to driver 580 solved all of the CUDA TIMEOUT errors I saw. Seeing it manifest in 580 is new. Have you tried 610? But some other teething problems seem to be happening with the 610 drivers…

@rbs_sci I could give 610 a shot and check if it solves it. The crashed are fairly reproducible, usually the jobs starts and crashes after 0.5-4h. Its also quite interesting that in your case 580 fixed it and that it only happened in motion correction. Motion correction runs fine on our setup.

Yes, I’ve been seeing more and more odd behaviour with nVidia drivers recently. We have four boxes in the office which are hardware identical, but at one point ran different driver versions, and saw behaviour with one driver which we could not reproduce on other systems running a different driver.

Takes me back to the early days of CUDA. :rofl:

Unfortunately, updating to 610 did not work. I still got the CUDA_ERROR_LAUNCH_TIMEOUT error. It’s a bit frustrating with the Blackwell GPUs. Ironically, the most reliable GPUs in our setup are a pair of RTX2080Ti’s, didn’t had issues for a very long time.

Fairly similar story for us - our most reliable GPUs are a pair of old 1080Tis, although the Ampere generation Quadros we have, have also been absolute troopers and extremely reliable.

Our four Blackwell cards have been a string of headaches.

2 Likes

@OleUns @rbs_sci Thanks for posting the issue and contributing to the discussion. We unfortunately do not know the cause of the problem, please update the forum with any new findings, for example related to your experiments with different driver versions.

Just in case, does setting

export CRYOSPARC_NO_PAGELOCK=false

inside cryosparc_worker/config.sh make a difference?

@wtempel I tested export CRYOSPARC_NO_PAGELOCK=false and it did not make any difference with 597.84-open.

If I find a solution I’ll update you in the forum.