Child process terminated unexpectedly with exit code -11 in patch motion correction

Hello cryoSPARC team,

For some time we have been having this error during patch motion correction:
Child process with PID XXXXXX terminated unexpectedly with exit code -11.

We have tried rolling back to CS 4.4, and out IT person troubleshooted alot of other things. This error happens in both of our worksations with different GPUs and on different datasets. The crashes always happen at a random movie, the GPUs crash sequentially (the remaining ones can run for several minutes after the first crashed) and sometimes it will run for 5 minutes, sometimes for 15 minutes.

Surprisingly, our cluster setup is fine and doesn’t show this error. I was hoping you could shed some light on this issue.
You can find below one of the logs:

/mnt/tesla/data/cryosparc/4.5.3/worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/cudadrv/nvrtc.py:257: UserWarning: NVRTC log messages whilst compiling kernel:

kernel(18): warning #177-D: variable "sd" was declared but never referenced

kernel(18): warning #177-D: variable "o" was declared but never referenced


  warnings.warn(msg)
/mnt/tesla/data/cryosparc/4.5.3/worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/dispatcher.py:536: NumbaPerformanceWarning: Grid size 12 will likely result in GPU under-utilization due to low occupancy.
  warn(NumbaPerformanceWarning(msg))
/mnt/tesla/data/cryosparc/4.5.3/worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/dispatcher.py:536: NumbaPerformanceWarning: Grid size 12 will likely result in GPU under-utilization due to low occupancy.
  warn(NumbaPerformanceWarning(msg))
gpufft: creating new cufft plan (plan id 4   pid 441338) 
	gpu_id  0 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 2917 0 
	istride 1 
	idist   17011944 
	onembed 5832 5834 0 
	ostride 1 
	odist   34023888 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 4   pid 441339) 
	gpu_id  1 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 2917 0 
	istride 1 
	idist   17011944 
	onembed 5832 5834 0 
	ostride 1 
	odist   34023888 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

/mnt/tesla/data/cryosparc/4.5.3/worker/cryosparc_compute/jobs/pipeline.py:59: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
  return self.process(item)
========= sending heartbeat at 2024-06-18 13:18:33.381849
/mnt/tesla/data/cryosparc/4.5.3/worker/cryosparc_compute/jobs/pipeline.py:59: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
  return self.process(item)
========= sending heartbeat at 2024-06-18 13:18:43.399320
========= sending heartbeat at 2024-06-18 13:18:53.419293
========= sending heartbeat at 2024-06-18 13:19:03.439589
<string>:1: RuntimeWarning: More than 20 figures have been opened. Figures created through the pyplot interface (`matplotlib.pyplot.figure`) are retained until explicitly closed and may consume too much memory. (To control this warning, see the rcParam `figure.max_open_warning`). Consider using `matplotlib.pyplot.close()`.
========= sending heartbeat at 2024-06-18 13:19:13.459194
========= sending heartbeat at 2024-06-18 13:19:23.479220
========= sending heartbeat at 2024-06-18 13:19:33.500499
========= sending heartbeat at 2024-06-18 13:19:43.521875
========= sending heartbeat at 2024-06-18 13:19:53.534256
========= sending heartbeat at 2024-06-18 13:20:03.551332
========= sending heartbeat at 2024-06-18 13:20:13.571479
========= sending heartbeat at 2024-06-18 13:20:23.592214
========= sending heartbeat at 2024-06-18 13:20:33.612311
========= sending heartbeat at 2024-06-18 13:20:43.633280
========= sending heartbeat at 2024-06-18 13:20:53.653677
========= sending heartbeat at 2024-06-18 13:21:03.673853
========= sending heartbeat at 2024-06-18 13:21:13.694523
========= sending heartbeat at 2024-06-18 13:21:23.707346
========= sending heartbeat at 2024-06-18 13:21:33.719378
========= sending heartbeat at 2024-06-18 13:21:43.739718
========= sending heartbeat at 2024-06-18 13:21:53.759719
========= sending heartbeat at 2024-06-18 13:22:03.779323
========= sending heartbeat at 2024-06-18 13:22:13.799345
========= sending heartbeat at 2024-06-18 13:22:23.818677
========= sending heartbeat at 2024-06-18 13:22:33.839321
========= sending heartbeat at 2024-06-18 13:22:43.859212
========= sending heartbeat at 2024-06-18 13:22:53.879008
========= sending heartbeat at 2024-06-18 13:23:03.900074
========= sending heartbeat at 2024-06-18 13:23:13.920176
========= sending heartbeat at 2024-06-18 13:23:23.939329
  ========= heartbeat failed at 2024-06-18 13:23:23.969495: 
========= sending heartbeat at 2024-06-18 13:23:33.979649
  ========= heartbeat failed at 2024-06-18 13:23:33.989203: 
========= sending heartbeat at 2024-06-18 13:23:43.999350
  ========= heartbeat failed at 2024-06-18 13:23:44.011054: 
 ************* Connection to cryosparc command lost. Heartbeat failed 3 consecutive times at 2024-06-18 13:23:44.011102.
/mnt/tesla/data/cryosparc/4.5.3/worker/bin/cryosparcw: line 150: 441301 Killed                  python -c "import cryosparc_compute.run as run; run.run()" "$@"

Let me know if you require any more information.

Thank you very much and best regards.

Welcome to the forum @Arpind.
Please can you describe the setup on which this error occurred:

  • is the GPU computer separate from the CryoSPARC master computer?
  • what are the outputs of these commands on the CryoSPARC master computer
    cryosparcm status | grep HOST
    free -h
    cat /sys/kernel/mm/transparent_hugepage/enabled
    

Thank you for your reply @wtempel .

This is a worksation setup so the GPUs and CryoSPARC master are all in the same computer.

Here are the outputs of the commands.

cryosparc@tesla:~$ ~/4.5.3/master/bin/cryosparcm status | grep HOST
export CRYOSPARC_MASTER_HOSTNAME="tesla.campus.mcgill.ca"

cryosparc@tesla:~$ free -h
                     total        used        free      shared  buff/cache   available
Mem:           503Gi       5.6Gi       185Gi       7.0Mi       312Gi       494Gi
Swap:          6.0Gi       229Mi       5.8Gi

cryosparc@tesla:~$ cat /sys/kernel/mm/transparent_hugepage/enabled
always [madvise] never

cryosparc@tesla:~$ 

Thanks for the information @Arpind.
What are the outputs of these commands (run on tesla):

host tesla.campus.mcgill.ca
cryosparcm status | grep PORT
ps -eo pid,ppid,start,rsz,vsz,cmd | grep -e cryosparc_ -e mongo

No problem @wtempel .

Here are the outputs :

cryosparc@tesla:~$ host tesla.campus.mcgill.ca
tesla.campus.mcgill.ca has address 132.206.28.252

cryosparc@tesla:~$ cryosparcm status | grep PORT
export CRYOSPARC_BASE_PORT=61000

cryosparc@tesla:~$ ps -eo pid,ppid,start,rsz,vsz,cmd | grep -e cryosparc_ -e mongo
 421461       1   Jun 18 20840  41124 python /mnt/tesla/data/cryosparc/4.5.3/master/deps/anaconda/envs/cryosparc_master_env/bin/supervisord -c /mnt/tesla/data/cryosparc/4.5.3/master/supervisord.conf
 421571  421461   Jun 18 829316 2308820 mongod --auth --dbpath /mnt/tesla/data/cryosparc/3.3.0/cryosparc_database --port 61001 --oplogSize 64 --replSet meteor --wiredTigerCacheSizeGB 4 --bind_ip_all
 421678  421461   Jun 18 92664 149280 python /mnt/tesla/data/cryosparc/4.5.3/master/deps/anaconda/envs/cryosparc_master_env/bin/gunicorn -n command_core -b 0.0.0.0:61002 cryosparc_command.command_core:start() -c /mnt/tesla/data/cryosparc/4.5.3/master/gunicorn.conf.py
 421679  421678   Jun 18 119700 915664 python /mnt/tesla/data/cryosparc/4.5.3/master/deps/anaconda/envs/cryosparc_master_env/bin/gunicorn -n command_core -b 0.0.0.0:61002 cryosparc_command.command_core:start() -c /mnt/tesla/data/cryosparc/4.5.3/master/gunicorn.conf.py
 422222  421461   Jun 18 92464 149380 python /mnt/tesla/data/cryosparc/4.5.3/master/deps/anaconda/envs/cryosparc_master_env/bin/gunicorn cryosparc_command.command_vis:app -n command_vis -b 0.0.0.0:61003 -c /mnt/tesla/data/cryosparc/4.5.3/master/gunicorn.conf.py
 422233  422222   Jun 18 271032 1304232 python /mnt/tesla/data/cryosparc/4.5.3/master/deps/anaconda/envs/cryosparc_master_env/bin/gunicorn cryosparc_command.command_vis:app -n command_vis -b 0.0.0.0:61003 -c /mnt/tesla/data/cryosparc/4.5.3/master/gunicorn.conf.py
 422246  421461   Jun 18 92728 149380 python /mnt/tesla/data/cryosparc/4.5.3/master/deps/anaconda/envs/cryosparc_master_env/bin/gunicorn cryosparc_command.command_rtp:start() -n command_rtp -b 0.0.0.0:61005 -c /mnt/tesla/data/cryosparc/4.5.3/master/gunicorn.conf.py
 422247  422246   Jun 18 227392 1028904 python /mnt/tesla/data/cryosparc/4.5.3/master/deps/anaconda/envs/cryosparc_master_env/bin/gunicorn cryosparc_command.command_rtp:start() -n command_rtp -b 0.0.0.0:61005 -c /mnt/tesla/data/cryosparc/4.5.3/master/gunicorn.conf.py
 422280  421461   Jun 18 144820 1152764 /mnt/tesla/data/cryosparc/4.5.3/master/cryosparc_app/nodejs/bin/node ./bundle/main.js
2137526 2057350 13:45:06  2360   9076 grep --color=auto -e cryosparc_ -e mongo

cryosparc@tesla:~$ 

Thank you for the support.

@Arpind Please can you identify the command_core-related log files that contain logs from the time window when the heartbeat failed by running the following command on tesla:

grep -l "^2024-06-18 1" /mnt/tesla/data/cryosparc/4.5.3/master/run/command_core*

then make copies of identified log file(s) and send us the compressed copy or copies as an email attachment. I will send you a private message with our email address.

Hi @wtempel , was just wondering if you remember whether there was a solution to this? We are having a pretty much identical issue on a fresh Ubuntu 22.04 / cryoSPARC 4.7.1 re-installation. Motion Correction jobs are failing following very similar warnings as above. The job failures come at seemingly random time points and each time cause the whole server to reboot. Maybe it’s a GPU memory issue / we haven’t set up the Nvidia drivers correctly?

Many thanks!

Louis

e.g.

================= CRYOSPARCW =======  2026-10-06 14:49:52.210872  =========
Project P8 Job J2
Master jzserver2.med.virginia.edu Port 39002
===========================================================================
MAIN PROCESS PID 12276
========= now starting main process at 2026-10-06 14:49:52.211911
motioncorrection.run_patch cryosparc_compute.jobs.jobregister
MONITOR PROCESS PID 12278
========= monitor process now waiting for main process
========= sending heartbeat at 2026-10-06 14:49:53.668895
***************************************************************
Transparent hugepages setting: always [madvise] never

Running job on hostname %s jzserver2.med.virginia.edu
Allocated Resources :  {'fixed': {'SSD': False}, 'hostname': 'jzserver2.med.virginia.edu', 'lane': 'default', 'lane_type': 'node', 'license': True, 'licenses_acquired': 4, 'slots': {'CPU': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23], 'GPU': [0, 1, 2, 3], 'RAM': [0, 1, 2, 3, 4, 5, 6, 7]}, 'target': {'cache_path': '/mnt/csparc_scratch', 'cache_quota_mb': None, 'cache_reserve_mb': 10000, 'desc': None, 'gpus': [{'id': 0, 'mem': 11346575360, 'name': 'NVIDIA GeForce RTX 2080 Ti'}, {'id': 1, 'mem': 11346575360, 'name': 'NVIDIA GeForce RTX 2080 Ti'}, {'id': 2, 'mem': 11346575360, 'name': 'NVIDIA GeForce RTX 2080 Ti'}, {'id': 3, 'mem': 11346575360, 'name': 'NVIDIA GeForce RTX 2080 Ti'}], 'hostname': 'jzserver2.med.virginia.edu', 'lane': 'default', 'monitor_port': None, 'name': 'jzserver2.med.virginia.edu', 'resource_fixed': {'SSD': True}, 'resource_slots': {'CPU': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63], 'GPU': [0, 1, 2, 3], 'RAM': [0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47]}, 'ssh_str': 'lab@jzserver2.med.virginia.edu', 'title': 'Worker node jzserver2.med.virginia.edu', 'type': 'node', 'worker_bin_path': '/srv2/home/lab/cryosparc/cryosparc_worker/bin/cryosparcw'}}
========= sending heartbeat at 2026-10-06 14:50:03.686498
gpufft: creating new cufft plan (plan id 0   pid 12319) 
	gpu_id  3 
	ndims   2 
	dims    768 768 0 
	inembed 768 770 0 
	istride 1 
	idist   591360 
	onembed 768 385 0 
	ostride 1 
	odist   295680 
	batch   70 
	type    R2C 
	wkspc   automatic 
	Python traceback:

gpufft: creating new cufft plan (plan id 0   pid 12317) 
	gpu_id  1 
	ndims   2 
	dims    768 768 0 
	inembed 768 770 0 
	istride 1 
	idist   591360 
	onembed 768 385 0 
	ostride 1 
	odist   295680 
	batch   70 
	type    R2C 
	wkspc   automatic 
	Python traceback:

gpufft: creating new cufft plan (plan id 0   pid 12316) 
	gpu_id  0 
	ndims   2 
	dims    768 768 0 
	inembed 768 770 0 
	istride 1 
	idist   591360 
	onembed 768 385 0 
	ostride 1 
	odist   295680 
	batch   70 
	type    R2C 
	wkspc   automatic 
	Python traceback:

gpufft: creating new cufft plan (plan id 0   pid 12318) 
	gpu_id  2 
	ndims   2 
	dims    768 768 0 
	inembed 768 770 0 
	istride 1 
	idist   591360 
	onembed 768 385 0 
	ostride 1 
	odist   295680 
	batch   70 
	type    R2C 
	wkspc   automatic 
	Python traceback:

gpufft: creating new cufft plan (plan id 1   pid 12319) 
	gpu_id  3 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 5834 0 
	istride 1 
	idist   34023888 
	onembed 5832 2917 0 
	ostride 1 
	odist   17011944 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 2   pid 12319) 
	gpu_id  3 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 5833 0 
	istride 1 
	idist   68036112 
	onembed 11664 11666 0 
	ostride 1 
	odist   136072224 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 3   pid 12319) 
	gpu_id  3 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 11666 0 
	istride 1 
	idist   136072224 
	onembed 11664 5833 0 
	ostride 1 
	odist   68036112 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/cudadrv/nvrtc.py:257: UserWarning: NVRTC log messages whilst compiling kernel:

kernel(18): warning #177-D: variable "sd" was declared but never referenced

kernel(18): warning #177-D: variable "o" was declared but never referenced


  warnings.warn(msg)
gpufft: creating new cufft plan (plan id 1   pid 12317) 
	gpu_id  1 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 5834 0 
	istride 1 
	idist   34023888 
	onembed 5832 2917 0 
	ostride 1 
	odist   17011944 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 2   pid 12317) 
	gpu_id  1 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 5833 0 
	istride 1 
	idist   68036112 
	onembed 11664 11666 0 
	ostride 1 
	odist   136072224 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 3   pid 12317) 
	gpu_id  1 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 11666 0 
	istride 1 
	idist   136072224 
	onembed 11664 5833 0 
	ostride 1 
	odist   68036112 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/cudadrv/nvrtc.py:257: UserWarning: NVRTC log messages whilst compiling kernel:

kernel(18): warning #177-D: variable "sd" was declared but never referenced

kernel(18): warning #177-D: variable "o" was declared but never referenced


  warnings.warn(msg)
/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/dispatcher.py:536: NumbaPerformanceWarning: Grid size 12 will likely result in GPU under-utilization due to low occupancy.
  warn(NumbaPerformanceWarning(msg))
/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/dispatcher.py:536: NumbaPerformanceWarning: Grid size 12 will likely result in GPU under-utilization due to low occupancy.
  warn(NumbaPerformanceWarning(msg))
gpufft: creating new cufft plan (plan id 1   pid 12318) 
	gpu_id  2 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 5834 0 
	istride 1 
	idist   34023888 
	onembed 5832 2917 0 
	ostride 1 
	odist   17011944 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 2   pid 12318) 
	gpu_id  2 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 5833 0 
	istride 1 
	idist   68036112 
	onembed 11664 11666 0 
	ostride 1 
	odist   136072224 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 3   pid 12318) 
	gpu_id  2 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 11666 0 
	istride 1 
	idist   136072224 
	onembed 11664 5833 0 
	ostride 1 
	odist   68036112 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/cudadrv/nvrtc.py:257: UserWarning: NVRTC log messages whilst compiling kernel:

kernel(18): warning #177-D: variable "sd" was declared but never referenced

kernel(18): warning #177-D: variable "o" was declared but never referenced


  warnings.warn(msg)
gpufft: creating new cufft plan (plan id 1   pid 12316) 
	gpu_id  0 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 5834 0 
	istride 1 
	idist   34023888 
	onembed 5832 2917 0 
	ostride 1 
	odist   17011944 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 2   pid 12316) 
	gpu_id  0 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 5833 0 
	istride 1 
	idist   68036112 
	onembed 11664 11666 0 
	ostride 1 
	odist   136072224 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 3   pid 12316) 
	gpu_id  0 
	ndims   2 
	dims    11664 11664 0 
	inembed 11664 11666 0 
	istride 1 
	idist   136072224 
	onembed 11664 5833 0 
	ostride 1 
	odist   68036112 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/cudadrv/nvrtc.py:257: UserWarning: NVRTC log messages whilst compiling kernel:

kernel(18): warning #177-D: variable "sd" was declared but never referenced

kernel(18): warning #177-D: variable "o" was declared but never referenced


  warnings.warn(msg)
/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/dispatcher.py:536: NumbaPerformanceWarning: Grid size 12 will likely result in GPU under-utilization due to low occupancy.
  warn(NumbaPerformanceWarning(msg))
/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/dispatcher.py:536: NumbaPerformanceWarning: Grid size 12 will likely result in GPU under-utilization due to low occupancy.
  warn(NumbaPerformanceWarning(msg))
========= sending heartbeat at 2026-10-06 14:50:13.695012
gpufft: creating new cufft plan (plan id 4   pid 12317) 
	gpu_id  1 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 2917 0 
	istride 1 
	idist   17011944 
	onembed 5832 5834 0 
	ostride 1 
	odist   34023888 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 4   pid 12319) 
	gpu_id  3 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 2917 0 
	istride 1 
	idist   17011944 
	onembed 5832 5834 0 
	ostride 1 
	odist   34023888 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 4   pid 12318) 
	gpu_id  2 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 2917 0 
	istride 1 
	idist   17011944 
	onembed 5832 5834 0 
	ostride 1 
	odist   34023888 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

gpufft: creating new cufft plan (plan id 4   pid 12316) 
	gpu_id  0 
	ndims   2 
	dims    5832 5832 0 
	inembed 5832 2917 0 
	istride 1 
	idist   17011944 
	onembed 5832 5834 0 
	ostride 1 
	odist   34023888 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/jobs/pipeline.py:59: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
  return self.process(item)
/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/jobs/pipeline.py:59: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
  return self.process(item)
/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/jobs/pipeline.py:59: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
  return self.process(item)
/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/jobs/pipeline.py:59: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
  return self.process(item)
========= sending heartbeat at 2026-10-06 14:50:23.712301
========= sending heartbeat at 2026-10-06 14:50:33.730857
========= sending heartbeat at 2026-10-06 14:50:43.748749
<string>:1: RuntimeWarning: More than 20 figures have been opened. Figures created through the pyplot interface (`matplotlib.pyplot.figure`) are retained until explicitly closed and may consume too much memory. (To control this warning, see the rcParam `figure.max_open_warning`). Consider using `matplotlib.pyplot.close()`.
========= sending heartbeat at 2026-10-06 14:50:53.766918
========= sending heartbeat at 2026-10-06 14:51:03.784713
========= sending heartbeat at 2026-10-06 14:51:13.802294
========= sending heartbeat at 2026-10-06 14:51:23.820702
========= sending heartbeat at 2026-10-06 14:51:33.838367
========= sending heartbeat at 2026-10-06 14:51:43.857397
========= sending heartbeat at 2026-10-06 14:51:53.876298
========= sending heartbeat at 2026-10-06 14:52:03.895988
========= sending heartbeat at 2026-10-06 14:52:13.914957
========= sending heartbeat at 2026-10-06 14:52:23.932300
========= sending heartbeat at 2026-10-06 14:52:33.951074
========= sending heartbeat at 2026-10-06 14:52:43.968286
========= sending heartbeat at 2026-10-06 14:52:54.035688
========= sending heartbeat at 2026-10-06 14:53:04.054281
========= sending heartbeat at 2026-10-06 14:53:14.072653
========= sending heartbeat at 2026-10-06 14:53:24.090910
========= sending heartbeat at 2026-10-06 14:53:34.106248
========= sending heartbeat at 2026-10-06 14:53:44.124427
========= sending heartbeat at 2026-10-06 14:53:54.141944
========= sending heartbeat at 2026-10-06 14:54:04.160292
========= sending heartbeat at 2026-10-06 14:54:14.179387
========= sending heartbeat at 2026-10-06 14:54:24.197296
========= sending heartbeat at 2026-10-06 14:54:34.216066
========= sending heartbeat at 2026-10-06 14:54:44.234309
========= sending heartbeat at 2026-10-06 14:54:54.252074
========= sending heartbeat at 2026-10-06 14:55:04.269857
========= sending heartbeat at 2026-10-06 14:55:14.288136
========= sending heartbeat at 2026-10-06 14:55:24.306319
========= sending heartbeat at 2026-10-06 14:55:34.325392
========= sending heartbeat at 2026-10-06 14:55:44.344276
========= sending heartbeat at 2026-10-06 14:55:54.362417
========= sending heartbeat at 2026-10-06 14:56:04.381593
========= sending heartbeat at 2026-10-06 14:56:14.399591
========= sending heartbeat at 2026-10-06 14:56:24.418208
========= sending heartbeat at 2026-10-06 14:56:34.429587
========= sending heartbeat at 2026-10-06 14:56:44.447820
========= sending heartbeat at 2026-10-06 14:56:54.466160
========= sending heartbeat at 2026-10-06 14:57:04.483295
========= sending heartbeat at 2026-10-06 14:57:14.499656
========= sending heartbeat at 2026-10-06 14:57:24.517705
========= sending heartbeat at 2026-10-06 14:57:34.535969
========= sending heartbeat at 2026-10-06 14:57:44.554292
========= sending heartbeat at 2026-10-06 14:57:54.571905
========= sending heartbeat at 2026-10-06 14:58:04.590297
========= sending heartbeat at 2026-10-06 14:58:14.609290
========= sending heartbeat at 2026-10-06 14:58:24.628105
========= sending heartbeat at 2026-10-06 14:58:34.645966
========= sending heartbeat at 2026-10-06 14:58:44.665394
========= sending heartbeat at 2026-10-06 14:58:54.684097
========= sending heartbeat at 2026-10-06 14:59:04.702555
========= sending heartbeat at 2026-10-06 14:59:14.720532
========= sending heartbeat at 2026-10-06 14:59:24.738648
========= sending heartbeat at 2026-10-06 14:59:34.756804
========= sending heartbeat at 2026-10-06 14:59:44.774956
========= sending heartbeat at 2026-10-06 14:59:54.793313
========= sending heartbeat at 2026-10-06 15:00:04.810331
========= sending heartbeat at 2026-10-06 15:00:14.828460
========= sending heartbeat at 2026-10-06 15:00:24.846665
========= sending heartbeat at 2026-10-06 15:00:34.864303
========= sending heartbeat at 2026-10-06 15:00:44.882759
========= sending heartbeat at 2026-10-06 15:00:54.898585
========= sending heartbeat at 2026-10-06 15:01:04.916784
========= sending heartbeat at 2026-10-06 15:01:14.935388
========= sending heartbeat at 2026-10-06 15:01:24.953604
========= sending heartbeat at 2026-10-06 15:01:34.971292
========= sending heartbeat at 2026-10-06 15:01:44.989802
========= sending heartbeat at 2026-10-06 15:01:55.007298
========= sending heartbeat at 2026-10-06 15:02:05.025293
========= sending heartbeat at 2026-10-06 15:02:15.043614
========= sending heartbeat at 2026-10-06 15:02:25.061018
========= sending heartbeat at 2026-10-06 15:02:35.079739
========= sending heartbeat at 2026-10-06 15:02:45.098284
========= sending heartbeat at 2026-10-06 15:02:55.116622
========= sending heartbeat at 2026-10-06 15:03:05.134456
========= sending heartbeat at 2026-10-06 15:03:15.152735
========= sending heartbeat at 2026-10-06 15:03:25.169474
========= sending heartbeat at 2026-10-06 15:03:35.182935
========= sending heartbeat at 2026-10-06 15:03:45.201725
========= sending heartbeat at 2026-10-06 15:03:55.219723
========= sending heartbeat at 2026-10-06 15:04:05.238888
========= sending heartbeat at 2026-10-06 15:04:15.256004
========= sending heartbeat at 2026-10-06 15:04:25.274405
========= sending heartbeat at 2026-10-06 15:04:35.292294
========= sending heartbeat at 2026-10-06 15:04:45.310776
========= sending heartbeat at 2026-10-06 15:04:55.329088
========= sending heartbeat at 2026-10-06 15:05:05.346731
========= sending heartbeat at 2026-10-06 15:05:15.365636
========= sending heartbeat at 2026-10-06 15:05:25.384022
========= sending heartbeat at 2026-10-06 15:05:35.402019
========= sending heartbeat at 2026-10-06 15:05:45.420300
========= sending heartbeat at 2026-10-06 15:05:55.439319
========= sending heartbeat at 2026-10-06 15:06:05.457293
========= sending heartbeat at 2026-10-06 15:06:15.475550
========= sending heartbeat at 2026-10-06 15:06:25.493479
========= sending heartbeat at 2026-10-06 15:06:35.511045
========= sending heartbeat at 2026-10-06 15:06:45.529133
========= sending heartbeat at 2026-10-06 15:06:55.547532
========= sending heartbeat at 2026-10-06 15:07:05.565664
========= sending heartbeat at 2026-10-06 15:07:15.583925
========= sending heartbeat at 2026-10-06 15:07:25.602581
========= sending heartbeat at 2026-10-06 15:07:35.619602
========= sending heartbeat at 2026-10-06 15:07:45.637487
========= sending heartbeat at 2026-10-06 15:07:55.655586
========= sending heartbeat at 2026-10-06 15:08:05.673917
========= sending heartbeat at 2026-10-06 15:08:15.692319
========= sending heartbeat at 2026-10-06 15:08:25.710302
========= sending heartbeat at 2026-10-06 15:08:35.728453
========= sending heartbeat at 2026-10-06 15:08:45.747808
========= sending heartbeat at 2026-10-06 15:08:55.766721
========= sending heartbeat at 2026-10-06 15:09:05.784304
========= sending heartbeat at 2026-10-06 15:09:15.802895
========= sending heartbeat at 2026-10-06 15:09:25.820355
========= sending heartbeat at 2026-10-06 15:09:35.837334
========= sending heartbeat at 2026-10-06 15:09:45.855534
========= sending heartbeat at 2026-10-06 15:09:55.873928
========= sending heartbeat at 2026-10-06 15:10:05.892292
========= sending heartbeat at 2026-10-06 15:10:15.910295
========= sending heartbeat at 2026-10-06 15:10:25.929519
========= sending heartbeat at 2026-10-06 15:10:35.947507
========= sending heartbeat at 2026-10-06 15:10:45.966238
========= sending heartbeat at 2026-10-06 15:10:55.985217
========= sending heartbeat at 2026-10-06 15:11:06.003300
========= sending heartbeat at 2026-10-06 15:11:16.022398
========= sending heartbeat at 2026-10-06 15:11:26.040825
========= sending heartbeat at 2026-10-06 15:11:36.060435
========= sending heartbeat at 2026-10-06 15:11:46.079173
========= sending heartbeat at 2026-10-06 15:11:56.098626
========= sending heartbeat at 2026-10-06 15:12:06.116296
========= sending heartbeat at 2026-10-06 15:12:16.135307
========= sending heartbeat at 2026-10-06 15:12:26.153299
========= sending heartbeat at 2026-10-06 15:12:36.172315
========= sending heartbeat at 2026-10-06 15:12:46.191949
========= sending heartbeat at 2026-10-06 15:12:56.208943
========= sending heartbeat at 2026-10-06 15:13:06.229796
========= sending heartbeat at 2026-10-06 15:13:16.248446
========= sending heartbeat at 2026-10-06 15:13:26.266627
========= sending heartbeat at 2026-10-06 15:13:36.285151
========= sending heartbeat at 2026-10-06 15:13:46.304275
========= sending heartbeat at 2026-10-06 15:13:56.322247
========= sending heartbeat at 2026-10-06 15:14:06.340188
========= sending heartbeat at 2026-10-06 15:14:16.358537
========= sending heartbeat at 2026-10-06 15:14:26.373610
========= sending heartbeat at 2026-10-06 15:14:36.391738
========= sending heartbeat at 2026-10-06 15:14:46.409950
========= sending heartbeat at 2026-10-06 15:14:56.428296
========= sending heartbeat at 2026-10-06 15:15:06.446974
========= sending heartbeat at 2026-10-06 15:15:16.465340
========= sending heartbeat at 2026-10-06 15:15:26.488508
========= sending heartbeat at 2026-10-06 15:15:36.506947
========= sending heartbeat at 2026-10-06 15:15:46.525892
========= sending heartbeat at 2026-10-06 15:15:56.543938
========= sending heartbeat at 2026-10-06 15:16:06.562293
========= sending heartbeat at 2026-10-06 15:16:16.581357
========= sending heartbeat at 2026-10-06 15:16:26.601154
========= sending heartbeat at 2026-10-06 15:16:36.620310
========= sending heartbeat at 2026-10-06 15:16:46.638338
========= sending heartbeat at 2026-10-06 15:16:56.656712
========= sending heartbeat at 2026-10-06 15:17:06.674898
========= sending heartbeat at 2026-10-06 15:17:16.693870
========= sending heartbeat at 2026-10-06 15:17:26.711485
========= sending heartbeat at 2026-10-06 15:17:36.729674
========= sending heartbeat at 2026-10-06 15:17:46.746281
========= sending heartbeat at 2026-10-06 15:17:56.763046
========= sending heartbeat at 2026-10-06 15:18:06.781288
========= sending heartbeat at 2026-10-06 15:18:16.799844
========= sending heartbeat at 2026-10-06 15:18:26.818797
========= sending heartbeat at 2026-10-06 15:18:36.835810
========= sending heartbeat at 2026-10-06 15:18:46.854590
========= sending heartbeat at 2026-10-06 15:18:56.871583
========= sending heartbeat at 2026-10-06 15:19:06.890151
========= sending heartbeat at 2026-10-06 15:19:16.907291
========= sending heartbeat at 2026-10-06 15:19:26.925526
========= sending heartbeat at 2026-10-06 15:19:36.937626
========= sending heartbeat at 2026-10-06 15:19:46.956015
========= sending heartbeat at 2026-10-06 15:19:56.973294
========= sending heartbeat at 2026-10-06 15:20:06.989926
========= sending heartbeat at 2026-10-06 15:20:17.005301
========= sending heartbeat at 2026-10-06 15:20:27.023719
========= sending heartbeat at 2026-10-06 15:20:37.036349
========= sending heartbeat at 2026-10-06 15:20:47.054379
========= sending heartbeat at 2026-10-06 15:20:57.073076
========= sending heartbeat at 2026-10-06 15:21:07.091312
========= sending heartbeat at 2026-10-06 15:21:17.109480
========= sending heartbeat at 2026-10-06 15:21:27.127108
========= sending heartbeat at 2026-10-06 15:21:37.145873
========= sending heartbeat at 2026-10-06 15:21:47.164042
========= sending heartbeat at 2026-10-06 15:21:57.182527
========= sending heartbeat at 2026-10-06 15:22:07.201376
========= sending heartbeat at 2026-10-06 15:22:17.219182
========= sending heartbeat at 2026-10-06 15:22:27.238194
========= sending heartbeat at 2026-10-06 15:22:37.256426
========= sending heartbeat at 2026-10-06 15:22:47.274565
========= sending heartbeat at 2026-10-06 15:22:57.292908
========= sending heartbeat at 2026-10-06 15:23:07.310920
========= sending heartbeat at 2026-10-06 15:23:17.329379
========= sending heartbeat at 2026-10-06 15:23:27.361656
========= sending heartbeat at 2026-10-06 15:23:37.380010
========= sending heartbeat at 2026-10-06 15:23:47.397299
========= sending heartbeat at 2026-10-06 15:23:57.416305
========= sending heartbeat at 2026-10-06 15:24:07.435827
========= sending heartbeat at 2026-10-06 15:24:17.455135
========= sending heartbeat at 2026-10-06 15:24:27.474725
========= sending heartbeat at 2026-10-06 15:24:37.493970
========= sending heartbeat at 2026-10-06 15:24:47.512736
========= sending heartbeat at 2026-10-06 15:24:57.531371
========= sending heartbeat at 2026-10-06 15:25:07.549293
========= sending heartbeat at 2026-10-06 15:25:17.560616
========= sending heartbeat at 2026-10-06 15:25:27.578752
========= sending heartbeat at 2026-10-06 15:25:37.595297
========= sending heartbeat at 2026-10-06 15:25:47.614293
========= sending heartbeat at 2026-10-06 15:25:57.632674
========= sending heartbeat at 2026-10-06 15:26:07.650737
========= sending heartbeat at 2026-10-06 15:26:17.669507
========= sending heartbeat at 2026-10-06 15:26:27.687289
========= sending heartbeat at 2026-10-06 15:26:37.706717
========= sending heartbeat at 2026-10-06 15:26:47.724641
========= sending heartbeat at 2026-10-06 15:26:57.742301
========= sending heartbeat at 2026-10-06 15:27:07.761122
========= sending heartbeat at 2026-10-06 15:27:17.779281
========= sending heartbeat at 2026-10-06 15:27:27.796714
========= sending heartbeat at 2026-10-06 15:27:37.813295
========= sending heartbeat at 2026-10-06 15:27:47.831303
========= sending heartbeat at 2026-10-06 15:27:57.849867
========= sending heartbeat at 2026-10-06 15:28:07.867751
========= sending heartbeat at 2026-10-06 15:28:17.881060
========= sending heartbeat at 2026-10-06 15:28:27.899236
========= sending heartbeat at 2026-10-06 15:28:37.917726
========= sending heartbeat at 2026-10-06 15:28:47.935989
========= sending heartbeat at 2026-10-06 15:28:57.956213
========= sending heartbeat at 2026-10-06 15:29:07.974449
========= sending heartbeat at 2026-10-06 15:29:17.988352
========= sending heartbeat at 2026-10-06 15:29:28.006755
========= sending heartbeat at 2026-10-06 15:29:38.071466
========= sending heartbeat at 2026-10-06 15:29:48.089508
========= sending heartbeat at 2026-10-06 15:29:58.107345
========= sending heartbeat at 2026-10-06 15:30:08.125635
========= sending heartbeat at 2026-10-06 15:30:18.143491
========= sending heartbeat at 2026-10-06 15:30:28.161844
========= sending heartbeat at 2026-10-06 15:30:38.180196
========= sending heartbeat at 2026-10-06 15:30:48.198290
========= sending heartbeat at 2026-10-06 15:30:58.216300
========= sending heartbeat at 2026-10-06 15:31:08.235566
========= sending heartbeat at 2026-10-06 15:31:18.253700
========= sending heartbeat at 2026-10-06 15:31:28.271886
========= sending heartbeat at 2026-10-06 15:31:38.290309
========= sending heartbeat at 2026-10-06 15:31:48.308338
========= sending heartbeat at 2026-10-06 15:31:58.325296
========= sending heartbeat at 2026-10-06 15:32:08.343296
========= sending heartbeat at 2026-10-06 15:32:18.364053
========= sending heartbeat at 2026-10-06 15:32:28.382724
========= sending heartbeat at 2026-10-06 15:32:38.402107
========= sending heartbeat at 2026-10-06 15:32:48.420305
========= sending heartbeat at 2026-10-06 15:32:58.438288
========= sending heartbeat at 2026-10-06 15:33:08.456558
========= sending heartbeat at 2026-10-06 15:33:18.467346
========= sending heartbeat at 2026-10-06 15:33:28.476318
========= sending heartbeat at 2026-10-06 15:33:38.494233

And is then killed by the cryosparc hearbeat monitor

.. suggests a lower-level (outside CryoSPARC software) hardware or software problem.

Out of curiosity: What was the reason for not installing CryoSPARC v5?

Hi @wtempel , you may be right. But is there any way to use the cryosparc logs to give any clues what might be going wrong? So far it’s been specifically Patch Motion Correction jobs and Non-Uniform Refinement jobs that trigger the crash (see below for last part of NU refine job log). Running the Patch Motion Correction jobs with only 3 out of 4 GPUs seemed to work for some reason, but NU refine only uses 1 GPU so difficult to bypass that way. I can try locking the NU refine to a particular GPU at a time to see if it’s one of the GPUs that’s gone bad?

As for the cryosparc version - we were trying to resurrect one of our servers after the OS got corrupted (perhaps due to cryosparc cache being mistakenly written into the root directory due to unmounted SSD). Out of paranoia, I was trying to restore everything to the same version as previously in case of compatibility issues. We previously didn’t have any issues like this as far as I know.

Many thanks!

Louis

P.s. this is a standalone installation.

========= sending heartbeat at 2026-10-08 14:22:01.940018
gpufft: creating new cufft plan (plan id 0   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 480 0 
	istride 1 
	idist   230400 
	onembed 480 480 0 
	ostride 1 
	odist   230400 
	batch   500 
	type    C2C 
	wkspc   automatic 
	Python traceback:

gpufft: creating new cufft plan (plan id 1   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 480 0 
	istride 1 
	idist   230400 
	onembed 480 480 0 
	ostride 1 
	odist   230400 
	batch   500 
	type    C2C 
	wkspc   automatic 
	Python traceback:

HOST ALLOCATION FUNCTION: using numba.cuda.pinned_array
/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/cudadrv/nvrtc.py:257: UserWarning: NVRTC log messages whilst compiling kernel:

kernel(35): warning #68-D: integer conversion resulted in a change of sign

kernel(44): warning #68-D: integer conversion resulted in a change of sign

kernel(17): warning #177-D: variable "N_I" was declared but never referenced


  warnings.warn(msg)
========= sending heartbeat at 2026-10-08 14:22:11.958542
<string>:1: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
========= sending heartbeat at 2026-10-08 14:22:21.977342
========= sending heartbeat at 2026-10-08 14:22:31.995872
/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/plotutil.py:571: RuntimeWarning: divide by zero encountered in log
  logabs = n.log(n.abs(fM))
========= sending heartbeat at 2026-10-08 14:22:42.016766
gpufft: creating new cufft plan (plan id 2   pid 169845) 
	gpu_id  1 
	ndims   3 
	dims    480 480 480 
	inembed 480 480 482 
	istride 1 
	idist   111052800 
	onembed 480 480 241 
	ostride 1 
	odist   55526400 
	batch   1 
	type    R2C 
	wkspc   automatic 
	Python traceback:

========= sending heartbeat at 2026-10-08 14:22:52.034949
gpufft: creating new cufft plan (plan id 3   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 482 0 
	istride 1 
	idist   231360 
	onembed 480 241 0 
	ostride 1 
	odist   115680 
	batch   500 
	type    R2C 
	wkspc   automatic 
	Python traceback:

gpufft: creating new cufft plan (plan id 4   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 482 0 
	istride 1 
	idist   231360 
	onembed 480 241 0 
	ostride 1 
	odist   115680 
	batch   500 
	type    R2C 
	wkspc   automatic 
	Python traceback:

/srv2/home/lab/cryosparc/cryosparc_worker/deps/anaconda/envs/cryosparc_worker_env/lib/python3.10/site-packages/numba/cuda/dispatcher.py:536: NumbaPerformanceWarning: Grid size 1 will likely result in GPU under-utilization due to low occupancy.
  warn(NumbaPerformanceWarning(msg))
========= sending heartbeat at 2026-10-08 14:23:02.047531
gpufft: creating new cufft plan (plan id 5   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 482 0 
	istride 1 
	idist   231360 
	onembed 480 241 0 
	ostride 1 
	odist   115680 
	batch   132 
	type    R2C 
	wkspc   automatic 
	Python traceback:

========= sending heartbeat at 2026-10-08 14:23:12.066076
<string>:1: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
========= sending heartbeat at 2026-10-08 14:23:22.078017
gpufft: creating new cufft plan (plan id 6   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 482 0 
	istride 1 
	idist   231360 
	onembed 480 241 0 
	ostride 1 
	odist   115680 
	batch   133 
	type    R2C 
	wkspc   automatic 
	Python traceback:

========= sending heartbeat at 2026-10-08 14:23:32.096322
========= sending heartbeat at 2026-10-08 14:23:42.115054
========= sending heartbeat at 2026-10-08 14:23:52.133360
========= sending heartbeat at 2026-10-08 14:24:02.147505
========= sending heartbeat at 2026-10-08 14:24:12.165839
========= sending heartbeat at 2026-10-08 14:24:22.183051
gpufft: creating new cufft plan (plan id 7   pid 169845) 
	gpu_id  1 
	ndims   3 
	dims    480 480 480 
	inembed 480 480 482 
	istride 1 
	idist   111052800 
	onembed 480 480 241 
	ostride 1 
	odist   55526400 
	batch   1 
	type    R2C 
	wkspc   manual 
	Python traceback:

========= sending heartbeat at 2026-10-08 14:24:32.202323
gpufft: creating new cufft plan (plan id 8   pid 169845) 
	gpu_id  1 
	ndims   3 
	dims    240 240 240 
	inembed 240 240 121 
	istride 1 
	idist   6969600 
	onembed 240 240 242 
	ostride 1 
	odist   13939200 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

<string>:1: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
========= sending heartbeat at 2026-10-08 14:24:42.220802
<string>:1: RuntimeWarning: invalid value encountered in true_divide
========= sending heartbeat at 2026-10-08 14:24:52.237018
gpufft: creating new cufft plan (plan id 9   pid 169845) 
	gpu_id  1 
	ndims   3 
	dims    480 480 480 
	inembed 480 480 241 
	istride 1 
	idist   55526400 
	onembed 480 480 482 
	ostride 1 
	odist   111052800 
	batch   1 
	type    C2R 
	wkspc   manual 
	Python traceback:

========= sending heartbeat at 2026-10-08 14:25:02.255382
========= sending heartbeat at 2026-10-08 14:25:12.273328
========= sending heartbeat at 2026-10-08 14:25:22.294324
========= sending heartbeat at 2026-10-08 14:25:32.310320
========= sending heartbeat at 2026-10-08 14:25:42.328335
========= sending heartbeat at 2026-10-08 14:25:52.354321
========= sending heartbeat at 2026-10-08 14:26:02.370309
<string>:1: DeprecationWarning: `np.bool` is a deprecated alias for the builtin `bool`. To silence this warning, use `bool` by itself. Doing this will not modify any behavior and is safe. If you specifically wanted the numpy scalar type, use `np.bool_` here.
Deprecated in NumPy 1.20; for more details and guidance: https://numpy.org/devdocs/release/1.20.0-notes.html#deprecations
========= sending heartbeat at 2026-10-08 14:26:12.388364
========= sending heartbeat at 2026-10-08 14:26:22.407336
========= sending heartbeat at 2026-10-08 14:26:32.425369
========= sending heartbeat at 2026-10-08 14:26:42.444336
========= sending heartbeat at 2026-10-08 14:26:52.463840
========= sending heartbeat at 2026-10-08 14:27:02.482310
========= sending heartbeat at 2026-10-08 14:27:12.503321
========= sending heartbeat at 2026-10-08 14:27:22.530189
<string>:1: DeprecationWarning: `np.bool` is a deprecated alias for the builtin `bool`. To silence this warning, use `bool` by itself. Doing this will not modify any behavior and is safe. If you specifically wanted the numpy scalar type, use `np.bool_` here.
Deprecated in NumPy 1.20; for more details and guidance: https://numpy.org/devdocs/release/1.20.0-notes.html#deprecations
========= sending heartbeat at 2026-10-08 14:27:32.547381
========= sending heartbeat at 2026-10-08 14:27:42.565987
========= sending heartbeat at 2026-10-08 14:27:52.584369
========= sending heartbeat at 2026-10-08 14:28:02.602370
========= sending heartbeat at 2026-10-08 14:28:12.621228
/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/sigproc.py:660: FutureWarning: `rcond` parameter will change to the default of machine precision times ``max(M, N)`` where M and N are the input matrix dimensions.
To use the future default and silence this warning we advise to pass `rcond=None`, to keep using the old, explicitly pass `rcond=-1`.
  x = n.linalg.lstsq(w.reshape((-1,1))*A, w*b)[0]
========= sending heartbeat at 2026-10-08 14:28:22.639773
========= sending heartbeat at 2026-10-08 14:28:32.658421
/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/plotutil.py:571: RuntimeWarning: divide by zero encountered in log
  logabs = n.log(n.abs(fM))
========= sending heartbeat at 2026-10-08 14:28:42.677330
/srv2/home/lab/cryosparc/cryosparc_worker/cryosparc_compute/plotutil.py:44: RuntimeWarning: invalid value encountered in sqrt
  cradwn = n.sqrt(cradwn)
========= sending heartbeat at 2026-10-08 14:28:52.695327
========= sending heartbeat at 2026-10-08 14:29:02.714259
========= sending heartbeat at 2026-10-08 14:29:12.728942
========= sending heartbeat at 2026-10-08 14:29:22.748174
========= sending heartbeat at 2026-10-08 14:29:32.767276
========= sending heartbeat at 2026-10-08 14:29:42.786409
========= sending heartbeat at 2026-10-08 14:29:52.806488
========= sending heartbeat at 2026-10-08 14:30:02.826987
========= sending heartbeat at 2026-10-08 14:30:12.846355
========= sending heartbeat at 2026-10-08 14:30:22.865300
========= sending heartbeat at 2026-10-08 14:30:32.884650
========= sending heartbeat at 2026-10-08 14:30:42.899966
gpufft: creating new cufft plan (plan id 10   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 482 0 
	istride 1 
	idist   231360 
	onembed 480 241 0 
	ostride 1 
	odist   115680 
	batch   136 
	type    R2C 
	wkspc   automatic 
	Python traceback:

<string>:1: UserWarning: Cannot manually free CUDA array; will be freed when garbage collected
========= sending heartbeat at 2026-10-08 14:30:52.919388
========= sending heartbeat at 2026-10-08 14:31:02.938836
========= sending heartbeat at 2026-10-08 14:31:12.958238
========= sending heartbeat at 2026-10-08 14:31:22.976768
========= sending heartbeat at 2026-10-08 14:31:32.996819
========= sending heartbeat at 2026-10-08 14:31:43.016735
========= sending heartbeat at 2026-10-08 14:31:53.035819
gpufft: creating new cufft plan (plan id 11   pid 169845) 
	gpu_id  1 
	ndims   2 
	dims    480 480 0 
	inembed 480 482 0 
	istride 1 
	idist   231360 
	onembed 480 241 0 
	ostride 1 
	odist   115680 
	batch   137 
	type    R2C 
	wkspc   automatic 
	Python traceback:

========= sending heartbeat at 2026-10-08 14:32:03.054153
========= sending heartbeat at 2026-10-08 14:32:13.072451
========= sending heartbeat at 2026-10-08 14:32:23.146138

A post was split to a new topic: CUDA_ERROR_OUT_OF_MEMORY in nonuniform refinement