Strange speed issues

This might be relevant if SLURM is configured with cgroup constraints on RAM and CPU resources, whereas a non-SLURM job would not be constrained in this way.

Have you tested and confirmed that this non-local cache actually improves particle read performance (versus reading particles directly from project directories)?

The [always] setting may lead to unpredictable performance problems. The effect of this setting may differ between nodes and may depend on external (other CryoSPARC or non-CryoSPARC) concurrent workloads. One may bypass this unpredictability with the [madvise] or [never] settings.

Not off the top of my head for v580. But in light of a post like RTX PRO 4000 Blackwell Xid 8 / CUDA_ERROR_LAUNCH_TIMEOUT in 2D Classification and Ab-initio - #3 by OleUns, the driver version should be a variable to consider.