Hello, and I have a question about Particle sets tool. I thought, in the “Balance half sets” mode, results should be the two half sets that have same number of particles. But in my case, result is a bit weird. Please check my case. I tried NU Ref. with initial 130k particles but result panel says It only used 32k.
@Randomnick Please can you post
- you CryoSPARC version
- a section of the job tree that includes the Particle Sets Tool and Nonuniform Refinement jobs you mentioned, with job cards in Outputs View.
I’m late, sorry. But I deleted the job, so I can’t attach the job tree…
I use ver4.7.1+251124.
The particles came from three multiclass-reconstruction jobs(used same input particles), and pooled into NU-refine(1). After that, 3D classification with 5 class was done. I chose a promising 130k particles from a class and conducted NU-refine(2) with a class volume output. But, only 30k particles were used, and I saw the caution about Half set size difference.. I think it was about 50% differ or more.
I attach the example

Hi @Randomnick! What you’re seeing here is the expected behavior. The “Balance half sets” mode of Particle Sets Tool drops particles from the larger set to make them the same size. This is because re-splitting particles (i.e., moving particles between half sets) but keeping the same reference violates the gold standard half-set independence assumption.
For moderate differences in size, it is probably fine to re-split the particles. You can do this by turning on Force re-do half-set split in a reconstruction or refinement. However, for a difference this large, I’d recommend first running an Ab-Initio Reconstruction to generate a new reference. This insures that resplitting the particles won’t introduce too much model bias.
Please let me know what questions you still have!
Is there an option to remove the previous half-set split? I collected particles from copies of the ab initio reconstruction and generated another initial volume for refinement, but refinement still says Set A is greater than Set B, as I mentioned above.
Hi @Randomnick, I’m having a bit of trouble following you here. Could you give me a bit more detail?
- What jobs did you run, in what order, to collect particles from copies of the ab initio reconstruction?
- How did you generate the initial volume?
- What particles did you plug into the refinement?
- What is the exact text of the error message?
-
I selected ‘good’ classes from each ab initio reconstruction job, which show clear features of my protein. Then I used them as input for the subsequent ab initio job. The original input particles had many duplicates, but they were removed automatically, yielding fewer input particles.
-
I used an ab initio volume that showed a clear feature from one of the ab initio copies.
-
As I wrote above.
-
Same as I attached before; the only differences are the number of particles and the percent difference.
Plus, it was solved when I turned on the force re-do GS split. Thanks for your concern.
Hi @Randomnick! Any time you select a single class from a classification there’s a chance it will have unbalanced half sets, since there’s no way to ensure that the half sets each contain an equal proportion of the particles which end up in a given class. Generally, though, we’d expect this difference to be relatively small. Moving forward with force re-do GS split sounds good!

