Processing a highly heterogeneous membrane protein dataset (~23k micrographs) – looking for workflow advice

Hi everyone,

I’m relatively new to cryo-EM image processing and would appreciate some advice on a challenging membrane protein dataset.

The target is a native membrane protein complex, purified in 1% DDM + 0.1%CHS, then exchanged to 0.02% GDN. The sample appears highly heterogeneous, and the exact composition/oligomeric state is unknown. Based on biochemical data and AlphaFold3 predictions, I expect the particle to be around 15 nm, although it could potentially be somewhat larger.

I have also processed the dataset in RELION, but unfortunately without better results.

Dataset

  • ~23,000 micrographs

  • Pixel size: 0.74 Å/px

  • Initial blob picking:

    • Minimum diameter: 90 Å

    • Maximum diameter: 200 Å

  • ~6 million particles extracted

  • Box size: 384 px

  • Fourier cropped (binned) to 192 px

Initial processing

To make processing manageable, I split the dataset into batches of ~500,000 particles and performed multiple rounds of 2D classification on each subset (parameters attached in screenshot).

After selecting reasonable classes, I merged the selected particles and performed additional rounds of 2D classification.

The resulting workflow was roughly:

Blob picker → Extract → Multiple rounds of 2D classification on subsets → Merge selected particles → Additional 2D classification

Template generation

I then generated ab initio volumes. Heterogeneous refinement and Non-Uniform refinement did not produce any high-quality maps, so I used one of the ab initio volumes that appeared most promising as a template (screenshot attached).

Screenshot 2026-06-22 at 13.55.37

I performed template picking using these volumes and repeated 2D-classification workflow with changes parameters (screenshot attached).

Current results

Some of the 2D classes look potentially interesting (attached screenshots). A few classes appear vaguely consistent with features expected from my AlphaFold3 model, but honestly it is difficult to say with confidence.

I also tried using AlphaFold3 model for template picking, but this did not improve the results.

Questions

  1. Does this overall workflow seem reasonable for a highly heterogeneous membrane protein sample?

  2. Are there any obvious processing steps or parameters that you would change?

  3. Would you recommend trying alternative approaches/software (Topaz, CryoDRGN, 3DVA, ISAC, etc.) at this stage?

  4. Is there anything in the attached 2D classes or ab initio volumes that suggests I may actually be looking at real particles rather than noise/contaminants?

  5. For a difficult membrane protein dataset like this, what would you try next?

Any feedback or suggestions would be greatly appreciated. I’m still learning and would be very grateful for an assessment of whether I’m following a sensible workflow or heading in the wrong direction.

Thank you!

Can’t see any protein-like features here - how do the micrographs look?

Cheers

Oli

Thanks Oli for the suggestion. I’ve attached representative micrographs. The top row shows the original micrographs and the bottom row shows denoised versions. The particles marked with red circles correspond to ATP synthase, which was identified as a contaminant during processing. My target is ~15 nm expected size, but I have not been able to confidently identify a distinct particle population corresponding to it. I’d be interested to hear whether anyone can spot a convincing particle population in these micrographs or suggest additional processing strategies.

yes there is clearly a globular, protein shape significantly smaller than the ATP synthase particles in your micrographs which is worth picking/investigating.
agree with Oli the 2D are largely not useful, though from them you could select a single favorite class and try reclassification to 30 classes with 1000 batchsize.
did you “inspect picks” to see if the 6million are reasonable?
yes, Topaz could possibly help - can you take the particles from your better ab initio class and run 2D with 1000 batchsize to see if you get a good class(es)? Seems that picking is not going well and that will make 2D not go well, though I wouldn’t expect template picking to yield this bad of results.
I can’t comment on ini class uncerainty or maximum res or re-center threshold, but defaults should work.

Thanks for the suggestions! :slight_smile:

I suspect the main challenge may indeed be the picking. This is a native membrane protein complex with relatively low expression levels, and ATP synthase appears to be a major contaminant. Based on the micrographs, I would estimate that ATP synthase-like particles account for a large fraction of the visible protein population, making it difficult to distinguish my target particles from contaminants, noise, or potentially detergent micelles.

As a result, I suspect that only a relatively small fraction of the extracted particles actually correspond to my target complex, which makes the picking and classification particularly challenging….

I have already started as you advised taking particles from the most promising ab initio class and running another round of 2D classification with a batch size of 1000 and selecting my best-looking 2D class and performing reclassification into 30 classes with a batch size of 1000. Let’s see if that helps.