How to process highly flexible 150 kDa protein complex?

Hi fellow structural biologists! I’m hoping to get wisdom/tips on this tough dataset on a relatively small (150 kDa), apparently highly flexible protein complex. This is a follow-up on a new collected dataset from the same protein complex that I’ve posted before ( HR-HAIR works well, but NU Refinement is failing? - #11 by olibclarke ). I think we’re so close to solving the structure that it hurts. I have tried many things, including local refinement. Some diagnostics below.

I have shown with size-exclusion chromatography that before/after vitrification, at least about half of the particles are the dimer (100 kDa) <> monomer (50 kDa) complex, after complex reconstituting from individual components. There seems to be an equilibrium of the two units.

This is the first round of 2D classification on blob-picked particles. I’m extracting with ~2x particle diameter for now, but I expect I’ll need to extract larger soon since my defocus range does seem on the higher end. (Average defocus around -1.1 but with a spread between -0.5 to -2.7.)

After multiple rounds of heterogenous refinement and ab-initio clean up, I was able to get the structure of the C2 empty dimer using C2-symmetry-imposed non-uniform refinement. To do this, I had to separate out the C2 empty dimer (which has very distinctive features, so that was doable) from the whole dimer:monomer complex. It appears rigid/stable enough to push to high resolution. This was done on 577,432 particles.

However, I am having trouble with the whole dimer:monomer complex. After C1-symmetry non-uniform refinement, I lose a lot of side chains I once saw in the “empty apo” C2 dimer. This was done on 400,910 particles.

I took this set of 400,910 particles and did a 2D classification with parameters as shown in @olibclarke’s HR-HAIR paper: 200 classes, maximum reconstruction resolution of 3 Å, initial classification uncertainty factor of 1, Circular mask diameter 80 Å, number of O-EM iterations 80, number of final full iterations 20, Batchsize per class 400. There clearly does appear to be some asymmetry induced by the monomer. The monomer does appear to “wobble” within the dimer.

To this end, I took this to 3D VA to see what might be going on. I used the particles/mask of the non-uniform refinement of the trimer. I used a filter resolution of 12 Å. (I screened a few other filter resolutions, but this one seemed best.)

Principal component 0:

output

Principal component 1:

output2

Principal component 2:

output2

Strangely enough, I when I try local refinement using masks to the monomer or dimer, I seem to get better resolution on the monomer, rather than the dimer. It does seem a little “noisy and spiky” though, so maybe this is some noise fitting. I have tried with/without recentering, using center as fulcrum or overriding the fulcrum at the monomer-dimer contact point that I can see visually. Recentering seems to have helped the most, though modestly (only about 0.1-0.2 Å resolution boost, but the map looks less “spiky” and noisy).

I am hoping the key is actually HR-HAIR. @olibclarke, I would love to hear your thoughts on what might be going on here. Here, I perform HR-HAIR in v4.7 with the ab-initio job (not the v5 implemented version, since my institution hasn’t helped us with updating our GPUs to handle v5 due to glibc…) and then take that set of aligned particles to a local refinement, masking the entire dimer:monomer trimer, the map looks much better. To note, I am starting at 5 Å resolution and going to a final 3 Å resolution. I wonder what might be going on here or if this is telling me something? I think once non-uniform refinement discards previous pose information, I lose a lot of resolution. But by doing local refinement on the entire region, I save that pose information from pseudo HR-HAIR ab initio.

Is there anything that I can do to help resolve this really flexible protein? I might not be able to get the whole protein complex, but if I could push it just enough to identify the contact sites to do biochemistry, that may be enough given the dimer is clearly resolved, and the helices on the dimer:monomer do seem to be there—just not as well-defined as in the only-dimer map…

Thank you all for your wisdom. I really appreciate it!

Hi,

thank you for all the detailed explanation, it is a very interesting case and I’m curious to what other’s suggestions might be.

It seems to me like having high-resolution reconstruction of your subunit will be very challenging. But it feels like you have a good enough reconstruction for the base C2 protein, and if that’s the case, your story will be about variability of the additional subunit, if it matches your biological data etc…

On the side, you tried 3DVA but maybe 3DFlex would give a better reconstruction, especially when the subunit has many poses with a wide angle of displacement/poses. Also, in such case, sorry for the off-topic/competition post, but perhaps give cryoDRGN a try as it seems to perform decently in such cases (although I haven’t been able to do so myself).

Best of luck

2 Likes

hi @MetabolicNerd, interesting! I am not the most expert here, but I would have tried a 3d classification with focus mask. try different class numbers, resolution (I would go something like 12Å first) to see if you can separate a subset that alignes. There is a beautiful tutorial on 3d classification of nucleosome bound particles here. Remember to try different masks (it may change a lot). Also try to optimize the mask iteratively, it may give better results from the initial particle subset. Also remember that the solvent mask should be big enough to “contain” your variability otherwise you will not classify it. HTH!

1 Like

I was about to suggest almost the same… except using 3DVA with a broad mask around the small protein. In clusters mode + NU-refinement there is hope that you improve resolution for that bit, but I am afraid 400 k ptcls won’t be enough as starting set with so much variability.

3DFlex might show the movement better if you use a custom mesh (one mesh for each protein, there are tutorials for this). I don’t like building models in 3DFlex maps, though (because they are deformations of an original map, not real reconstructions like in 3DVA clusters mode). But still, you might get useful insights from it.

1 Like

Worked on similar case of small protein right on the symmetry axis breaking symmetry. Used symmetry refinement to get high res then 3dclass to separate out the particles for the small protein in different orientations then aligned up the orientations. In your case how do 100kda part of maps compare, the 577k monomerless to 400k monomer containing? If the 100kda dimer part matches rather well, id suggest going backwards; combine to 977k, do C2NUR (try 400k and 577k as staring maps for 977k), should get higher res of the 100kda (if res does not improve might need to clean up the 400k stack more or the 577k vs 400k are too different).Note* Then your 400k 100kda has good particle orientations its just separating that out and sticking with local refinement with minor movements and separating 50kda poses. You already know the 400k so you can just particle tools intersect and keep the orientations from the 977k job for the 400k particles with homo reconstruct or local refinement pose gaussian small and then its to separate the two poses with 3dclass. Can always get better separation though so 3Dclass/3DVar on the 977k, often similar or better success then hetero at splitting. Your current 400k 3DVar PC2 shows 50kda monomer interacting to both side of 100kda dimer, so likely the 50kda is in mixed orientations with one being more populated but still merged in final result, 3dclass/3dvar want a mask made of the combination of the current C1NUR 400k and rotated 180deg and maximized, such mask cover both interfaces of the 50kda if doesn’t already. This approach uses the C2 sym of the 100kda to get the orientations such hopefully goes to better resolution upon local refine than the NUR trying to align it up.

Back to 400k 100kda monomer is in mixed orientations, do also C2 sym expand (hopefully 400k monomer bound is still C2), mask on half of 100kda dimer and closest position for 50kda, fulcrum on the interface, do local refinement pose gaussian small. Then same mask for 3dclass/3dvar Looking for 2 classes (try more classes as additional jobs) in the C2 expanded those with 50kda and those without. Those without are mostly other pose for 50kda but likely some without 50kda, see note below about aligning particles to other position, then repeat local refine and 3dclass/3dvar to see if can remove more of the 50kda less.

Note*
Alternative is to take the your current 400k C1 and align up maps/particles to 577k map then do local refinement with mask on the 100kda for the 977k as alternative approach to C2NUR, may need to subtract the 50 kda part of your 400k map and use that as the map to align to the 577k to update particle orientations, usually the best way to get align to work is to take 577k map in chimera, align to 400k and resample the 577k on 400k, such 577k is now 400k position. Take that new resampled 577k as the position of the 400k and will align perfectly with the 577k original. If cant get align to work, then just redo 577k C2NUR with the 400k map as input and particles should be close enough position.

1 Like

Thank you all for the great suggestions. I’ll give all these suggestions a go and hope to come back with promising updates!

I have already tried @carlos’s suggestions. (3DVA on just the small 50 kDa C1 protein) I think there’s possibly something dynamic happening here with “docking,” but I can’t be too sure yet. I’m also trying not to read so into it at this low-resolution stage to avoid biasing my interpretation of this potentially really cool biology. This job was also done with filter resolution 8 Å. (Rationale for 8 Å being that this is where I see a “dip” in resolution when (i) combining the two particle stacks for a NU-Refine as @T_Bird suggested and (ii) also just doing a NU refine. Is this maybe the correct interpretation of “when” variability happens based on the GSFSC curves?)

Principal component 0:

output0

Principal component 1:

output1

Principal component 2:

output2

And here’s a quick NU-Refine with C2 symmetry relaxation on all the particles. (Probably still needs some particle curation when using this many particles.) The dimer is “empty” and doesn’t have the monomer. There doesn’t seem to be many particles that underwent the symmetry relaxation, as shown in the pose-difference plot.

Yes, the bumps in FSC curves are very common when there is conformational heterogeneity. I makes sense that you have good alignment in low and high resolutions, with a disagreement where things move, like a “fulcrum in reciprocal space”. But from what I’ve read from other people, the bumps are not to be taken too seriously - FSC is a global measure of all particles in reciprocal space, so we can’t really know what is happening in real space. Now for this complex in particular, I do believe you need more particles. Like 5 or 10 times more, if you can.

Considering the results from masked 3DVA: yes, you possibly have that kind of movement of the small protein, but you can not assume that the large bit is not moving. Another possibly useful experiment is to local refine masking one lobe of the large protein, then running 3DVA masking the small protein. Just to see how much the movement of the small protein relates to the other lobe. This would be so much nicer if you could make reliable reconstructions in high resolution, but that I am afraid will only be possible with a much larger particle set.