Recurring amino acid clustering pattern in >160,000 protein structures: looking for independent visual confirmation
Disclosure: I am part of the team behind this project. Posting here to reach independent observers who can help validate or challenge the finding.
Big-data analysis of over 160,000 crystallographic structures from the PDB suggests a previously unreported general pattern: beyond the well-known hydrophobic core, amino acids appear to cluster by chemical family (polar, acidic, basic, special) into groups following specific size (approximately 8 residues) and shape.
We call this the Mosaic Q pattern, quantified by a parameter Q based on Euclidean distances between same-type residues.
Resources:
- Preprint
- GitHub
- Dataset (HuggingFace)
- Image repository (FairSharing)
- Q descriptor software (bio.tools) / PyPI
Note: the preprint predates the current version of the analysis. The updated methodology and results are documented in the GitHub repository.
Why we're posting here
The pattern can also be assessed visually using Jmol. We are building an open repository of protein structure images contributed by independent observers. Dozens of contributors have already submitted images, and so far the pattern appears consistently across them. The more people render and submit their own images, the stronger (or weaker) the overall evidence becomes. We are particularly interested in:
- Cases where the pattern is not apparent (potential counterexamples are as valuable as confirmations)
- Methodological feedback on the Mosaic Q parameter or the statistical analysis
- Suggestions for proteins or structural families worth prioritizing
Participation requires no special background beyond basic familiarity with protein structure visualization. Instructions and the image submission form are at proteins-mosaic-q.org.
We welcome skeptics as much as supporters: independent scrutiny is the point.
0 answers
No answers yet.
Log in to answer this question.
Genuine question on the Mosaic Q parameter: have you tested whether the ~8-residue grouping holds when you control for protein family size and functional class? If the clustering is invariant across function, that's a much stronger structural claim than if it tracks with family-level evolutionary pressure.
I get your point, the fact is given that the analysis was based on the >160,000 crystallographic PDB structures available at the time had a strong signal (R² = 0.979) we reckon the property must be relatively conserved accross families. Nevertheless, an analysis per family would indeed shed more insight into this structural feature.
On the cluster size, we tested several hypothetical configurations against the real data, and the model that better reproduced the empirical equation was that of the hydrophobic core along with clusters of ~8 residues according to the chemical type (polar, acidic, basic, special), which seems confirmed by visual inspection of hundreds of structures. However, more research is still needed.