Hosted Demo: https://cpop-v2-designer-txmbwpw7qifbcuvcatby2m.streamlit.app/?token=cpop_vip Feedback Form (Pharma/Biotech): https://forms.gle/qDi1JSoTr3c582VF9 Feedback Form (Academic): https://forms.gle/PS84eW3Hn7GVGwKN7
Hey biostars community,
My name is Joshua Haigler and I’m looking for technical feedback on CPOP V2, a revised computational platform for the precision design of artificial miRNases. The goal of this model is to speed up the lead optimization process of miRNA-based therapeutics that catalyitcally act on mRNA molecules to influence diseases like neurological disorders and cancers at significantly lower dosages than is currently possible.
Disclaimer: This post serves as a timestamped public disclosure of methodology and preliminary results. Reported metrics reflect validation on the full N=2,114 dataset; ablation results on core experimental records are available in the hosted demo at the time of posting this. A formal manuscript is in preparation.
The Evolution from V1 to V2 My initial version suffered from a small dataset of n=72, which was partially compensated for using LOOCV. CPOP V2 addresses this through a significantly expanded dataset that utilizes transfer learning.
Key Innovations in V2
- Expanded Dataset (N=2,114): The dataset moved from 72 to 2,114 records using a three-tier strategy involving literature mining, surrogate transfer from ribozyme/DNAzyme assays, and physics-augmented data generated via NUPACK-driven thermodynamic simulations.
- Dual-Architecture Approach: The system now utilizes a GATv2 Graph Neural Network (GNN) to encode molecular designs as graphs, capturing structural nuances that flat feature vectors miss. This is paired with a four-model ML ensemble (Random Forest, XGBoost, Gradient Boosting, and MLP) for performance prediction. Uncertainty Quantification: A Gaussian Process (GP) head provides calibrated confidence intervals, flagging designs in underexplored chemical spaces.
Current Performance Metrics Validation of the ensemble model using a 5-fold cross-validation protocol shows significant improvements over the baseline:
Metric
Accuracy (R²) 0.72 (V1) vs 0.842 (V2)
MAE 12.4% (V1) vs 7.15% (V2)
Reliability Low (Overfit) (V1) vs High (Generalizable) (V2)
Call to Action I would greatly appreciate the community's eyes on my methodology and performance. You can interact with the current model state and provide feedback through the links at the top of this post.
Specific Questions for the Community Methodology: Does the transition to a GATv2 GNN backbone effectively address the limitations of V1’s small dataset in drug design, or are there alternative architectures (e.g., Transformer-based) that might better capture the 3D interaction geometry of oligo-peptide conjugates? Pareto Optimization: My current Pareto Front balances target cleavage against the 2,656 human miRNAs in miRBase v22. Are there additional biological "noise" factors or off-target repositories you would recommend integrating? Symmetry Index Predictions: I am modeling the relationship between bilateral symmetry and catalytic rate as a sigmoid response curve. Based on your experience with enzyme kinetics, does this plateauing effect at high symmetry align with biological expectations, or should I consider a different functional form?
Regardless of the outcome, thank you for your time and expertise.
Joshua Haigler UNC Charlotte
0 answers
No answers yet.
Log in to answer this question.