Hi all, I’m looking for feedback on whether this type of work is realistically publishable as a speculative, hypothesis-generating study, rather than as definitive biological truth. We would be extremely conservative in our claims and explicitly frame this as proposing a mechanistic hypothesis rather than proving one.
Background
I’m studying a historically rare but increasingly frequent subtype of liver cancer that appears resistant to the standard drug used for more common liver cancers. The original goal was to identify candidate pathways that might plausibly explain this resistance and then validate them experimentally.
We initially planned to conduct cell culture and qPCR validation, but funding cuts eliminated this possibility. The available human bulk microarray cohorts and TCGA data are so poorly annotated that meaningful clinical validation isn’t possible. I contacted a group with semi-annotated data, but legal restrictions prevented further data sharing.
Despite this, my PI would like to pursue publication, specifically as a computational, hypothesis-generating paper, rather than a validation study. I'm the only computational person in the lab, and most of what I do is beyond her scope, so she's given me some time to brainstorm and figure out a viable path forward.
Analysis overview
Because human datasets for the rare cancer are extremely limited, I used mouse model scRNA-seq datasets, which have been shown in the literature to closely resemble human liver cancer transcriptional programs and are commonly used as stand-ins when human data are unavailable.
1. Ortholog mapping & cell selection
- Mouse genes were mapped to human orthologs using
orthogene. - Cell types were annotated, and the analysis was restricted to hepatocytes.
2. Cross-species integration
- Mouse and human scRNA-seq datasets were integrated using scANVI (semi-supervised) on the top 6,000 HVGs.
- This produced a corrected counts matrix.
- Correlation and PCA analysis on raw versus corrected counts showed a broadly similar structure, supporting preservation of the biological signal.
3. Pseudobulk DE and pathway analysis
- Hepatocyte-only pseudobulk DE was performed using limma-voom, followed by GSEA.
(Hepatocytes are of particular interest to the lab as key resistance drivers and are the most easily validatable with cell culture at a later date.) - The corrected counts matrix was used. The intent here was not to claim definitive DE, but to identify candidate pathways that differ between conditions on a comparable expression scale.
4. Internal consistency/support analyses
- To test whether the identified resistance pathways showed preferential activation (and whether known drug-target pathways were suppressed), FDR-corrected Spearman correlations were computed between pathway gene signatures and pseudobulk-aggregated raw hepatocyte counts within each original dataset.
- Genes outside the 6,000 HVGs could still emerge if they showed significant correlation with the pathway signature.
- Strong negative correlations aligned with known drug-action pathways.
- GSEA on FDR-significant genes ranked by signed correlation coefficients further supported the internal coherence of the hypothesized resistance program.
5. Biological plausibility
- Key regulators of this pathway are known to be mutated specifically in the rare cancer subtype, but their downstream transcriptional effects have not been explored.
- No direct DE comparison between these cancer subtypes has been published.
- A prior microarray meta-analysis reported upregulation of a broad pathway class consistent with our findings, although it did not explicitly identify this pathway.
What I’m asking
- Is a clearly labeled, hypothesis-generating, cross-species scRNA-seq study like this publishable at all without wet-lab or clinical validation?
- Are there aspects of this approach (e.g., ortholog mapping, scANVI correction, pseudobulk DE) that reviewers are likely to reject even for a speculative paper?
- Would this be better framed as a brief report / computational hypothesis / methods-forward paper, or is the lack of validation still likely to be a hard stop?
I’d really appreciate honest, even blunt, feedback so I can decide whether to proceed or pivot while there’s still time.
0 answers
No answers yet.
Log in to answer this question.
Lets define first which journals you have in mind. MDPI-style, yes almost certainly will go anywhere. Decent journals, probably not due to lack of any validation. Problem with single-cell datasets is that you can create a lot of meaningfully-looking figures, but lots of results can be entirely nonsense. You need at least some sort of validation. What do you have in mind?
That’s fair. This was originally an undergraduate project, and the likely target would be Life (MDPI), which the lab publishes in regularly, rather than a high-bar mechanistic journal.
The intent would be to frame this strictly as hypothesis-generating. In the absence of wet-lab or well-annotated clinical data, I was hoping that internal consistency checks (e.g., pathway-level Spearman correlations in raw data) together with concordance with prior studies and hypotheses in reviews could serve as supporting evidence rather than validation.
Given those constraints, I’d be very interested in how you would approach this in my situation, or whether you think there’s a better way to structure or scope the analysis.
I hear you that you want to sell it as hypothesis-generating, but I personally would not really accept this as a reviewer as an excuse for validation of any kind. If we did, then every poorly-validated study with wild claims could use this tactic. These internal sanity checks anyway should be in the supplement to some extend, I don't see how this increases credibility.
Maybe try something like PLoS Computational Biology first, rather than MDPI. While not high impact at all, at least PLoS journals are sound with no bitter taste like MDPI.
We had it once that a similar situation occurred and relevant data could not be shared. We ended up providing them with code, genesets and a to-do list for analysis. They ran it, never sharing the actual raw / patient data, and we received results lists and plots. Would that be an option? Or alternatively, could you visit them for a week and do the analysis on-site, so there would legally be no sharing in terms of sending data out? If such analysis turns out useful it would strengthen the paper quite a lot.
Thanks, I guess I'll add PLoS Comp. Bio to the top of the preference list. As for visiting them on site it's out of the question as they are in another country. The code sending idea sounds pretty good though, I'll reach out again and see what they think. I'll let you know how that goes, thanks kind stranger!
This is all good advice. I'll add on to it to say you should try to find other collaborators that might be willing to do some validation experiments for you.
This can be an easy sell when you can come to them with your initial results/hypotheses and say "we think X, Y, and Z might be interesting, we'd like to do experiments A and B to help validate it". Often, you can find such collaborators at your institution, particularly for the relatively simple experiments you list.
Otherwise, I agree the value of a "hypothesis"-generating paper is quite low, even if you can demonstrate your findings in orthogonal datasets. Any reviewer worth their weight will request at least basic validation in an appropriate model. Then again, plenty of journals publish flat out fraudulent crap, so I am sure you could find a landing spot if you were willing to pay the APC. But it seems like you care about doing actual meaningful science, so I'd also recommend shooting for a journal where you'll get real reviews.
Thanks, this sounds very helpful! I'll propose this to my supervisor and see what she thinks and maybe why she didn't immediately propose this. Going to a different lab to do the validation myself also seems like a good way to network towards my master's as well.
I found my current position by visiting another lab and doing some experiments there for one of their projects as part of a collaboration. So yes, getting a network is great.