Thanks for your reply. Yes I suppose you're right about using the tumour sequences. I was hoping to avoid the effects of normal cell contamination in the solid tumour biopsies, but these will of course still be present if I sample CNA events from the solid tumour in any case. Do you have any recommendations for how to admix my samples at varying ratios in silico? It would especially be helpful if it is possible to choose which tumour reads to include so I can have control over exactly how many events are included in the synthetic dataset, and their sizes/locations.
On the topic of cfDNA fragment lengths, there is of course remarkable consistency in fragment lengths for non-tumour derived cfDNA (166bp), whereas tumour-derived cfDNA has been shown to have shorter fragments (and a few much longer ones, according to some publications). It would be interesting to include in the synthetic dataset if possible, since many variant callers rely on fragment lengths for ctDNA identification, but I am assuming this will be too difficult to simulate.