This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Troubleshooting High cis/tran ratio in Hi-C

Hello everyone,

I recently prepared several Arima Hi‑C libraries for a small plant genome (~250 Mb) and ran an initial ~20 Gb sequencing test. After processing with HiCUP (Arima mode), I obtained unexpectedly high cis/trans ratios:

cis/trans = 3.71–5.71

Valid pairs: 67–70%

Mapped: 25–28%

Deduplicated unique pairs: 71–79%

These values are much higher than what I’ve seen in my previous plant Hi‑C datasets (typically cis/trans ~0.9–1.4), so I’m trying to understand whether this is normal for Arima libraries on compact genomes, or whether it could indicate an issue during library preparation.

I am wondering if these libraries will affect negatively long‑range or inter‑chromosomal interaction resolution?

Regards,

sequencing hi-c

1 answer

Some thoughts come to mind:

On their own, those cis/trans values don’t sound like problems. Assuming the library prep was not dominated by short-range cis interactions (e.g., <= 1 kb), those could be considered excellent numbers (e.g., at least in my experience working with mammalian model organisms, human cell lines, and yeast).

To really interpret/dig into that, it’d help to know the numbers and proportions for longer-range cis interactions, e.g., within 1–10 kb (or 1–20 kb) and >10 kb (or >20 kb). Also, what do the cis/trans ratios look like when excluding very close cis interactions (<= 1 kb)?

It’s good to check those things b/c short-range cis pairs can be inflated by, e.g., incomplete digestion or unligated/poorly processed fragments.

What also stands out to me is the mapped fraction (only 25–28%), which seems a bit low to me. That would make me want to look at (at least some) the following:

  • reference/assembly
  • contaminant content (if any)
  • read trimming (probably fine though)
  • alignment/junction handling (it’s been a long time since I used HiCUP, so maybe it’s been updated, but the field has come a long way since the “style” of alignment and junction handling in the early implementations of HiCUP)
  • species mismatch (perhaps a stretch though)
  • mitochondrial and/or chloroplast (or other organellar) DNA
  • whether HiCUP settings were appropriate for the plant genome (kind of tied to the above)

Also, when I see the initial cis/trans ratios (around 1), those strike me as low. But again, this is a different organism versus what I have the most experience with.

Generally, from the bench side, artificially elevated trans signal can arise when random inter-molecular ligation events aren’t controlled/suppressed during library prep. Indeed, this was often more of an issue in earlier iterations of Hi-C bench protocols, where the chemistry and protocol design were generally less effective at minimizing such artifactual trans interactions. Arima is at least intended to reduce/suppress these issues, both in terms of the empirical data they and many labs report, and in my anecdotal experience running the Arima benchwork. Did you also use Arima for your initial sample preps?

Log in to answer this question.