I have multiple data that in all of them the first 3rd columns are the same and I like to merge those data based on these three columns. I can do it with merge command in R but I like to do it in Linux (I used joint command but it does not work well).
data1:
chr1 724060 724400 SK chr1 724206 725561 peak_1 24 . 2 194
chr1 729399 731900 sun . -1 -1 . . . . 0
data2:
chr1 724060 724400 sk . -1 -1 . . . . 0
chr1 729399 731900 sun chr1 724206 725561 peak_10 24 . 5 104
output:
chr1 724060 724400 SK chr1 724206 725561 peak_1 24 . 2 194 . -1 -1 . . . . 0
chr1 729399 731900 sun . -1 -1 . . . . 0 chr1 724206 725561 peak_10 24 . 5 104
1 answer
I'm not sure I fully understand the question, but I'm not thinking mega clearly. Is this the desired outcome?
I'm confused by your desired output because I don't see the string "RA" anywhere in the input files.
1.
First I had to manipulate the whitespace in your example data so that it was properly tabulated.:
perl -p -e 's/ +/\t/g' file1.txt > file1.tsv
# and the same for file2
2.
If your data is already correctly ordered top-to-bottom (you haven't stated), then this will work I think:
paste file1.tsv <(cat file2.tsv | cut -d$'\t' -f 4-)
which yields:
$ paste file1.tsv <(cat file2.tsv | cut -d$'\t' -f 4-)
chr1 724060 724400 SK chr1 724206 725561 peak_1 24 . 2 194 sk . -1 -1 . . . . 0
chr1 729399 731900 sun . -1 -1 . . . . 0 sun chr1 724206 725561 peak_10 24 . 5 104
Log in to answer this question.
Include exact
joincommand you have tried.Just I did , join data1.bed data2.bed > data.bed
For
jointo work, the files must be sorted (in the same order), and you have to tell join whichfieldyou want it to do the joining by.