Hi Kenosis and thanks for your reply. I test your solution and it works just fine for the example that I porivide. But unfortunatly when I tried will a different exemple :
chr01 8947 15
chr01 8972 10
chr01 9075 1
chr01 9101 2
chr01 9172 12
chr01 9505 1
chr01 11554 1
chr01 11751 1
chr01 11947 1
chr01 12174 1
chr01 12350 1
chr01 12415 2
chr01 12504 1
chr01 12606 1
chr01 12792 1
chr01 12898 1
chr01 13081 1
chr01 13355 3
chr01 13381 3
chr01 13415 1
chr01 13444 1
chr01 13585 18
chr01 13610 1
The result is :
Lowest, most frequent position : 13585
Highest, most frequent position: 13585
instead of :
Lowest, most frequent position : 8947
Highest, most frequent position: 13585
Please do you have an idea about that ? Tnaks again
You should clarify what you mean by most frequent lowest and highest values. It looks like 1 should be the lowest occurrence value and 19 should be your highest?
I don't understand what are the rules for getting 8848 and 11761. Could you explain why it's not 8755 or 11776? If this is because you take into account occurrence, why it's not 9198 and 9268 ? Thanks for clarification.
Sorry, maybe I was not very clair about that. I would like to select 8848 because even if it's less frequent than 9198 and 9268 it's smallest and most frequent than 8755 (which is the min value). The same logic will be true zith 11761. Please let me know if it's still not clear. Thanks guys
I still don't know what the criteria are for choosing these positions. Why is 8755 the minimum value? You mean minimum base position?
Thanks Damian to take time for this ! the selection must verify tow important criteria the first one is the poisition and the scond is the occuerence of this position. It this case the minimum value (column 2) is 8755. But it's not the most frequent ! So I would like to select 8848 because it's the smallest AND most frequent values in my data. And this will be the same for the bigest values. Hope that I'm more clear.
You should rethink these criteria. Your criteria of smallest by base position and most frequent by occurrence does not allow you to distinguish between 8848 and 9198 unless you weigh one criteria over the other somehow. And if you do choose to weigh position over occurrence, you need to justify why you are doing that.
Like Damian said, you should give some weight. It seems your acceptable range for high or low value in 1000 bp, is that right. and like manu said, why it's not 8755 or 11776? why it's not 9198 and 9268 ? what is that black magic. BTW, do you use R, R may be the best way to do these things than scripting languages.
I have a csv file containing three fields :
I want to filter scaffold with high frequency in respective chromosome in new filtered file..
Any suggestions?
I think you should post it as a separate question & show some data for people to understand what you are asking for.