This is a test version of Biostars. For the public version, visit https://www.biostars.org.
What should I consider as FASTA for dataset?

I am trying to make a dataset from rcsb for my research paper.

I have seen many pdbs have less residues in a chain of a protein than the full FASTA sequence. Most likely, the cause is that they were unmodeled due to its going missing during the crystallization phase. For example, 1ZM1. Here in chain B, the last couple of residues were unmodeled.

Should I use the shortened FASTA from pdb or should I use the full FASTA for my dataset?

pdb fasta

Should I use the shortened FASTA from pdb or should I use the full FASTA for my dataset?

What analysis are you trying to use the data for?

Protein attribute prediction

0 answers

No answers yet.

Log in to answer this question.