I'm trying to develop a predominantly automatic annotation tool. The biological interpretation of sequences can be influenced by a variety of factors. Factors which can be explored in more detail manually. I was hoping for a generic approach that could fit most/all sequences without special treatment requirements. I felt that options 2 and 3 were cop outs, but I don't want to potentially mislead anyone that might use the data. Recategorizing questionable data should fix this problem.
Edit1 - I'll also look more into how questionable data might be stored within Chado (a biological database schema I'm using).