You're correct, that all Ensembl transcripts go back to a cDNA or protein (and no XP or XMs). The 'known' and 'novel' status is actually determined after the genebuild. Once the Ensembl transcript set has been "built", the transcripts are compared against scientific, public databases. If there is a sequence match to a protein or cDNA for the same species, the transcript is classified as "known". If there is no match, the transcript is classified as "novel". We also have a "known by projection" classification, which are Ensembl transcripts with a sequence match for another species. This classification is more common in species where not much cDNA or protein is available- a homology build had to occur.
I'm not sure where you are seeing "predictions". These might be our Genscan predictions. The genebuilders start with genscan predictions, which overpredict, in their initial alignments of cDNA and proteins to the genome. The genscan predictions are not used as supporting evidence, and do not lead to transcripts in the Ensembl gene sets. They are only there to support the annotation pipeline:
http://www.ensembl.org/info/docs/genebuild/genome_annotation.html
Where exactly in the database did you find them (can you show me your query)?
This type of question is great for our helpdesk (helpdesk[?]ensembl.org). Feel free to continue this discussion there, if you wish to send your script or query along that would help.