To take the best from a CAGE experiment, I recommend to call peaks from the 5′ ends of the aligned CAGE reads (for instance with paraclu, and run the differential expression analysis at the peak level. Alternatively, one can just intersect the 5′ ends with promoter regions, for instance using the FANTOM5 peaks, or a region flanking the start site of GENCODE, RefSeq or the FANTOM Cat. More complicated approaches also exist, for instance RECLU.
I would not compare directly a CAGE library to a RNA-seq library: they have different purposes, and the most appropriate tools to process them differ. However, if you have a full CAGE dataset matched with a full RNA-seq dataset, you can compare the results of each analysis. In many cases, I would expect them to cross-validate, for instance when « gene A is induced by treatment T ». But you can also see a differential promoter analysis with CAGE that is not reflected in RNA-seq, or a differential splicing highlighted with RNA-seq but not reflected at the promoter level with CAGE...
As far as I know there is no influence of the transcript length on the read counts in CAGE-Seq, so therefore FPKM normalization would be inappropriate.