Thanks for your comments, The assumtion is a large portion of the genes are uniformely expressed, after scaling you can use your prefer method for DEG detection.
Regarding your second question, my problem with spike-ins it's the few sequences that you test, but perhaps we can try to see the methods perfomance with such data sets.
And yes, there are previous publications with critics to house-keeping genes, some going further saying that there are not such genes.