Thanks a lot for your comprehensive explanation. My concern is that I have to compare two statistical models for my thesis and since I don't have the house keeping genes, I can't decide which model is working better. That's why I suggested using a benchmark method as a standard to compare the other two models. If I could simulate data, that would be easy to compare the models, but I don't know how to simulate data.
A good approach would be to test the data with more than 2 algorithms and either take the intersection of results of the N algorithms used or the union of the intersections between at least 2 softwares
If I understand you well, you said I can find the common DE genes in two or more benchmark methods and then use those genes as the index genes and check how many of them are detected by each one of my models, right? It seems good.