I think the important thing is to have realistic expectations. Lifestyle choices- what you eat, how you sleep, your happiness have incomparably higher impact on chronic diseases than just about any data-oriented approach.
This is not to discourage you from studying bioinformatics, just to understand what bioinformatics actually is.
The way life works, even for a simple bacterium, is still immensely mysterious.
Once you get to multicellular organisms, it becomes mind-boggling. How can a fly process all that information and orient itself that quickly? How can it produce all that energy to keep itself off the ground? Now take a microscope and look closer at that fly; note the immense detail and sophistication of how it is all put together. Imagine it all came from a single cell, and billions of other flies were and are being produced the same way every year.
A larger mammal is made up of trillions of cells, each dividing and replacing each other along the way, but all came from a single cell. How is that possible? In ten days of division of a human embryo, you can tell which way the head will be.
Long story short, we don't even understand the basic modes of operation of the simplest organisms.
To start in bioinformatics start simple, get data for simple, seemingly trivial organisms. Study the data, try to dive deeper. You'll understand the real challenges much sooner that way than diving into a complex field like cancer research where the science can also be distorted by intense competition for funding, recognition, and prestige.
If you really want to contribute to actual biological research then I don't see how that would go without being actively part of a lab. The idea of home-based research is delusional (despite I applaud your motivation) as none of what you bring to the table is fundamentally new. Bioinformatics, computational biology, systems biology, or whatever label you give it is actively applied to cancer (any pretty much any) research on a daily basis. In 2026 more than ever it is not about the new fancy method, but about asking the right question with the right data, and expanding on the decades of existing research.
In particular:
1) You can dive into the many tutorials towards OMICS analysis to get the lingo and conventions of the field going. You can practice by replicating existing analysis, to get technical experience. But replication is "simple" (not easy but simple like non-complex) because the paper already tells you the answer to the hard part: The biological interpretation (if the paper bothered to have any). Beyond that, consider internships in a reputable lab.
2) I don't think "quickly" applies to anything in life science if done properly.
3) I don't think actual research in terms of contributing to biological inference can really be open source. Things like preprocessing and pipeline stitching very well can be -- but that by itself doesn't find new biology, despite having an indirect impact of course.
4) Becoming a scientist in a lab with a relevant project. I don't think there are shortcuts that would produce any meaningful insights. After all, the generations of scientists working on cancer for the last decades are all not lazy or incompetent, it's just a highly complex heterogeneous topic that advances very slowly. In cases such as AML, even with all these modern OMICS-assisted projects, if you really break it down we have not really madesubstantial treatment improvement in decades. It's still largely chemotherapy and stem cell transplantation. That just as an example.
That said, I applaud your motivation, but advise to not overestimate your potential impact, to avoid disappoitments.
To follow up on this, the slightly cynical voice in my head says that a “lonely” bioinformatician’s greatest contribution to cancer research might not be discovering something new, but simply establishing a baseline for what percentage of the already published cancer bioinformatics is irreproducible, ad hoc p-hacking.
Istvan Albert, Thank you very much for reviewing my initial post.
ATpoint, Thank you for reply on my letter) I didn't mean, by and large, solo work, of course) (With the exception of the idea put forward by Istvan Albert with the verification of p-hacking, which I will now think about). In my letter, I was thinking about how to get started, what basic steps to take to join the lab within a reasonable period of time (e.g. 1 year) and not turn it into a process of endless learning for the sake of learning.
In my experience you need a lab to get real experience on real world data, any home exercise is only good for the technical basics. You need to get in touch with the lingo and the way a life scientist thinks. THe p-hacking think is something I recommend against, because there won't imo be much output other than "you're all wrong" with no consequence. I guess internships are what you need. We collaborated with mathamaticians and hardcode CS people before, but they were often lost in the analysis because either they did not understand what to analyze, or how to analyze it in a biological meaningful way. That's way I keep saying that detail knowledge of CS and stats is at best a basic for a hands-on scientist who looks at data. Far more important is knowing the field, the system/biology, the literature and acquiring a feel how to put things into a coherent narrative. This you won't get from home.
I don't fully agree with the above. While it may be true that you can't do "classical" bioinformatics at home and alone, I think outstanding bioinformatics value can come from doing the work NOBODY is currently doing. Namely, figuring out which papers are complete nonsense and which ones are not.
In my opinion, what the field would be greatly enriched by are coherent, simple, clear explanations of why data in some influential papers are correct and which ones are incorrect or wrong, and what is wrong with them. Doing that job is probably much harder to do than the original analysis itself. Proving something wrong is a far more difficult. Not to mention the uphill battles that go with it:
For example:
Notably, Salzberg identified fundamental flaws in the original study as early as 2020. Yet it took four years for his critique to be published and for Nature to retract the study. His analysis showed that two major errors at the core of the original methodology were necessary for the conclusions. Even so, it took a long time and a hugely influential individual to get the retraction through.