This is a test version of Biostars. For the public version, visit https://www.biostars.org.
Tool: An automation and orchestration tool - Need honest feedback

Hi dear community! I am coming from Systems Developing and DevSecOps role with the education of M.Sc. Chem. I'm trying to enter Bioinformatics for some time now but there are no opportunities in my country. However, as I really like what and how BI research and answers I've continued to learn, read and practice (on my own). As my entrance project I decided to create some automation and orchestration tool as I've read a lot about missing such tools enforcing bioinformaticians to work in many tools, transfer INs/OUTs etc.

I've created tool which keeps researchers in one tool providing the possibility to pipeline the OUT of one tools as IN for another tool. I have not re-invented a hot water, just used a lot tools and made them to work in a orchestrated mode.

AI-refined description of my LOCUS (https://locus.web-host.rs):

Folding (/folding) - protein structure & drug discovery Structure prediction from an amino-acid sequence via real ESMFold (own ml_service Python microservice). Mutation effect / variant impact scoring - ESM2 masked-marginal scoring of point mutations. Structural alignment/superposition - RMSD between any two of a user's predictions (BioPython Kabsch algorithm). Domain/motif annotation - real InterProScan5 (Pfam, PANTHER, PROSITE, etc.) via EBI's public REST API. Ligand binding-site prediction - fpocket (Voronoi/alpha-sphere pocket detection). Molecular docking - a user-supplied ligand (SMILES) docked with AutoDock Vina against a detected pocket, to actually score binding, not just locate a cavity. pgvector similarity search across predicted structures.

NGS Orchestrator (/ngs) - DNA sequencing pipelines Full FASTQ -> BAM -> VCF pipeline (bwa, samtools, GATK) against a real GRCh38-derived reference. Coverage/depth reports (samtools coverage). BAM QC metrics (samtools stats). Structural variant / CNV calling (CNVkit, single-sample mode). Multi-sample / cohort joint variant calling (bcftools mpileup/call across 2+ samples at once). CWL export of any completed analysis job, for portability outside LOCUS.

VariantLens (/variants) - clinical variant interpretation VCF import and annotation against ClinVar, UniProt, gnomAD (population frequency), and VEP (SIFT/PolyPhen-2 in-silico predictors). ACMG/AMP classification - a real 7-code evidence subset (PVS1, PM2, PP3, BA1, BS1, BP4, BP7, plus trio-derived PM3/PM6) combined per the actual Richards et al. 2015 rules, with ClinVar concordance flagging. Trio analysis - de novo, dominant/inherited, recessive-candidate, and compound-het Mendelian classification from a trio's genotypes. Pharmacogenomics (PGx) - CPIC-curated star-allele calling across 6 genes (CYP2C19, CYP2C9, TPMT, DPYD, SLCO1B1, VKORC1) with live PharmGKB guideline links. Tiered variant report, viewable in-app and exportable as a PDF report (now including the PGx section).

RNA-Seq (/rna) - transcript-level analysis Real transcript quantification via salmon against a canonical-transcript reference (BRCA1/BRCA2/BRAF). Differential expression between two sample groups via real DESeq2 (negative binomial GLM, with a dispersion-estimation fallback for small references). Single-cell RNA-seq run support.


I would appreciate if anyone interested in this tool can register (free, no spamming emails except the first one which activates an account) and take a look on it. Did I properly recognized the gaps in analyzing tools? Am I on right way in my intention to help researchers to save time they waste on pivoting on many tools? Is that a real issue? It is hosted on my own server (96GB RAM, not so much) as a demo. My plan is to provide a commercial license to labs and eventually free for academia.

Any feedback, advice, objection .. whatever is highly appreciated in advance!

orchestration automation

I clicked the link but I can't tell what the service does - instead of informing and showing me what it is it asks me to create an organization to log in ...

Hi Istvan! Correct, it is prepared for multitenancy so an 'organization' is what distincts the tenants.

You can write whatever you want as this is demo not, not an official one (still).

Once you create an account (and an org.) you'll get an email to activate your account (please check spam, although it is not a spam but as it is hosted at my home, mail address has no reputation). Thank you for willing to try it. I am working on it every day when get a chance.

But why would someone go through all that trouble - time and attention are valuable - when they don't even know what is on the other end,

There has to be some incentive, piquing someone's interest,

Right now, you are asking for honest feedback, basically work and some effort - but you are not making that process easy or approachable

All I have are some vague sentences - the kind I have read many times in my life, and most often these are oversold

You should have a demo account, a clickable link, or a snapshot of the reports, or something where I can get a sense of what I am getting without going through a confusing activation ... misleadingly titled "Register Organization: ...

You're right. I'll create a video/demo explaining what platform can do. I think I was announcing the tool too early (I planned a video already) as a result of kind of euphoria due to being happy and could not resist to announce my first BI work. My bad.

A video is coming as soon as I get chance to create it.

Instead of the video, I recommend the following:

  • Share the source code on github, with clear instructions how to deploy
  • Remove the login requirement for a test user or at least allow signup with OpenID

3 answers

Follow the "do one thing well rule". All that you describe already has dozens of publically available and largely or fully automated solutions, with active communties at its back to maintain it. nf-core for example, and others. You try to break into a saturated niche.

Thank you for this response. That is what I expected!

I agree with the “do one thing well” principle. LOCUS is not trying to replace nf-core or Nextflow, or similar workflow ecosystems. It focuses on one practical gap: helping small genomics labs to go from FASTQ to reproducible, clinician-friendly results with minimal DevOps burden by orchestrating proven pipelines.

While I cannot comment on the usability and tuility of your app, I'd doubt that that 'practical' gap really exists. Also the problem with these LLM statements ("reproducible, clinician-friendly results with minimal DevOps burden by orchestrating proven pipelines") is that they sound very convincing at first. It's less clear what that actually means and how the tool can make good on these promises. Also, your slogan states that the app uses "real tools" (do other pipelines use "fake/false tools"?), a typical give-away for LLM generated text.

Thank you, Michael! I stated this is AI driven description, so - yes, it's LLM. English is not my native language, but in the technology (IT at least) there are well known sintagma which describe tech aspects very good, yet sounds like cheap commercials.

reproducible - yes, clinician-friendly results - by the form; by the quality - a serious testing/fixing cycle required.

with minimal DevOps burden by orchestrating proven pipelines - exactly that, the average user creates his own pipeline without asking a tech guys for help to pipeline this to this, redirect this, copy that etc.

app uses "real tools" - yes, it relies on many community proven existing tools. Not a custom ones.

So, not the platform production ready, but an idea reviewing ready.

helping small genomics labs to go from FASTQ to reproducible, clinician-friendly results with minimal DevOps burden by orchestrating proven pipelines

That goal is laudable but there are certain limits you will need to be aware of. The issue will be with validation of results your tool is producing. Where clinical data is concerned, utmost care/rigor needs to be exercised, since the results can directly impact lives (if results are used as is). Depending on region of the world one is from, legal requirements (and testing) for software and systems that produce clinical results can be very strict. e.g. computer systems validation (CSV), CFR 21 compliance in US.

By your own admission you are coming to this

I am coming from Systems Developing and DevSecOps role with the education of M.Sc. Chem.

so I hope your website provides a disclaimer to make it clear that that the results are for "research" use only.

Thank you, valuable comment. This website (rather a platform as it is actually interconnected and nested chain of a tools with various purpose, specifically configured databases) is just a demo. The platform is meant to be installed and configured on (for example) lab hardware and therefore stay in a internal circle, and and there is no reason data to leave the physical location.

Legal part - still far away. My post was more like an ask for intention (and also an extrapolation) of my idea if it worth at all or is non sense from the POV of real bioinformaticians.

No reason to start a legal part if the idea is non sense.

The platform is meant to be installed and configured on (for example) lab hardware and therefore stay in a internal circle, and and there is no reason data to leave the physical location.

Thanks for the clarification. It may be made clear on the platform site but it is not clear in original post above. In any case, you as the developer would carry responsibility for the functionality/content, no matter where the software is installed.

Don't know if you have considered the fact that bioinformatics tools are dynamic. They keep changing and sometimes in drastic ways. You would need to be ready to provide long term support for these changes, if you expect people to use your software.

No reason to start a legal part if the idea is non sense.

It is better to be safe than sorry. You may never know if the idea is non-sense. What some consider "junk", can be a treasure for others.

You may want to take a look at the discussion in this related thread: AI for bioinformatics pipeline development. Is it a good idea?

Thank you for the benevolent advices. Yes, my bad for poor description.

And yes, I am aware of the fact that tools are live and constantly updated. My platform uses dockerised external tools (on demand) and only when needed, with strict versioning which means every platform version will work forever, and if not updated (with stored tools versions).

I think it worth to access and look in there.

From what I've seen, the problem isn't that bioinformatics lacks tools—it's that researchers constantly jump between them, each with different formats, dependencies, and interfaces. An orchestration layer can definitely reduce friction if it's reliable and reproducible. My only caution is that adoption depends less on features and more on validation, documentation, containerization, workflow portability, and integration with existing pipelines like Galaxy, Nextflow, or Snakemake. If your platform complements those rather than replacing them, labs may be more willing to try it. I'd also ask a few active researchers to test real projects and prioritize their workflow pain points over adding more features.

Thank you, Henry! Great comment!

Genuine bioinformatics depth (not just tool wrappers):

  • Real ACMG/AMP variant classification - implements the actual Richards et al. 2015 evidence-code table (PVS1, PM2, PM3, PM6, PP3, BA1, BS1, BP4, BP7), including trio-derived de novo (PM6, correctly distinguished from PP1) and compound-het evidence (PM3) from real genotype data - not a made-up severity score.
  • Real statistical differential expression - DESeq2's actual negative-binomial GLM for RNA-seq comparisons, with a real fallback
    (per-transcript dispersion) when the reference is too small to fit a trend curve, exactly as DESeq2 itself recommends.
  • Real structural bioinformatics chain - ESMFold structure prediction -> ESM2 zero-shot variant effect scoring -> Kabsch superposition/RMSD -> InterProScan domain annotation -> fpocket geometric pocket detection -> AutoDock Vina docking (search box centered on real pocket residue coordinates) -> PLIP protein-ligand interaction fingerprinting -> OpenMM molecular dynamics. Each step feeds the next with real coordinates/data, not placeholder connections.
  • Trio (Mendelian) analysis - de novo, dominant/inherited, recessive-candidate, and compound-het-candidate classification
    directly from parsed VCF genotype columns.
  • Multi-source variant annotation - ClinVar, gnomAD population frequency, Ensembl VEP (SIFT/PolyPhen), Pangolin splice-effect
    scores, PharmGKB - cross-referenced, not siloed; Pangolin scores even feed directly into the ACMG classifier's PP3/BP4 codes.

Every pipeline step records exact tool + version, every job has a real Docker-executed command trail, and results are downloadable as runnable CWL - reproducible outside LOCUS entirely, not just a report screenshot.

Hi guys!

For anyone interested, I have updated the login requirements and now it is possible to login without registration with your Google or Microsoft account. The platform will automatically create an organization for you and you're in!

The high(est) level platform organization is: Case:

  • Folding (and related)
  • NGS Orchestrator (and related)
  • VariantLens (and related)
  • RNA-seq (and related)

enter image description here

I appreciate every feedback!

I think you made some progress towards being more accessible, but still not enough.

People don't want to have to register just to see what something does.

I would recommend creating a demo account that people can log in directly, an account that has precomputed results and runs for exploration.

Then we can chat, having to fork over my Google info on a blank page ... it is a turnoff for a lot of people

Thanks Istvan!

Finally, I added a demo account, and now the following auth methods are available:

  • Manual registration
  • Google/Microsoft login
  • Demo account with already persistent data/completed analysis:
    • locus@demo.org (pass: iEgbCp4gMp0i8iQq)

Once again, url: https://locus.web-host.rs

Add a link on the page that, when clicked takes people to the demo project.

Your site has substantial usability problem, it focuses a lot on methodology, instead of results.

The number of clicks one has to make are too many. For example when one generates variants, they want to see the variants first and foremost. Show the results as soon as possible, right away in fact.

The pipeline is secondary, think about someone that runs one hundred of these, do they want to see each time the lenghty pipeline steps ...

usability is the hardest to get right

That was what I wanted to hear. Thank you, Istvan.

This is DevOps speaking out of myself :) That part of me is highly interested in pipelines, likes to see them and to has ability to tweak, reconfigure, see the progress.

I would like more people to try this platform out and see the main projection for change/inprovement. However, you're, this must be more result-centric. Will work in that direction, thank you a lot for your time and the willingness!

Log in to answer this question.