jiz this is very anthropocentric .. ;)
Hi everyone,
Whenever I teach an introductory class about Bioinformatics, I like to use this word cloud - I feel it gives a quick glimpse of the field. However, it is outdated (from 2011).
So I want to create an up-to-date figure. I understand I can manually copy the tag counts from the tag content and use only the first few pages, because the word counts quickly drop below 15 or so.
However, ideally I would like to have the data for each year separately, to show how topics change over time. Unfortunately, I am a new user (first post here) and do not have the privileges to download the database as shown in the blog post above.
Would anyone with access be able / willing to fetch and share these data?
Thank you very much!
2 answers
instead of using biostar, use pubmed an the mesh terms. Here is an example using my tools http://lindenb.github.io/jvarkit/PubmedDump.html and http://lindenb.github.io/jvarkit/XsltStream.html
Pierre, thanks for pointing this out!
Please, see my reply above about wanting to reflect community questions.
The blog post What is bioinformatics about? also has a recipe for creating word clouds from abstracts. At this moment, the blog is down to me, though.
What tool creates the cloud itself? It has cool looking styling, what parameters does it need to make it look like that? Now that I have played a bit with word clouds I think that figuring out the right styling is a separate challenge onto its own.
after googling: https://wordart.com/
Love the standalone "High".
looks like R is missing, second most common tag, probably because it is one letter long.
Another word cloud, this time based on the words in the 1000 most highly voted post titles:
and using the https://wordart.com service.

I used these data and made a figure to look like the one from the original post. RNA-seq pretty much overwhelms everything else.

Same data with different scaling.

Log in to answer this question.

cc Istvan Albert , Devon Ryan and Pierre Lindenbaum
Nice application! Tags are very often used incorrectly though. An alternative approach could be to use the title/abstract of recent (bioinformatics) papers? Although filtering those terms obviously requires some more work to get the bioinformatics terminology out.
Well, using paper title/abstracts would certainly be useful, although in a different way.
For my teaching purposes, I actually want to use community-based information, in the sense that it reflects user needs. I expect my students to face many of the questions that other users experience. For instance, note that the tag software error has been used 1838 times - I would then address the value of resources such as Biostars. It also shows trends in language use (e.g., noticeable drop in perl, increasing importance of python and especially R).
Some noise due to incorrect tag usage should not be a big deal. In any case, I will filter for only the most used ones, so the signal will still be there.
Thanks for the input!
Gha, funny that you pick the example of
software error. Because it is used so frequently it will become the first suggestion users get when making a new post, leading it to be used more and more. Often it's actually a user error :)Even better!
Nice to know this, because in classes students show similar behavior. They often jump the gun and call anything a software error, when it is usually just a typo. This will make a good example!
Here is a list of tags with counts:
http://data.biostarhandbook.com/data/biostar-tags.txt