Showing posts with label data mining. Show all posts
Showing posts with label data mining. Show all posts

5 Sept 2010

What is the scientific paper? 4: Access

This is a guest post by Joe Dunckley
Completing the series exploring the question "what is the scientific paper?", reposted from my old blog, and originally written following Science Online 2009. As I reminded people at the time, these were just my own half-thought through ideas, not the policy or manifesto of anyone or anything I'm affiliated with.
A friend of mine once told me how much she hated "the proliferation of these bioinformatics papers." All these simulations and models of what happens in real life. All of it utterly useless -- since when was the stuff that comes out of a computer worth anything? None of it even remotely reflects anything that happens in real life. And the methodology papers -- the endless methodology papers. They're making yet another neural network and modifying a bayesian something-or-other, when they haven't even found where they left the markov models yet! How can you have so many of these methodology papers? Clearly they can be no more than incremental advances. (Of course, BLAST is an exception -- it's old enough to have been around and heard of when we were undergrads, and is therefore a perfectly legitimate and mainstream molecular biology tool.)
Similarly, some people still voice their skepticism about the need for open access. Access isn't really a problem, is it? These open access advocates are just making facile arguments about the how the people who pay for scientific research should have some kind of say regarding its dissemination.[1] Come on, really, show me, who is in want of access? Everyone (everyone who matters) already has subscriptions, right? Access isn't a problem. And the open access "movement" isn't an ideology. It's just another business model.
And then, yesterday afternoon m'colleague shouted for advice handling an author of a scientific manuscript who was questioning the need to deposit her not inextensive collection of genomes in a database. I don't blame the author for wanting to get out of the chore—she had a lot of data, and depositing it will be a dull repetitive task. M'colleage was trying to write a letter and struggling to put into words the reason why we mandate deposition of sequence data, and why merely including them as supplementary MS Word files isn't good enough.
These attitudes, you will have noticed, have one particular thing in common: they all completely miss the fact that the biomedical sciences have moved on in the past quarter century. In almost every field (lets not wake the poor taxonomists) the science being done and the science being published today are not quite like that of 25 years ago. Even if the science of today were like that of 25 years ago the case for open data sharing would be strong enough; as it is, it's simply absurd to think that open sharing of data isn't worth doing.
--
Individual scientific papers -- the basic units of scientific research -- are rarely exciting; rarely even interesting. Where nerds get excited about science, it's where science offers a beautiful explanation for how the world works. And scientific papers don't do that. They offer some speculative interpretations of data on obscure problems in obscure systems. It is the literature as a whole -- hundreds of dull papers put together -- which tells a complete and exciting story. The sum is more than the parts -- the theory is more than the data.
In the field I know best -- cancer cell biology -- 99 in 100 papers published are tedious details, discovered with a science-by-numbers formula. The (anti-)proliferative effect of one abbreviation interacting with another abbreviation in three-letter-acronym-and-a-number cells, concluding with a suggestion that the authors' work might have implications for cancer treatment and a note that further work is necessary. Or even better, the complete lack of anything interesting at all happening when the first abbreviation interacts with the second. The abbreviations and their effects have been studied, in combination with others, in all of the most widely used three-letter-acronym-and-a-number cell-types, and somebody is scraping the barrel.
But the tedious details put together add up to an understanding of how the cell works and how it goes wrong. The details could be put together by a human, going through the thousands of papers on the topic, assembling the facts and finding the trends. Or, more plausibly, given the amount of tedious details out there, they could be assembled by a computer, with a database and a clever algorithm. Except that four in every five of those tedious details, discovered at great expense to taxpayers, will be inaccessible to that clever algorithm. They will be locked away in the basements of university libraries, hidden in human-readable prose that humans will never read. The results of billions of pounds of work searching for an understanding of cancer and a better chance at defeating it will be worthless, because they will never be amongst the parts that add up to the greater whole.
So I told m'colleague to explain to her author that unless she deposits her genome sequences, the last three years of her professional life will ultimately have been wasted. An average paper in a high-volume mid-tier journal that will be glanced at by a few colleagues when published. Another bullet point on a CV. They will never further science beyond that. They won't contribute any important discovery or real advance to the field. They will be forgotten. Nobody will seek them out when the time comes to make the leap forward.
That's just where biology is at these days: lots of tiny fragments of data, spread thin through the literature. The most interesting and important unanswered questions will require the synthesis of that work. The most interesting and important questions can't be answered without the heap of data that has already been produced, but which is locked away.
On machine readable data, Mike Ellis says, "at some point in the future, you'll want to do "something else" with your content. Right now you have no idea whatsoever what that something else might be." This is especially true in science: at some point in the future, tedious data obtained at great expensive, as part of the bigger picture, will finally be important and valuable. Right now, you can have no idea how important.
Publishers are allowed to get away with keeping science closed, holding it back, and wasting public money because there are still sufficient numbers of scientists who let them -- who have themselves failed to grasp that the world and science have changed.

8 Sept 2007

Multiple Stab Wounds May Be Harmful To Monkeys

Multiple Stab Wounds May Be Harmful To Monkeys. Repeatedly stabbing monkeys with sharpened objects may have an adverse effect on their health, according to a new study.

In other news, four CIA agents are trapped in a dating mining disaster.



30 Jan 2007

Tools to search the literature, and PubReMiner plugin

Recently I came across PubMed PubReMiner, created by Jan Koster. I've been very struck by this tool, which I think is pretty much the best way to search PubMed.

I have previously tried a number of different tools (see my list of Tools to search the literature in the sidebar), and Google Scholar by far outstrips a standard PubMed search due to the use of the PageRank algorithm to pull the most prestigious work to the top of the results. The PageRank algorithm doesn't just look at citations, rather it weights them by how often that referring articles has itself been cited. A citation from a source that is itself heavily cited counts more than one from a source that nobody has ever cited.

I've tried out Kfinder, which takes an abstract or other text as the input, and suggests keywords based on the frequency of occurrence of improbable words. You select keywords, and it returns researchers who match that search in Medline at least twice. Kfinder is quite slow and limited to Medline, but it is intuitive, and a good start to selecting keywords if you haven't had much practice.

PubNet from the Gerstein lab looks really promising. It visualizes the network resulting from a query to Medline. The network to the left, focussed around Howard Ochman and Emmanuel Lerat, clearly shows a network of collaborating colleagues, but I only ran that search because I knew of the network already. It could be useful, but I've not found the time to devote to exploring its possibilities, and it takes a while to generate the visualization at times. If it were quicker and easier to navigate the results, I might use it.

I've only had a quick play with Authoratory, and while the concept is excellent (automatically mining information from the results of PubMed searches), the delivery is lacking. When Deborah Saltman, our Editorial Director for Medicine, tried it she found that she was missing, and the keyword search doesn't take Boolean searches yet. A definite work-in-progress.

e-Biosci is clever in that it accepts any text as input (an abstract, or even a whole manuscript, although it was quite sluggish!) and calculates the concepts contained within. You can add and remove concepts to refine your search, and weight how important they are, and then search using these concepts in Medline abstracts and some full text, including BioMed Central's. The advantage of this approach is that you never need to think about appropriate keywords or search terms; the disadvantage is that some concepts are quite diverse. A good example is that an abstract about physician uncertainty in medical decision-making returned some physics articles near the top! I find that it can return items that you probably wouldn't have found otherwise, and can be very accurate at times.

eTBLAST is one of the big hitters in the field. It runs searches against Medline automatically when given an input of text, much as e-Biosci does, and returns a list of related articles. You can then get list of experts in the field, journals to submit to, the history of publishing in this field and several more features. eTBLAST does all the thinking for you, but it does take its precious time. It can take minutes for the results to be returned, which makes me think that the option to have the results emailed is the only way it will get routinely used.

But, as I said at the top, PubMed PubReMiner is my current favourite. Why? Well, it takes standard PubMed queries, which makes it very easy to start using. It is quick and unfussy, and returns the results in easy-to-read columns: a list of the most common journals in the results, a list of the authors who appear most often, and a list of words that most commonly appear in the abstracts, as well as MeSH terms, affiliations and the publications by year. It is simple, but highly effective.

I liked it so much, that I made a Firefox search plugin for it. After vainly following a tutorial, I found that searchplugins.net has a plugins generator, which I've used to create one for PubReMiner, complete with a logo. It is set to the default of a 1000 abstract limit. You can view the source code, and search for it under PubMed or PubReMiner. You can also install it now.

28 Jan 2007

Mashups, mirrors, mining and open access

The Creative Commons Attribution License under which open access articles are made available by both BioMed Central and PLoS allows others to create sites that incorporate the content of these articles, so long as the original source is clearly acknowledged.

Two ways to do this are mashups and mirrors. According to Wikipedia, a mashup is a site that "combines content from more than one source into an integrated experience". A mirror is an exact copy of a website.

BioMed Central officially has four mirrors to which we feed content, at INIST in France, University of Potsdam in Germany, PubMed Central at the NIH, and the National Library of the Netherlands' e-Depot. I've come across some unofficial mirrors in specific areas like genomics and bioinformatics in the past.

PLoS ONE already has its own unofficial mirror, created by the people behind HubMed: PLoS Too. Rather than displaying the articles as they appear on the publisher's site, this is a pared-down view of the articles and it has a couple of good new features - auto-generated tags for each article, and a very quick live search box.

On the mashup side, Free Biomedical Images has made open access images available in a searchable database, mainly (entirely?) taken from BioMed Central articles, and fully attributed. Users can comment on the images, rate them, email them to a friend and jump to the published article.

A key feature of open access is that we don't hide away the full text of our articles. The entire 'corpus' of our open access research articles is available on our data mining page for anyone to download. Gerry Rubin has said that "the most important reason for Open Access is data mining".

The idea of mashups, scripts and extensions is just beginning to reach the bioinformatics community. A bioinformatics mashup by Alfonso Valencia is iHOP (Information Hyperlinked over Proteins), which links information about genes and proteins to text from PubMed. Not satisfied with just a mashup, Mark Wilkinson has created a Greasemonkey userscript called iHOPerator that enhances the iHOP website with tag clouds. You can read about in his BMC Bioinformatics article. Two other Greasemonkey userscripts link PubMed to social bookmarking sites, one to CiteULike, the other to Connotea. A third links Google Scholar to CiteULike. The iSpecies search engine pulls together information about any species you enter from disparate sources, including scores of biomedical databases and even Yahoo! Image search.

Mashups, mirrors and mining are definitely the future of science publishing.