Dan has posted previously on how difficult it is for authors to get negative studies published. Perhaps this is the real reason why the STAR*ICU study took 4+ years to make it to press. I suspect that if the study was completed the exact same way that it was but found a benefit for barrier precautions, it would have appeared in press around 2009 or even earlier. Just a guess.
Mike has posted at least twice on Ben Goldacre and his blog/book called Bad Science (part 1 and part 2). Ben has a new post in the Guardian that discusses how medicine, academia and popular culture all favor positive, eye-catching and potentially spurious trial results and ignore important negative studies. His discussion centers around a paper published last year that seemed to provide evidence of precognition - you know it before it actually happens. That "positive" paper received tons of press, while a new negative study can't see the light of day. I think this sort of bias plays a large role in infection prevention research - it is so much easier to publish a positive quasi-experimental study showing a benefit than a negative study. This is why it was so great that after 4+ years of waiting the STAR*ICU study, which was a negative study, was published at the same time as the VA study, which showed a benefit. This way, we could have a rational discussion with the positive/negative evidence receiving "almost" equal billing.
Ben Goldacre "Backwards step on looking into the future" Guardian 4/23/2011
Pondering vexing issues in infection prevention and control
Showing posts with label Positive outcome bias. Show all posts
Showing posts with label Positive outcome bias. Show all posts
Monday, April 25, 2011
Tuesday, December 14, 2010
More on the "Truth Wearing Off" and my advice to epidemiologists of all ages
Andrew Gelman, a Professor of Statistics at Columbia, has a new post discussing the New Yorker article I mentioned last week. I highly recommend that you look at the the article that he wrote in American Scientist discussing the statistical challenges in estimating small effects.
Thus, when small 'underpowered' studies actually find an effect, it has to be a very large effect to reach statistical significance. So, small studies report overestimates of the effect.
One thing we know about before-after, quasi-experimental studies, which are commonly used in assessing infection prevention interventions, is that they are underpowered and over-estimate the effect compared to randomized trials. Power is derived from sample size, effect size AND study design, among other things.
QE studies in our field also suffer from publication bias since many have been completed by clinicians who won't go through the trouble of reporting negative studies. How many papers have you read in ICHE/AJIC/CID that mentioned ADI for MRSA not working? Even if ADI for MRSA is the greatest control measure ever, which it might be, given a normal distribution of benefit, you would expect some studies to be negative, would you not? Where are they?
Even, when negative studies do appear (e.g. Harbarth JAMA 2008 or Charlie Huskins hopefully soon to be published STAR-ICU trial) they are often not believed or even thought to be flawed! Why? Nothing works 100% of the time and a negative study is NOT an erroneous result. A negative study is certainly not prima facie evidence of a flawed study. You must use all of the data, assess it based on quality and power and look for publication bias. (this is my advice to epidemiologists of all ages)
So, are we over-estimating the benefits of ADI and other interventions used in infection prevention?
Link: Gelman's post
Gelman and Weakliem American Scientist, 2009 (PDF)
My favorite passage: "Statistical power refers to the probability that a study will find a statistically significant effect if one is actually present. For a given true effect size, studies with larger samples have more power. As we have discussed here, “underpowered” studies are unlikely to reach statistical significance and, perhaps more importantly, they drastically overestimate effect size estimates. Simply put, the noise is stronger than the signal."
Thus, when small 'underpowered' studies actually find an effect, it has to be a very large effect to reach statistical significance. So, small studies report overestimates of the effect.
One thing we know about before-after, quasi-experimental studies, which are commonly used in assessing infection prevention interventions, is that they are underpowered and over-estimate the effect compared to randomized trials. Power is derived from sample size, effect size AND study design, among other things.
QE studies in our field also suffer from publication bias since many have been completed by clinicians who won't go through the trouble of reporting negative studies. How many papers have you read in ICHE/AJIC/CID that mentioned ADI for MRSA not working? Even if ADI for MRSA is the greatest control measure ever, which it might be, given a normal distribution of benefit, you would expect some studies to be negative, would you not? Where are they?
Even, when negative studies do appear (e.g. Harbarth JAMA 2008 or Charlie Huskins hopefully soon to be published STAR-ICU trial) they are often not believed or even thought to be flawed! Why? Nothing works 100% of the time and a negative study is NOT an erroneous result. A negative study is certainly not prima facie evidence of a flawed study. You must use all of the data, assess it based on quality and power and look for publication bias. (this is my advice to epidemiologists of all ages)
So, are we over-estimating the benefits of ADI and other interventions used in infection prevention?
Link: Gelman's post
Gelman and Weakliem American Scientist, 2009 (PDF)
Tuesday, December 7, 2010
Where did all of the significant findings go?
There is a really interesting piece in the New Yorker (Dec 13, 2010). Worth tracking down a copy at your neighbors or dentist's office since the free online version is limited to the abstract. Jonah Lerher describes in The Truth Wears Off, that initial studies often report large benefits from treatments or large associations between a disease and a specific risk factor which then can't be validated in future studies. There are many potential reasons for this 'decline effect' including publication bias - only publishing positive findings, especially in major or high-impact journals. Dan had a nice post discussing positive outcome bias a couple weeks ago. Another issue might be selective reporting of results by investigators desperate to find strong associations so that they can get published and then get re-funded. Certainly regression to the mean is important - ye olde bell-shaped curve. One thing they don't mention is confirmation bias, which I think drives both NIH funding and publication decisions and could be responsible for some of the reduced effect sizes seen in later vs. earlier publications.
This 'decline effect' is troubling given what it says about the scientific process. One wonders if changes in how science is funded and reported could impact this?
Thursday, November 18, 2010
Accentuate the positive, eliminate the negative….
It sounds like a good idea when Mr. Bing Crosby croons it, but it’s not good for science. Negative studies are every bit as important as positive ones, but they are much less likely to be published. There are many reasons for this, but one of them (positive outcome bias by manuscript reviewers) is carefully examined in an interesting study found in the November 22 issue of the Archives of Internal Medicine (Emerson GB, et al. Arch Intern Med 2010;170:1934).
The investigators fabricated two versions of a manuscript (a “positive” version and a “no-difference” version), purposefully placing minor errors in each. They randomly sent one version or the other to over 200 peer reviewers. Their findings? Not only was the positive version more likely to be recommended for publication, but reviewers were less likely to detect the errors in the positive version, and more likely to score the methods section of the positive version higher (despite the fact that the methods sections were identical).
So listen up, authors, reviewers and journal editors: stop accentuating the positive and eliminating the negative! Submit those high-quality negative studies, review them fairly, and get them out there for all to see.
One thing Bing Crosby and I can agree on, though...you still shouldn’t mess with Mr. In-Between.
The investigators fabricated two versions of a manuscript (a “positive” version and a “no-difference” version), purposefully placing minor errors in each. They randomly sent one version or the other to over 200 peer reviewers. Their findings? Not only was the positive version more likely to be recommended for publication, but reviewers were less likely to detect the errors in the positive version, and more likely to score the methods section of the positive version higher (despite the fact that the methods sections were identical).
So listen up, authors, reviewers and journal editors: stop accentuating the positive and eliminating the negative! Submit those high-quality negative studies, review them fairly, and get them out there for all to see.
One thing Bing Crosby and I can agree on, though...you still shouldn’t mess with Mr. In-Between.
Subscribe to:
Posts (Atom)
OSHA! OSHA! OSHA!
In many parts of the country, as rates of COVID-19 are declining and vaccination coverage is increasing (albeit with substantial variati...
-
Back on clinical service again and having more thoughts on poor hospital design. Last month I wondered why there were no stethoscope wipe...
-
Those that follow me on twitter or the blog have probably noticed my recent focus on trying to understand the emergence of compulsory influe...
-
REUTERS/Athit Perawongmetha With the assistance of a great supply management team, we have been able to outfit all of our clinical staff wit...
