Showing posts with label intervention. Show all posts
Showing posts with label intervention. Show all posts

Thursday, March 21, 2013

Blogging as post-publication peer review: reasonable or unfair?








In
a previous blogpost, I criticised a recent paper claiming that playing action
video games improved reading in dyslexics. In a series of comments below the
blogpost, two of the authors, Andrea Facoetti and Simone Gori, have responded
to my criticisms. I thank them for taking the trouble to spell out their views
and giving readers the opportunity to see another point of view. I am, however,
not persuaded by their arguments, which make two main points. First, that their
study was not methodologically weak and so Current Biology was right to publish
it, and second, that it is unfair, and indeed unethical, to criticise a
scientific paper in a blog, rather than through the regular scientific
channels.


Regarding the study
methodology, as noted above, the principal problem with the study by
Franceschini et al was that it was underpowered, with just 10 participants per
group.  The authors reply with an
argument ad populum, i.e. many other studies have used equally small samples.
This is undoubtedly true, but it doesn’t make it right. They dismiss the paper
I cited by Christley (2010) on the grounds that it was published in a low
impact journal. But the serious drawbacks of underpowered studies have been
known about for years, and written about in high- as well as low-impact
journals (see references below).


The response by Facoetti
and Gori illustrates the problem I had highlighted. In effect, they are saying
that we should believe their result because it appeared in a high-impact
journal, and now that it is published, the onus must be on other people to
demonstrate that it is wrong. I can appreciate that it must be deeply
irritating for them to have me expressing doubt about the replicability of
their result, given that their paper passed peer review in a major journal and
the results reach conventional levels of statistical significance. But in the
field of clinical trials, the non-replicability of large initial effects from
small trials has been demonstrated on numerous occasions, using empirical data
- see in particular the work of Ioannidis, referenced below. The reasons for
this ‘winner’s curse’ have been much discussed, but its reality is not in
doubt. This is why I maintain that the paper would not have been published if
it had been reviewed by scientists who had expertise in clinical trials
methodology. They would have demanded more evidence than this.


The response by the
authors highlights another issue: now that the paper has been published, the
expectation is that anyone who has doubts, such as me, should be responsible
for checking the veracity of the findings. As we say in Britain, I should put
up or shut up. Indeed, I could try to get a research grant to do a further
study. However, I would probably not be allowed by my local ethics committee to
do one on such a small sample and it might take a year or so to do, and would
distract me from my other research. Given that I have reservations about the
likelihood of a positive result, this is not an attractive option. My view is
that journal editors should have recognised this as a pilot study and asked the
authors to do a more extensive replication, rather than dashing into print on
the basis of such slender evidence. In publishing this study, Current Biology
has created a situation where other scientists must now spend time and
resources to establish whether the results hold up.


To establish just how
damaging this can be, consider the case of the FastForword intervention,
developed on the basis of a small trial initially reported in Science in 1996.
After the Science paper, the authors went directly into commercialization of
the intervention, and reported only uncontrolled trials. It took until 2010 for
there to be enough reasonably-sized independent randomized controlled trials to
evaluate the intervention properly in a meta-analysis, at which point it was
concluded that it had no beneficial effect. By this time, tens of thousands of
children had been through the intervention, and hundreds of thousands of
research dollars had been spent on studies evaluating FastForword.


I appreciate that those
reporting exciting findings from small trials are motivated by the best of
intentions – to tell the world about something that seems to help children. But
the reality is that, if the initial trial is not adequately powered, it can be
detrimental both to science and to the children it is designed to help, by
giving such an imprecise and uncertain estimate of the effectiveness of
treatment.


Finally, a comment on
whether it is fair to comment on a research article in a blog, rather than
going through the usual procedure of submitting an article to a journal and
having it peer-reviewed prior to publication. The authors’ reactions to my
blogpost are reminiscent of Felicia Wolfe-Simon’s response to blog-based
criticisms of a paper she published in Science: "The items you are
presenting do not represent the proper way to engage in a scientific
discourse”. Unlike Wolfe-Simon, who simply refused to engage with bloggers,
Facoetti and Gori show willingness to discuss matters further, and present
their side of the story, but they nevertheless it is clear they do not regard a
blog as an appropriate place to debate scientific studies. 



I could not disagree
more. As was readily demonstrated in the Wolfe-Simon case, what has come to be
known as ‘post-publication peer review’ via the blogosphere can allow for new
research to be rapidly discussed and debated in a way that would be quite
impossible via traditional journal publishing. In addition, it brings the
debate to the attention of a much wider readership. Facoetti and Gori feel I
have picked on them unfairly: in fact, I found out about their paper because I
was asked for my opinion by practitioners who worked with dyslexic children.
They felt the results from the Current Biology study sounded too good to be
true, but they could not access the paper from behind its paywall, and in any
case they felt unable to evaluate it properly. I don’t enjoy criticising
colleagues, but I feel that it is entirely proper for me to put my opinion out
in the public domain, so that this broader readership can hear a different
perspective from those put out in the press releases. And the value of blogging
is that it does allow for immediate reaction, both positive and negative. I
don’t censor comments, provided they are polite and on-topic, so my readers
have the opportunity to read the reaction of Facoetti and Gori. 


I should emphasise that I
do not have any personal axe to grind with the study's authors, who I do not
know personally. I’d be happy to revise my opinion if convincing arguments are
put forward, but I think it is important that this discussion takes place in
the public domain, because the issues it raises go well beyond this specific
study.






References


Button, K. S., Ioannidis,
J. P. A., Mokrysz, C., Nosek, B. A., Flint, J., Robinson, E. S. J., &
Munafo, M. R. (2013). Power failure: why small sample size undermines the
reliability of neuroscience. Nature Reviews Neuroscience, advance online publication.
doi: 10.1038/nrn3475


Ioannidis, J. P. A. (2005).
Why most published research findings are false. PLoS Medicine, 2(8), e124. doi:
10.1371/journal.pmed.0020124


Ioannidis, J. P. (2008).
Why most discovered true associations are inflated. Epidemiology 19(5),
640-648.


Ioannidis JP, Pereira TV,
& Horwitz RI (2013). Emergence of large treatment effects from small
trials--reply. JAMA : the journal of the American Medical Association, 309 (8),
768-9 PMID: 23443435




Saturday, March 9, 2013

High-impact journals: where newsworthiness trumps methodology





Here’s a paradox: Most scientists would give their eye teeth to get a paper in a high impact journal, such as Nature, Science, or Proceedings of the National Academy of Sciences. Yet these journals have had a bad press lately, with claims that the papers they publish are more likely to be retracted than papers in journals with more moderate impact factors. It’s been suggested that this is because the high impact journals treat newsworthiness as an important criterion for accepting a paper. Newsworthiness is high when a finding is both of general interest and surprising, but surprising findings have a nasty habit of being wrong.



A new slant on this topic was provided recently by a paper by Tressoldi et al (2013), who compared the statistical standards of papers in high impact journals with those of three respectable but lower-impact journals. It’s often assumed that high impact journals have a very high rejection rate because they adopt particularly rigorous standards, but this appears not to be the case. Tressoldi et al focused specifically on whether papers reported effect sizes, confidence intervals, power analysis or model-fitting. Medical journals fared much better than the others, but Science and Nature did poorly on these criteria. Certainly my own experience squares with the conclusions of Tressoldi et al (2013), as I described in the course of discussion about an earlier blogpost.



Last week a paper appeared in Current Biology (impact factor = 9.65) with the confident title: “Action video games make dyslexic children read better.” It's a classic example of a paper that is on the one hand highly newsworthy, but on the other, methodologically weak. I’m not usually a betting person, but I’d be prepared to put money on the main effect failing to replicate if the study were repeated with improved methodology. In saying this, I’m not suggesting that the authors are in any way dishonest. I have no doubt that they got the results they reported and that they genuinely believe they have discovered an important intervention for dyslexia. Furthermore, I’d be absolutely delighted to be proved wrong: There could be no better news for children with dyslexia than to find that they can overcome their difficulties by playing enjoyable computer games rather than slogging away with books. But there are good reasons to believe this is unlikely to be the case.



An interesting way to evaluate any study is to read just the Introduction and Methods, without looking at Results and Discussion. This allows you to judge whether the authors have identified an interesting question and adopted an appropriate methodology to evaluate it, without being swayed by the sexiness of the results. For the Current Biology paper, it’s not so easy to do this, because the Methods section has to be downloaded separately as Supplementary Material. (This in itself speaks volumes about the attitude of Current Biology editors to the papers they publish: Methods are seen as much less important than Results). On the basis of just Introduction and Methods, we can ask whether the paper would be publishable in a reputable journal regardless of the outcome of the study.



On the basis of that criterion, I would argue that the Current Biology paper is problematic, purely on the basis of sample size. There were 10 Italian children aged 7 to 13 years in each of two groups: one group played ‘action’ computer games and the other was a control group playing non-action games (all games from Wii's Rayman Raving Rabbids - see here for examples). Children were trained for 9 sessions of 80 minutes per day over two weeks. Unfortunately, the study was seriously underpowered. In plain language, with a sample this small, even if there is a big effect of intervention, it would be hard to detect it. Most interventions for dyslexia have small-to-moderate effects, i.e. they improve performance in the treated group by .2 to .5 standard deviations. With 10 children per group, the power is less than .2, i.e. there’s a less than one in five chance of detecting a true effect of this magnitude. In clinical trials, it is generally recommended that the sample size be set to achieve power of around .8. This is only possible with a total sample of 20 children if the true effect of intervention is enormous – i.e. around 1.2 SD, meaning there would be little overlap between the two groups’ reading scores after intervention. Before doing this study there would have been no reason to anticipate such a massive effect of this intervention, and so use of only 10 participants per group was inadequate. Indeed, in the context of clinical trials, such a study would be rejected by many ethics committees (IRBs) because it would be deemed unethical to recruit participants for a study which had such a small chance of detecting a true effect.



But, I hear you saying, this study did find a significant effect of intervention, despite being underpowered. So isn’t that all the more convincing? Sadly, the answer is no. As Christley (2010) has demonstrated, positive findings in underpowered studies are particularly likely to be false positives when they are surprising – i.e., when we have no good reason to suppose that there will be a true effect of intervention. This seems particularly pertinent in the case of the Current Biology study – if playing active computer games really does massively enhance children’s reading, we might have expected to see a dramatic improvement in reading levels in the general population in the years since such games became widely available.



The small sample size is not the only problem with the Current Biology study. There are other ways in which it departs from the usual methodological requirements of a clinical trial: it is not clear how the assignment of children to treatments was made or whether assessment was blind to treatment status, no data were provided on drop-outs, on some measures there were substantial differences in the variances of the two groups, no adjustment appears to have been made for the non-normality of some outcome measures, and a follow-up analysis was confined to six children in the intervention group. Finally, neither group showed significant improvement in reading accuracy, where scores remained 2 to 3 SD below the population mean (Tables S1 and S3): the group differences were seen only for measures of reading speed.



Will any damage be done? Probably not much – some false hopes may be raised, but the stakes are not nearly as high as they are for medical trials, where serious harm or even death can result from wrong results. There is concern, however, that quite apart from the implications for families of children with reading problems, there is another issue here, about the publication policies of high-impact journals. These journals wield immense power. It is not overstating the case to say that a person’s career may depend on having a publication in a journal like Current Biology (see this account – published, as it happens, in Current Biology!). But, as the dyslexia example illustrates, a home in a high-impact journal is no guarantee of methodological quality. Perhaps this should not surprise us: I looked at the published criteria for papers on the websites of Nature, Science, PNAS and Current Biology. None of them mentioned the need for strong methodology or replicability; all of them emphasised “importance” of the findings.



Methods are not a boring detail to be consigned to a supplement: they are crucial in evaluating research. My fear is that the primary goal of some journals is media coverage, and consequently science is being reduced to journalism, and is suffering as a consequence.



References



Brembs, B., & Munafò, M. R. (2013). Deep impact: Unintended consequences of journal rank. arXiv:1301.3748.



Christley, R. M. (2010). Power and error: increased risk of false positive results in underpowered studies. The Open Epidemiology Journal, 3, 16-19.



Halpern, S. D.,  Karlawish, J. T, & Berlin, J. A. (2002). The continuing unethical conduct of underpowered clinical trials. Journal of the American Medical Association, 288(3), 358-362. doi: 10.1001/jama.288.3.358



Lawrence, P. A. (2007). The mismeasurement of science. Current Biology, 17(15), R583-R585. doi: 10.1016/j.cub.2007.06.014



Tressoldi, P., Giofré, D., Sella, F., & Cumming, G. (2013). High Impact = High Statistical Standards? Not Necessarily So. PLoS ONE, 8 (2) DOI: 10.1371/journal.pone.0056180