Showing posts with label research. Show all posts
Showing posts with label research. Show all posts

Tuesday, June 17, 2014

The University as big business:


The case of King's College London



 










© www.CartoonStock.com




King's College London is in the news for all the wrong reasons. In a document full of weasel words ('restructuring', 'consultation exercise'), staff in the schools of medicine and biomedical sciences, and the Institute of Psychiatry were informed last month that 120 of them were at risk of redundancy. The document was supposed to be confidential but was leaked to David Colquhoun who has posted a link to it on his blog.  This isn't the first time KCL has been in the news for its 'robust' management style. A mere four years ago, a similar though smaller purge was carried out at the Institute of Psychiatry, together with a major divestment in Humanities at KCL.



Any tale of redundancies on such a scale is a human tragedy, whether it be in a car factory or a University. But the two cases are not entirely parallel. For a car factory, the goal of the business is to make a profit. A sensible employer will try to maintain a cheerful and committed workforce, but ultimately they may be sacrificed if it proves possible to cut costs by, for instance, getting machines to do jobs that were previously done by people. The fact that a University is adopting that approach – sacking its academic staff to improve its bottom line – is an intellectual as well as a human tragedy. It shows how far we have moved towards the identification of universities with businesses.



Traditionally, a university was regarded as an institution whose primary function was the furtherance of learning and knowledge. Money was needed to maintain the infrastructure and pay the staff, but the money was a means to an end, not an end in itself. However, it seems that this quaint notion is now rejected in favour of a model of a university whose success is measured in terms of its income, not in terms of its intellectual capital.



The opening paragraph of the 'consultation document' is particularly telling: "King’s has built a reputation for excellence and has established itself as a world class university. Our success has been built on growing research volumes in key areas, improving research quality, developing our resources and offering quality teaching to attract the best students in an increasingly competitive environment." Note there is no mention of the academic staff of the institution. They are needed, of course, to "grow research volumes" (ugh!), just as factory workers are needed to manufacture cars. But they aren't apparently seen as a key feature of a successful academic institution. Note too the emphasis is on increasing the amount of research rather than research quality.



The most chilling feature of the document is the list of criteria that will be used to determine which staff are 'at risk'.  You are safe if you play a key role in teaching, or if you have grant income that exceeds a specified amount, dependent on your level of seniority.

What's wrong with this? Well, here are four points just for starters:



1. KCL management justifies its actions as key for "maintaining and improving our position as one of the world’s leading institutions". Sorry, I just don't get it. You don't improve your position by shedding staff, creating a culture of fear, and deterring research superstars from applying for positions in your institution in future.



2. The 'restructuring' treats individual scientists as islands. The Institute of Psychiatry has over the years built up a rich research community, where there are opportunities for people to bounce ideas off each other and bring complementary skills to tackling difficult problems. Making individuals redundant won't just remove an expense from the KCL balance sheet – it will also affect the colleagues of those who are sacked. 



3. As I've argued previously, the use of research income as a proxy measure of research excellence distorts and damages science. It provides incentives for researchers to get grants for the sake of it – the more numerous and more expensive the better. We end up with a situation where there is terrific waste because everyone has a massive backlog of unpublished work.

 

4. I suspect that part of the motivation behind the "restructuring" is in the hope that new buildings and infrastructure might reverse the poor showing of KCL in recent league tables of student satisfaction. If so, the move has backfired spectacularly. The student body at KCL has started a petition against the sackings, which has drawn attention to the issue worldwide.I urge readers to sign it.



Management at KCL just doesn't seem to get a very basic fact about running a university: Its academic staff are vital for the university's goal of achieving academic excellence. They need to be fostered, not bullied. One feels that if KCL were falling behind in a boat race, they'd respond by throwing out some of the rowers.


Sunday, May 11, 2014

Changing the landscape of psychiatric research:


What will the RDoC initiative by NIMH achieve?










©CartoonStock.com




There's a lot wrong with current psychiatric classification. Every few years, the American Psychiatric Association comes up with a new set of labels and diagnostic criteria, but whereas the Diagnostic and Statistical Manual used to be seen as some kind of Bible for psychiatrists, the latest version, DSM5, has been greeted with hostility and derision. The number of diagnostic categories keeps multiplying without any commensurate increase in the evidence base to validate the categories. It has been argued that vested interests from pharmaceutical companies create pressures to medicalise normality so that everyone will sooner or later have a diagnosis (Frances, 2013). And even excluding such conflict of interest, there are concerns that such well-known categories as schizophrenia and depression lack reliability and validity (Kendell & Jablensky, 2003).



In 2013, Tom Insel, Director of the US funding agency, National Institute of Mental Health (NIMH), created a stir with a blogpost in which he criticised the DSM5 and laid out the vision of a new Research Domain Criteria (RDoC) project. This aimed "to transform diagnosis by incorporating genetics, imaging, cognitive science, and other levels of information to lay the foundation for a new classification system."



He drew parallels with physical medicine, where diagnosis is not made purely on the basis of symptoms, but also uses measures of underlying physiological function that help distinguish between conditions and indicate the most appropriate treatment. This, he argued, should be the goal of psychiatry, to go beyond presenting symptoms to underlying causes, reconceptualising disorders in terms of neural systems.



This has, of course, been a goal for many researchers for several years, but Insel expressed frustration at the lack of progress, noting that at present: "We cannot design a system based on biomarkers or cognitive performance because we lack the data". That being the case, he argued, a priority for NIMH should be to create a framework for collecting relevant data. This would entail casting aside conventional psychiatric diagnoses, working with dimensions rather than categories, and establishing links between genetic, neural and behavioural levels of description.



This represents a massive shift in research funding strategy, and some are uneasy about it. Nobody, as far as I am aware, is keen to defend the status quo, as represented by DSM.  As Insel remarked in his blogpost: "Patients with mental disorders deserve better". The issue is whether RDoC is going to make things any better. I see five big problems.



1. McLaren (2011) is among those querying the assumption that mental illnesses are 'disorders of brain circuits'. The goal of the RDoC program is to fill in a huge matrix with new research findings. The rows of the matrix are not the traditional diagnostic categories: instead they are five research domains: Negative Valence Systems, Positive Valence Systems, Cognitive Systems, Systems for Social Processes, Arousal/Regulatory Systems, each of which has subdivisions: e.g. Cognitive Systems is broken down into Attention, Perception, Working memory, Declarative memory, Language behavior and Cognitive (effortful) control. The columns of the matrix are Genes, Molecules, Cells, Circuits, Physiology, Behavior, Self-Reports, and Paradigms. Strikingly absent is anything about experience or environment.



This seems symptomatic of our age. I remember sitting through a conference presentation about a study investigating whether brain measures could predict response to cognitive behaviour therapy in depression.  OK, it's possible that they might, but what surprised me was that no measures of past life events or current social circumstances were included in the study. My intuitions may be wrong, but it would seem that these factors are likely to play a role. My impression is that some of the more successful interventions developed in recent years are based not on neurobiology or genetics, but on a detailed analysis of the phenomenology of mental illness, as illustrated, for example, by the work of my colleagues David Clark and Anke Ehlers. Consideration of such factors is strikingly absent from RDoC.



 2. The goal of the RDoC is ultimately to help patients, but the link with intervention is unclear. Suppose I become increasingly obsessed with checking electrical switches, such that I am unable to function in my job. Thanks to the RDoC program, I'm found to have a dysfunctional neural circuit. Presumably the benefit of this is that I could be given a new pharmacological intervention targeting that circuit, which will make me less obsessive. But how long will I stay on the drug? It's not given me any way to cope with the tendency of checking the unwanted thoughts that obtrude into my consciousness, and they are likely to recur when I come off it.  I'm not opposed to pharmacological interventions in principle, but they tend not to have a 'stop rule'. 



There are psychological interventions that tackle the symptoms and the cognitive processes that underlie them more directly.  Could better knowledge of neurobiological correlates help develop more of these?  I guess it is possible, but my overall sense is that this translational potential is exaggerated – just as with the current hype around 'educational neuroscience'. The RDoC program embodies a mistaken belief that neuroscientific research is inherently better than psychological research because it deals with primary causes, when in fact it cannot capture key clinical phenomena. For instance, the distinction between a compulsive hand-washer and a compulsive checker is unlikely to have a clear brain correlate, yet we need to know about the specific symptoms of the individual to help them overcome them.



3. Those proposing RDoC appear to have a naive view of the potential of genetics to inform psychiatry.  It's worth quoting in detail from their vision of the kinds of study that would be encouraged by NIMH, as stated here:



Recent studies have shown that a number of genes reported to confer risk for schizophrenia, such as DISC1 (“Disrupted in schizophrenia”) and neuregulin, actually appear to be similar in risk for unipolar and bipolar mood disorders. ... Thus, in one potential design, inclusion criteria might simply consist of all patients seen for evaluation at a psychotic disorders treatment unit. The independent variable might comprise two groups of patients: One group would be positive and the other negative for one or more risk gene configurations (SNP or CNV), with the groups matched on demographics such as age, sex, and education. Dependent variables could be responses to a set of cognitive paradigms, and clinical status on a variety of symptom measures. Analyses would be conducted to compare the pattern of differences in responses to the cognitive or emotional tasks in patients who are positive and negative for the risk configurations.



This sounds to me like a recipe for wasting a huge amount of research funding. The effect sizes of most behavioural/cognitive genetic associations are tiny and so one would need an enormous sample size to see differences related to genotype. Coupled with an open-ended search for differences between genotypes on a battery of cognitive measures, this would undoubtedly generate some 'significant' results which could go on to mislead the field for some time before a failure to replicate was achieved (cf. Munafò, & Gage, 2013).



The NIMH website notes that "the current diagnostic system is not informed by recent breakthroughs in genetics". There is good reason for that: to date, the genetic findings have been disappointing. Such associations as are found either indicate extremely rare and heterogeneous mutations of large effect and/or involve common genetic variants whose small effects are not of clinical significance. We cannot know what the future holds, but to date talk of 'breakthroughs' is misleading.



4. Some of the entries in the RDoC matrix also suggest a lack of appreciation of the difference between studying individual differences versus group effects.  The RDoC program is focused on understanding individual differences. That requires particularly stringent criteria for measures, which need to be adequately reliable, valid and sensitive to pick up differences between people.  I appreciate that the published RDoC matrices are seen as a starting-point and not as definitive, but I would recommend that more thought goes into establishing the psychometric credibility of measures before embarking on expensive studies looking for correlations between genes, brains and behaviour. If the rank ordering of a group of people on a measure is not the same from one occasion to another, or if there are substantial floor or ceiling effects, that measure is not going to be much use as an indicator of an underlying construct. Furthermore, if different versions of a task that are supposed to tap into a single construct give different patterns of results, then we need a rethink – see e.g. Foti et al, 2013; Shilling et al, 2013, for examples.  Such considerations are often ignored by those attempting to move experimental work into a translational phase. If we are really to achieve 'precision medicine' we need precise measures.



5. The matrix as it stands does not give much confidence that the RDoC approach will give clearer gene-brain-behaviour links than traditional psychiatric categories.



For instance, BDNF appears in the Gene column of the matrix for the constructs of acute threat, auditory perception, declarative memory, goal selection, and response selection. COMT appears with threat, loss, frustrative nonreward, reward learning, goal selection, response selection and reception of facial communication. Of course, it's early days. The whole purpose of the enterprise is to flesh out the matrix with more detailed and accurate information. Nevertheless, the attempts at summarising what is known to date do not inspire confidence that this goal will be achieved.



After such a list of objections to RDoC, I do have one good thing to say about it, which is that it appears to be encouraging and embracing data-sharing and open science. This will be an important advance that may help us find out more quickly which avenues are worth exploring and which are cul-de-sacs. I suspect we will find out some useful things from the RDoC project: I just have reservations as to whether they will be of any benefit to psychiatry, or more importantly, to psychiatric patients.



References

Foti, D., Kotov, R., & Hajcak, G. (2013). Psychometric considerations in using error-related brain activity as a biomarker in psychotic disorders. Journal of Abnormal Psychology, 122(2), 520-531. doi: 10.1037/a0032618



Frances, A. (2013). Saving normal: An insider's revolt against out-of-control psychiatric diagnosis, DSM-5, big pharma, and the medicalization of ordinary life. New York: HarperCollins.



Kendell, R., & Jablensky, A. (2003). Distinguishing between the validity and utility of psychiatric diagnoses. American Journal of Psychiatry, 160, 4-12.



McLaren, N. (2011). Cells, Circuits, and Syndromes: A Critical Commentary on the NIMH Research Domain Criteria Project Ethical Human Psychology and Psychiatry, 13 (3), 229-236 DOI: 10.1891/1559-4343.13.3.229



Munafò, M. R., & Gage, S. H. (2013). Improving the reliability and reporting of genetic association studies. Drug and Alcohol Dependence(0). doi: http://dx.doi.org/10.1016/j.drugalcdep.2013.03.023



Shilling, V. M., Chetwynd, A., & Rabbitt, P. M. A. (2002). Individual inconsistency across measures of inhibition: an investigation of the construct validity of inhibition in older adults. Neuropsychologia, 40, 605-619.





This article (Figshare version) can be cited as:

 Bishop, Dorothy V M (2014): Changing the landscape of psychiatric research: What will the RDoC initiative by NIMH achieve?. figshare. http://dx.doi.org/10.6084/m9.figshare.1030210 

Sunday, January 12, 2014

Why does so much research go unpublished?







As described in my last blogpost, I attended an excellent symposium on waste in research this week. A recurring theme was research that never got published. Rosalind Smyth described her experience of sitting on the funding panel of a medium-sized charity. The panel went to great pains to select the most promising projects, and would end a meeting with a sense of excitement about the great work that they were able to fund. A few years down the line, though, they'd find that many of the funds had been squandered. The work had either not been done, or had been completed but not published.



In order to tackle this problem, we need to understand the underlying causes. Sometimes, as Robert Burns noted, the best-laid schemes go wrong. Until you've tried to run a few research projects, it's hard to imagine the myriad different ways in which life can conspire to mess up your plans. The eight laws of psychological research formulated by Hodgson and Rollnick are as true today as they were 25 years ago.



But much research remains unpublished despite being completed. Reasons are multiple, and the strategies needed to overcome them are varied, but here is my list of the top three problems and potential solutions.




Inconclusive results




Probably the commonest reason for inconclusive results is lack of statistical power. A study is undertaken in the fond hope that a difference will be found between condition X and condition Y, and if the difference is found, there is great rejoicing and a rush to publish. A negative result should also be of interest, provided the study was well-designed and adequately motivated. But if the sample is small, then we can't be sure whether our failure to observe the effect is because it is absent: a real but small effect could be swamped by noise. 



I think the solution to this problem lies in the hands of funding panels and researchers: quite simply, they need to take statistical power very seriously indeed and to consider carefully whether anything will be learned from a study if the anticipated effects are not obtained. If not, then the research needs to be rethought. In the fields of genetics and clinical trials, it is now recognised that multicentre collaborations are the way forward to ensure that studies are conducted with sufficient power to obtain a conclusive result.




Rejection of completed work by journals




Even well-conducted and adequately powered studies may be rejected by journals if the results are not deemed to be exciting. To solve this problem, we must look to journals. We need recognition that - provided a study is methodologically strong and well-motivated - negative results can be as informative as positive ones. Otherwise we are doomed to waste time and money pursuing false leads.  As Paul Glasziou has emphasised, failure is part of the research process. It is important to tell people about what doesn't work if we are not to repeat our mistakes.



We do now have some journals that will publish negative results, and there is a growing move toward pre-registration of studies, with guaranteed publication if the methods meet quality criteria. But there is still a lot to be done, and we need a radical change of mindset about what kinds of research results are valuable.




Lack of time




Here, I lay the blame squarely on the incentive structures that operate in universities. To get a job, or to get promoted, you need to demonstrate that you can pull in research income. In many UK institutions this is quite explicit, and promotions criteria may give a specific figure to aim for of X thousand pounds research income per annum. There are few UK universities whose strategic plan does not include a statement about increasing research funding. This has changed the culture dramatically;  as Fergus Millar put it: "in the modern British university, it is not that funding is sought in order to carry out research, but that research projects are formulated in order to get funding".



Of course, for research to thrive, our Universities need people who can compete for funding to support their work. But the acquisition of funding has become an end in itself, rather than a means to an end. This has the pernicious effect of driving people to apply for grant after grant, without adequately budgeting for the time it takes to analyse and write up research, or indeed to carefully think about what they are doing.  As I argued previously, even junior researchers these days have an 'academic backlog' of unwritten papers.



At the Lancet meeting there were some useful suggestions for how we might change incentive structures to avoid such waste. Malcolm MacLeod argued researchers should be evaluated not by research income and high-impact publications, but by the quality of their methods, the extent to which their research was fully reported, and the reproducibility of findings. An-Wen Chan echoed this, arguing for performance metrics that recognise full dissemination of research and use of research datasets by other groups. However, we may ask whether such proposals have any chance of being adopted when University funding is directly linked to grant income, and Universities increasingly view themselves as businesses.



I suspect we would need revised incentives to be reflected at the level of those allocating central funding before vice-chancellors took them seriously.  It would, however, be feasible for behaviour to be shaped at the supply end, if funders adopted new guidelines. For a start, they could look more carefully at the time commitments of those to whom grants are given: in my experience this is never taken into consideration, and one can see successful 'fat cats' accumulating grant after grant, as success builds on success. Funders could also monitor more closely the outcomes of grants: Chan noted that NIHR withholds 10% of research funds until a paper based on the research has been submitted for publication. Moves like this could help us change the climate so that an award of a grant would confer responsibility on the recipient to carry through the work to completion, rather than acting solely to embellish the researcher's curriculum vitae.





References



Chan, A., Song, F., Vickers, A., Jefferson, T., Dickersin, K., Gotzsche, P., Krumholz, H. M., Ghersi, D., & van der Worp, H. B. (2014). Increasing value and reducing waste: addressing inaccessible research Lancet (8 Jan ) : 10.1016/S0140-6736(13)62296-5





Macleod, M. R., Michie, S., Roberts, I., Dirnagl, U., Chalmers, I., Ioannidis, J. P. A., . . . Glasziou, P. (2014). Biomedical research: increasing value, reducing waste. Lancet, 383(9912), 101-104.

Thursday, January 9, 2014

Off with the old and on with the new: the pressures against cumulative research


 

Yesterday I escaped a very soggy Oxford to make it down to London for a symposium on "Increasing value, reducing waste" in Research. The meeting marked the publication of a special issue of the Lancet containing five papers and two commentaries, which can be downloaded here.



I was excited by the symposium because, although the focus was on medicine, it raised a number of issues that have much broader relevance for science, including several that I have raised on this blog, including pre-registration of research, criteria used by high-impact journalsethics regulation, academic backlogs, and incentives for researchers. It was impressive to see that major players in the field of medicine are now recognizing that there is a massive problem of waste in research. Better still, they are taking seriously the need to devise ways in which this could be fixed.



I hope to blog about more of the issues that came up in the meeting, but for today I'll confine myself to one topic that I hadn't really thought about much before, but which I see as important, namely the importance of doing research that builds on previous research, and the current pressures against this.



Iain Chalmers presented one of the most disturbing slides of the day, a forest plot of effect sizes found in medical trials for a treatment to prevent bleeding during surgery.




Based on Figure 3 of Chalmers et al, 2014

Time is along the x-axis, and the horizontal line corresponds to a result where the active and control treatments do not differ. Points which are below the line and whose fins do not cross it show a beneficial effect of treatment. The graph shows that the effectiveness of the treatment was clearly established by around 2002, yet a further 20 studies including several hundred patients were reported in the literature after that date. Chalmers made the point that it is simply unethical to do a clinical trial if previous research has already established an effect. The problem is that researchers often don't check the literature to see what has already been done, and so there is wasteful repetition of studies. In the field of medicine this is particularly serious because patients may be denied the most effective treatment if they enrol in a research project.



Outside medicine, I'm not sure this is so much of an issue. In fact, as I've argued elsewhere, in psychology and neuroscience I think there's more of a problem with lack of replication. But there definitely is much neglect of prior research. I lose count of the number of papers I review where the introduction presents a biased view of the literature that supports the authors' conclusions. For instance, if you are interested in the relation between auditory deficit and children's language disorders, it is possible to write an introduction presenting this association as an established fact, or to write one arguing that it has been comprehensively debunked. I have seen both.



Is this just lazy, biased or ignorant authors? In part, I suspect it is. But I think there is a deeper problem which has to do with the insatiable demand for novelty shown by many journals, especially the high-impact ones. These journals typically have a lot of pressure on page space and often allow only 500 words or less for an introduction. Unless authors can refer to a systematic review of the topic they are working on, they are obliged to give the briefest account of prior literature. It seems we no longer value the idea that research should build on what has gone before: rather, everyone wants studies that are so exciting that they stand alone. Indeed, if a study is described as 'incremental' research, that is typically the death knell in a funding committee.



We need good syntheses of past research, yet these are not valued because they are not deemed novel. One point made by Iain Chalmers was that funders have in the past been reluctant to give grants for systematic reviews. Reviews also aren't rated highly in academia: for instance, I'm proud of a review on mismatch negativity that I published in Psychological Bulletin in 2007. It not only condensed and critiqued existing research, but also discovered patterns in data that had not previously been noted. However, for the REF, and for my publications list on a grant renewal, reviews don't count.



We need a rethink of our attitude to reviews. Medicine has led the way and specified rigorous criteria for systematic reviews, so that authors can't just cherrypick specific studies of interest. But it has also shown us that such reviews are an invaluable part of the research process. They help ensure that we do not waste resources by addressing questions that have already been answered, and they encourage us to think of research as a cumulative, developing process, rather than a series of disconnected, dramatic events.



Reference

Chalmers, Iain, Bracken, Michael B., Djulbegovic, Ben, Garattini, Silvio, Grant, Jonathan, Gülmezoglu, A. Metin, Howells, David W., Ioannidis, John P. A., & Oliver, Sandy (2014). How to increase value and reduce waste when research priorities are set Lancet : 10.1016/S0140-6736(13)62229-1

Thursday, September 12, 2013

Evaluate, evaluate, evaluate







© www.CartoonStock.com


When I was
starting out on a doctorate, I’d look at the senior people in my field and
wonder if I’d ever be like them. It must be great, I thought, to reach the
advanced age of 40. By then you’d have learned everything you needed to know to
do great science, and you could just focus on doing it. I suspect today’s crop
of grad students are a bit more savvy than I was, but all the same, I wonder if
they realise just how wrong that picture is – for two reasons.



First, you never stop learning. The field moves on. Instead of getting easier, it gets harder. I
remember when techniques such as functional brain imaging first came along. The
most competent people in that area were either those who had developed the
methods, or young people who learned them as grad students. If you were of the
generation above, you had three choices: ignore the methods, spend time
learning them, or hire junior people who knew what they were doing. As the
methods evolve, they get ever more complex, and meanwhile, your own brain
starts to shrink. So if you are anticipating making it to a tenured post and
then settling down in your armchair, think again.



Second, the more
senior you get, the more of your time is spent, not on doing your own research,
but on evaluation. You learn that an email entitled ‘invitation’ should not
make your spirits rise: it’s just a desperate attempt to put a positive spin on
a request for you to do more work for no reward. You get regular ‘invitations’ to review
papers and grants, write job references, appraise promotion bids, sit on
interview panels and examine theses. If you are involved in teaching, you’ll
also be engaged in numerous other forms of appraisal.



I was prompted to
think about this when someone asked on an electronic forum what was a
reasonable number of doctoral theses to examine each year. The general consensus was two: though it will
obviously depend on what other commitments someone has. It also varies from
country to country. There are some jolly
places in Europe where a PhD viva is just an excuse for a boozy party with a
lot of dressing up in funny gowns and hats. In UK psychology, the whole thing
is no fun at all: you have to read a document of 50,000-70,000 words reporting
a body of work based on a series of experimental studies. You then write a
report on it and see the candidate for a face-to-face viva, which is typically
2 to 3 hours long. Although failure is uncommon, it is not assumed that the
candidate will pass (unlike in the viva-as-party countries), and weeping or
catatonic candidates are not unheard of. Taking into account travel, etc., if
you are going to do a proper job, you are probably talking about three days’
work. For this you get paid around the minimum wage – the fee for examining is
typically somewhere between £120 and £200.



So why do we do
it? The major reason is because the entire academic enterprise depends on
reciprocity: we want people to examine our students and review our papers and
grants. In addition, it’s important to maintain standards, and to ensure that
degrees, promotions, publications and grants go to those who merit them. But the demands keep growing. In the 37 weeks of this year I’ve been asked
to review 76 papers and six grants. I agreed to review 16 papers and three of
the grants. This, of course, is nothing compared with being a journal editor or
serving on a grants board, something that most of us will do at some point.



Clearly, if I
agreed to do everything I was asked, I’d have no time for anything else. Of course, one learns to say no. But
awareness of these pressures has made me look with rather a critical eye at how
we use evaluation. There is, for instance, research suggesting that job interviews aren’t very useful at identifying good candidates:  we tend to be seduced by immediate
impressions, which may not be a good indicator of a person’s suitability. Like
most people, I’d be reluctant to take on an employee I hadn’t interviewed, but
if Daniel Kahneman is to be believed, this is just because I am a victim of the
Illusion of Validity.



I’m a supporter of the peer review system used by
journals, and here I feel  I’m on more
solid ground, because I can point to instances where my papers have been
improved by input from reviewers. Nevertheless, where reviewing is used simply
to reject/accept papers or grant proposals, 
and where fine-grained decisions have to be made between many
high-quality submissions, agreement between experts may be little better than chance
(e.g. Fogelholm et al, 2012). Nevertheless, we stick with it, because it’s hard
to know what to put in its place.



I’ve written a fair bit about that expensive and time-consuming evaluation process that UK academics
engage in, the REF. It requires experts to make
judgements of whether, for instance, papers are of 3* or 4* quality, a
distinction based on whether the research is “world leading” or “internationally
excellent…. but falls short of the highest standards of excellence.” The reliability of such judgements has not, to my knowledge, been evaluated, yet large amounts of funding depend on them. Those on REF committees are in the same situation as Pavlov’s poor dogs, having
to make distinctions that are on the one hand impossible (discriminating
circles and ellipses that become increasingly similar) and on the other hand
very important (get it wrong and you get a shock).



There is one good
thing about doing so much evaluation. You have the opportunity to see what
others are doing – you may be the first person to read an important new paper,
or examine a ground-breaking thesis. You may be forced to engage with different
ways of thinking, and confronted with new topics and ideas. You may be able to provide useful input to authors. And since you
yourself will be evaluated, it can be useful to see life from the other side of
the table, as the person doing the evaluating. But all too often, even these
advantages fail to compensate for the fact that as a senior academic you will
spend more and more time on evaluation of others and less and less doing your
own research.



Reference

Fogelholm, Mikael, Leppinen, Saara, Auvinen, Anssi, Raitanen, Jani, Nuutinen, Anu, & Väänänen, Kalervo (2012). Panel discussion does not improve reliability of peer review for medical research grant proposals Journal of Clinical Epidemiology, 65 (1), 47-52 DOI: 10.1016/j.jclinepi.2011.05.001