27 September 2016

The last word on the well-being and genomics saga (or maybe not)

It looks like the dust may be settling on the long-running saga of the Fredrickson et al. studies of genomics and well-being, and the Brown et al. reanalyses of the same.  We have probably arrived at the end of the discussion in the formal literature.  Both sides (of course) think they have won, but the situation on the ground probably looks like a bit of a mess to the casual observer.

I last blogged about this over two years ago.  Since then, Fredrickson et al. have produced a second article, published in PLoS ONE, partly re-using the data from the first, and claiming to have found "the same results" --- except that their results were also different (read the articles and decide for yourself) --- with a new mathematical model.  We wrote a reply article, which was also published in PLoS ONE.  Dr. Fredrickson wrote a formal comment on our article, and we wrote a less-formal comment on that.

I could sum up all of the above articles and comments here, but that would serve little purpose.  All of the relevant evidence is available at those links, and you can evaluate it for yourself. However, I thought I would take a moment here to write up a so-far unreported aspect of the story, namely how Fredrickson et al. changed the archived version of one of their datasets without telling anybody.

In the original version of the GSE45330 dataset used in Fredrickson et al.'s 2013 PNAS article, a binary categorical variable, which should have contained only 0s and 1s, contained a 4. This, of course, turned it into basically a continuous variable when it was thrown as a "control" into the regressions that were used to analyse the data. We demonstrated that fixing this variable caused the main result of the 2013 PNAS article --- which was "supported" by the fact that the two bars in Figure 2A were of equal height but opposite sign --- to break; one of the bars more than halved in size.(*) For reasons of space, and because it was just a minor point compared to the other deficiencies of Fredrickson et al.'s article, this coding error was not covered in the main text of our 2014 PNAS reply, but it was handled in some detail in the supporting information.

Fredrickson et al. did not acknowledge their coding error at that time. But by the time they re-used these data with a new model in their subsequent PLoS ONE article (as the "Discovery" sample, which was pooled with the "Confirmation" sample to make a third dataset), they had corrected the coding error, and uploaded the corrected version to the GEO repository, causing the previous version to be overwritten without a trace.
This means that if, today, you were to read Fredrickson et al.'s 2013 PNAS article and download the corresponding dataset, you would no longer be able to reproduce their published Figure 2A; you would only be able to generate the "corrected"(*) version.

The new version of the GSE45330 dataset was uploaded on July 15, 2014 --- a month after our PNAS article was accepted, and a month before it was published.  When our article appeared, it was accompanied by a letter from Drs. Fredrickson and Cole (who would certainly have received --- probably on the day that our article was accepted --- a copy of our article and the supporting information, in order to write their reply), claiming that our analysis was full of errors.  Their own coding error, which they must have been aware of because /a/ we had pointed it out, and /b/ they had corrected it a month earlier, was not mentioned.

Further complicating matters is the way in which, early in their 2015 PLoS ONE article, Fredrickson et al. attempted to show continuity between their old and new samples in their Figure 1C.  Specifically, this figure reproduced the incorrect bars from their PNAS article's Figure 2A (i.e., the bars produced without the coding error having been corrected). So
Fredrickson et al. managed to use both the uncorrected and corrected versions of the data in support of their hypotheses, in the same PLoS ONE article.  I would like to imagine that this is unprecedented, although very little surprises me any more.

We did manage to get PLoS ONE to issue a correction for the figure problem.  However, this only shows the final version of the image, not "before" and "after", so here, as a public service, is the original (left) and the corrected version (right).  As seems to be customary, however, the text of Fredrickson et al.'s correction does not accept that this change has any consequences for the substantive conclusions of their research.(*)

Alert readers may have noticed that this correction leaves a problem with Fredrickson et al.'s 2013 PNAS article, which still contains the uncorrected Figure 2A, illustrating the authors' (then) hypotheses that hedonic and eudaimonic well-being had equal and opposite parts to play in determining gene expression in the immune system:
CTRA gene expression varied significantly as a function of eudaimonic and hedonic well-being (Fig. 2A). As expected based on the inverse association of eudaimonic well-being with depressive symptoms, eudaimonic well-being was associated with down-regulated CTRA gene expression (contrast, P = 0.0045). In contrast, CTRA gene expression was significantly up-regulated in association with increasing levels of hedonic well-being (p. 13585)
But as we have seen, the corrected version of the figure shows a considerable difference between the two bars, representing hedonic and eudaimonic well-being, especially if one considers that the bars represent log-transformed numbers.  This implies that the 2013 PNAS article is now severely flawed(*); Figure 2A needs to be replaced, as does the claim about the opposite effects of hedonic and eudaimonic well-being.  We contacted PNAS, asking for a correction to be issued, and were told that they consider the matter closed.  So now, both the corrected and uncorrected figures are in the published literature, and two different and contradictory conclusions about the relative effects of hedonic and eudaimonic well-being on gene expression are available to be cited, depending on which fits the narrative at hand.  Isn't science wonderful?

There seems to be one remaining question, which is exactly how unethical it was for the alterations to the dataset to have been made.  We made a complaint to the Office of Research Integrity, and it went nowhere.  It could be argued, I suppose, that the new version of the data was better than the old one. But we certainly didn't feel that Drs. Fredrickson and Cole had acted in an open and transparent manner.  They read our article and supporting information, saw the coding error that we had found, corrected it without acknowledging us, and then published a letter saying that our analyses were full of errors.  I find this, if I may use a little British understatement for a moment, to be "not entirely collegial".  If this is the norm when critiques of published work are submitted through the peer-review system, as psychologists were recently exhorted to do by a senior figure in the field, perhaps we should not be surprised when some people who discover problems in published articles decide to use less formal methods to comment.



(*) Running through this entire post, of course, is the assumption that the reader has set aside for the moment our demonstration of all of the other flaws in the Fredrickson et al. articles, including the massive overfitting and the lack of theoretical coherency.  Arguably, those flaws make the entire question of the coding error moot, since even the "corrected" version of the figures very likely fails to correspond to any real effect.  But I think it's important to look at this aspect of the story separately from all of the other noise, as an example of how difficult it can be to get even the most obvious errors in the literature corrected.

19 August 2016

It's a small world

I am a co-author on an article that was published (open access!) yesterday (2016-08-18) in the Journal of Social and Political Psychology, along with Stephan Lewandowsky, Michael Mann, and Harris Friedman.  It has an amusing twist to it that illustrates how small the world is.

The idea for this article was floated by Stephan Lewandowsky back in 2013.  He got in touch with Harris Friedman after our article (Brown, Sokal, & Friedman, 2013; full text here) was published, causing some ripples in psychological circles, in American Psychologist.  Steve saw the story of the BSF article as a good example of how people from outside science ought to go about trying to correct problems in the literature, in contrast to the ways in which certain people attack scientists, verbally or even physically, especially when it comes to controversial areas such as research using animals, global warming, genetically-modified organisms, nuclear power, and vaccines.

For various reasons, it took a while to get the drafting process started, but I'm pleased the article has been published now, and not just because it includes Monty Python's The Meaning of Life in the references section.  (I have previously cited This is Spinal Tap; if anyone has any good ideas for ways to cite either Wayne's World or Pulp Fiction, I'm all ears.)

Actually, I didn't know much at all about Michael Mann until I saw his name included in the e-mails at the start of the project.  I was aware that there was something controversial in climate science to do with hockey sticks, but I tend to steer clear of the global warming debate anyway; there are many other people working on it, and I feel I can be of more use (to whomever) elsewhere.  As I read Mike's faculty page, though, a light bulb fizzled into life at the back of my brain; I was sure I'd seen that name before.  So I went searching and found what I had dimly remembered, in the form of the name of the conservative blogger, Mark Steyn. I won't go into any more detail because that's what Google's for, but here's something you definitely won't find there(*): As well as authoring with Michael Mann, I have also authored with Mark Steyn.  We were exact high school contemporaries (although only he could tell you how he went from a grammar school in Birmingham, England to worldwide fame as Canada's leading neocon blogger), and in 1973, in what would be about the eighth grade in the U.S. system, he and I collaborated on a cartoon strip for a school magazine, about a superhero called "Mini-Man".  Mark drew the pictures and I contributed some of the "humour".  One thing I remember is that Mini-Man's height was specified very precisely; it probably wasn't 2.9013 inches, but it was something rather close to that.

So yeah, it's a really small world.



(*) Until about an hour after this blog post appears, of course.

16 August 2016

Misusing science to further an agenda risks harming both

Anyone who has anything to do with science will have had a conversation with someone whose attitude can be summarised as, "Huh. Scientists. What do they know? Last year they said eating butter/smoking cigarettes/injecting heroin/playing frisbee with a lump of plutonium was bad for us, now they say it's good."

Several years ago, I would tell such people that it wasn't the scientists who were the problem; rather, it was the journalists who were distorting things to get a cool story.  Then I got a bit closer to science, and I started to ask myself some questions.  It seemed like, in many cases, the scientists were not entirely innocent.  It turned out that researchers themselves, or their institutions' press departments, will often spin a piece of research into a cute story; in some cases, I suspect that the press release is written even before the first participant is recruited.

But this particular story takes me back to the old days.  Terrible reporting of an innocent study, just to fill column inches (or, more likely these days, to provoke clicks).

The study in question is Market Signals: Evidence on the Determinants and Consequences of School Choice from a Citywide Lottery, by Steven Glazerman and Dallas Dotter.  (You can download the full article as a PDF file from the page I linked to.)  The authors examined the behaviour of parents whose children were about to enter, or change school within, the school system of Washington, DC.  Basically, not everybody can get to go to their first choice of school, so parents rank a selection of schools in descending order of preference, and then a computer tries to assign as many people as possible to a choice that is as high on their list as possible.

Parents didn't give reasons for their choice of rank ordering, but Glazerman and Dotter reasoned that it might be possible to examine their choices and see what factors were influencing them.  For example, it seems reasonable that the further a school is from your home, the less likely you are going to be to want to send your child there, all other things being equal.  On the other hand, if there's a good bus service, that might offset the distance factor, perhaps especially for older kids who can ride the bus on their own.

These kinds of studies can often provide useful information for people who are planning educational and other resources.  Indeed, Glazerman and Dotter were interested in seeing what factors actually drive parental preference for schools, as a way to help school systems plan where to put their schools, how large to make them, etc.  For example, if they were to discover that distance actually has a very small effect if there is a good bus service, that might allow planners to feel better about moving a school to a greenfield site some way removed from where people live, and provide extra buses, rather than trying to expand the school in a limited space in its current location.  It's all very wonkish, numerical stuff --- indeed, the article comes from an organisation called "Mathematica Policy Research".

Now, one of the factors that Glazerman and Dotter examined was ethnicity (or race, or whatever you want to call it).  In the study, parents and their children were categorised as "White", "African American", or "Hispanic".  (For the purposes of this post, I'll ignore awkward questions about mixed-race families, or indeed the meaning of race and ethnicity; this post isn't really about that, although of course as a white person I have my own baggage here.)  Also, data were available on the ethnic mix of the children already attending each school.  So one of the factors that the authors were able to tease out from their data was the extent to which the proportion of students of ethnicity X in a school affected the preference of parents of ethnicity X for that school.


I've taken the liberty of reproducing Table 7 from the article here (apologies to people reading this on a mobile device).  To see how the model works, look at the first section, "Convenience", and the first line within that, "Distance (miles)".  For each of three school age ranges, and for each of three ethnicities, there is a number showing the effect of distance from home to school on parents' likelihood of choosing any given school.  All of the numbers are negative, which means that the model appears to be working: A greater distance has a negative effect on your willingness to choose that school.  And as a bonus, this effect is larger for elementary school, which makes sense (to me, anyway) --- it's more important that your smaller kids' elementary school is closer to your home than their big siblings' high school.  (The actual numbers in the table are standardised, so they don't have any meaning outside the table; just remember that bigger numbers mean a stronger positive or negative preference.)

Now look at the section entitled "School Demographics".  It gets a little complicated here because the authors found that a quadratic relation between demographics and likelihood of choosing provided a slightly better fit to the data, but basically, the same rules hold: A positive number means a preference for the same ethnicity, and a larger number means a stronger preference.  The quadratic terms are not very large, so for the purposes of this post, we can look at just the first line in this section, "Own-race percentage/10". In contrast to home-to-school distance, the results for ethnicity are not very consistent.  For White parents, there is a coefficient of 0.109 (i.e., an apparent preference) for a larger number of White students in their kids' elementary school, and the stars next to this value mean that it is statistically significant, suggesting that there was little variability among parents on this measure.  On the other hand, African American parents have a statistically significant coefficient of 0.188 for their preference for seeing more students of the same ethnicity in middle school, and for Hispanic parents, the coefficient for their preference for more Hispanic kids in high school is even higher at 0.485.

These numbers don't immediately seem to make a lot of sense to me.  Maybe there are some other factor driving them.  Remember, parents didn't explicitly state "I want my kid to go to a school with lots of people who look like him/her"; this was inferred from their expressed preferences of school, and the ethnic makeup of that school.  It might be that there are other factors driving these choices that the authors didn't (or couldn't) measure, or it could be that there is a lot of noise in their model.  The article is only a "Working Paper", meaning it hasn't been published in a peer-reviewed academic journal yet.

However, here's how this was written up in Slate by Dana Goldstein: "One Reason School Segregation Persists: White parents want it that way."  I encourage you to read that piece after first reading Glazerman and Dotter's carefully-written study.  The Slate article is a collection of cherry-picked items designed to support an agenda.  Here's the cherry-picking in full:
Across race and class, a middle-school parent was 12 percent more likely to choose a school where his child’s race made up 20 percent of the study body, compared with a school with similar test scores where his child’s race made up only 10 percent of the study body. White and higher-income applicants had the strongest preferences for their children to remain in-group, while black elementary school parents were essentially “indifferent” to a school’s racial makeup, the researchers found. The findings for Hispanic elementary and middle school parents were not statistically significant.
Let's unpack that.  The first statement doesn't tell us anything about ethnic bias, other than the rather unsurprising news that parents of all races would apparently slightly prefer their kids to be in a 20% minority versus a 10% minority.  (After all, Everyone's a Little Bit Racist.) The second sentence is a masterpiece of careful drafting.  First, note "White and higher-income applicants".  Everyone knows that White people tend to have higher incomes, so this is just rhetorical double-dipping, hiding the fact that higher-income African American and Hispanic parents also had a preference for their child to "remain in-group".  That might tell us something about well-off people (perhaps a follow-up article is in the works, telling us about the evils of rich, as opposed to White, people), but it's utterly irrelevant to the claims that this phenomenon is being driven by White people's prejudices.  Second, did you spot that "black elementary school parents were essentially 'indifferent' to a school’s racial makeup"?  That's indeed what the data show.  But Goldstein chose not to tell us that African American parents were apparently very concerned about the racial makeup of middle schools.  And finally, look at the last sentence.  It's also true, but it omits the fact that the coefficient of ethnic preference for Hispanic parents of high school students was statistically significant (and large).  But the net result is clear: The scene is set for the author to tear into the barely-unconscious sins of (only) White parents.

Perspective is everything.  Back in the Cold War, there was a joke that went like this:  The American ambassador to the United Nations challenged the Soviet ambassador to a running race.  The New York Times reported the result: "U.S. ambassador beats Soviet ambassador". Pravda reported: "Soviet ambassador finishes heroic second in race; U.S. ambassador next to last". 

So, let's get some perspective here.  These parents are residents of Washington DC, a city that is 48% Black and 44% White; probably one of the most ethnically mixed cities in the United States, I'm guessing.  It's surrounded by the leafy suburbs of Maryland and northern Virginia, which, from what I've seen on tourist visits to those areas, is where a lot of White people who commute to work in DC tend to live; and they were not part of Glazerman and Dotter's study, which covered District of Columbia residents only.  Those White people who have not become part of the "white flight" to the suburbs are, I suggest, likely to be pretty tolerant of people from other ethnicities.  Indeed, Glazerman and Dotter's results suggest that the percentage of White students at which the attractiveness of ethnic similarity for a middle school peaked was just 26% (i.e., less White than the city as a whole).  This does not suggest some kind of supremacist attitude towards the fellow students of these parents' 11-14 year old children.  (My bet, for what it's worth, is that noise is the best explanation of a lot of these findings, but I'm not here to critique Glazerman and Dotter's study, which I found interesting and informative.)

This could get political, and I don't want it to.  Racism is a bad thing, and mixing ethnicities in schools seems to me to be a good idea.  But journalists with an agenda to find bad things happening ought not to cherry-pick scientific reports in which those bad things have not, in fact, been discovered.  It provides ammunition for the kind of people who use words like "libtard" on social media, and it does a disservice to those who are very likely not part of the problem.  There are any number of other sources of racial disharmony that it would be much more productive to investigate.

I asked Steve Glazerman, one of the authors of the study, for a comment on this.  He replied: "Misinterpretation is an occupational hazard that we occasionally face as researchers”.  Science, especially social science, has plenty of problems right now.  In its efforts to get away from confirmation bias, it doesn't need lazy journalism, demonstrating exactly the same bias, to create false narratives with potentially damaging consequence for public policy.

Dana Goldstein concluded her article with "Because research—and history—show that left to their own devices, parents won’t desegregate schools."  I can't comment on the "history" part of that, although I suspect that it's true, albeit complicated.  But this research says no such thing.  Falsely adopting the legitimacy conferred by "SCIENCE" is dangerous, no matter how well-meant one's agenda might be.

04 July 2016

Old stereotypes

For an assortment of reasons, I found myself reading this article one day: This Old Stereotype: The Pervasiveness and Persistence of the Elderly Stereotype by Amy J.C. Cuddy, Michael I. Norton, and Susan T. Fiske (Journal of Social Issues, 2005).

The premise was (roughly) that elderly people are stereotyped as "warmer" to the extent that they are also perceived as incompetent (as in "Grandma's adorable, but she is a bit doddery").  The authors wrote:

We might expect a competent elderly person to be seen as less warm than a reassuringly incompetent elderly person. The open question is whether this predicted loss of warmth is offset by increases in perceived competence, or whether efforts to gain competence may backfire, decreasing rated warmth without corresponding benefits in competence(*).

The experimental scenario was fairly simple.  There were 55 participants in three conditions.  In the Control condition, participants read a neutral story about an elderly man, named George.  In the High Incompetence (hereafter, just High) condition, the story had extra information suggesting George was rather forgetful.  In the Low Incompetence (hereafter, just Low) condition, by contrast, the story had extra information suggesting George had a pretty good memory for his age.   The dependent variable was a rating of how warmly participants felt towards George: whether they thought he was warm, friendly, and good-natured.  Each of those was measured on a 1-9 scale.

Here is the results section:

Let's see.  The three warmth ratings were averaged, and then a one-way ANOVA was performed.  This was statistically significant, but of course that doesn't tell us exactly where the differences are coming from.  You might expect to see this investigated with standard ANOVA post-hoc tests (such as Tukey's HSD), but in this case, the authors apparently chose to report simple t tests --- "Paired comparisons" (**) --- comparing the groups.  Between High and Low, the t value was reported as 5.03, and between High and Control, it was 11.14.  These values are always going to be statistically significant; for 5.03 with 35 dfs this is a p of around .00001 and for 11.14 with 34 dfs, the p value is bordering on the homeopathic, certainly far below .00000001.

Hold on a minute.  The overall 3x1 ANOVA was just about significant at p < .03, but two of the three possible t tests were slam-dunk certainties?  That doesn't feel right.

Let's plug those means and SDs into a t test calculator.  There are several available online (e.g., this one), or you can build your own in a few seconds with Excel: put the means in A1 and B1, the Ns in C1 and D1, the SDs in E1 and F1, and then put this formula in G1:
  =(A1-B1)/SQRT((E1*E1/C1)+(F1*F1/D1))
(That just gives you the Student's t statistic; adding p values is left as an exercise for the reader, as is the extension to Welch's t test.)

Before we can run our t test, though, we need the sizes of each sample.  We know that nHigh + nLow + nControl equals 55.  Also, the t test for High/Low had 35 dfs, meaning nHigh + nLow equals 37, and the t test for High/Control had 34 dfs, meaning nHigh + nControl equals 36.  Putting those together gives us 18 for nHigh, 19 for nLow, and 18 for nControl.

OK, now we can do our calculations.  Here's what we get:
High/Low: t(35) = 1.7961, p = .0811
High/Control: t(34) = 3.2874, p = .0024
Low/Control: t(35) = 0.7185, p = .4772 (just for completeness)

So there is no statistically significant difference between the High and Low conditions.  And, while the High/Control comparison is significant, its strength is far less than what was reported. If you ran this experiment, you might conclude that the intervention was maybe doing something, but it's not clear what.  Certainly, the authors' conclusions seem to need substantial revision.

But wait... there's more.  (Alert readers will recognise some of the ideas in what follows from our GRIM preprint).

Remember our sample sizes: nHigh = 18, nLow = 19, nControl = 18.  And the measure of warmth was the means of three items on a 1-9 scale.  So the possible total warmth scores across the 18 or 19 participants, when you add up the three-item means, were (18.000, 18.333, 18.666, ..., 161.666, 162.000) for High and Control, and (18.000, 18.333, 18.666, ..., 170.666, 171.000) for Low.

Now, the mean of the High scores was reported as 7.47.  Multiply that by 18 and you get 134.46.  Of course, 7.47 was probably rounded, so we need to look at what it could have been rounded from.  The candidate total scores either side of 134.46 are 134.333 and 134.666.  But when you divide 134.333 (recurring) by 18, you get 7.46296, which rounds (and truncates) to 7.46, not 7.47.  And when you divide 134.666 (recurring) by 18, you get 7.48148, which rounds (and truncates) to 7.48, not 7.47.

Let's look at the Low scores.  The mean was reported as 6.85.  Multiply that by 19 and you get 130.15.  Candidate total scores in that range are 130.000 and 130.333.  But when you divide 130.000 by 19, you get 6.84211, which rounds (and truncates) to 6.84, not 6.85.  And when you divide 130.333 (recurring) by 19, you get 6.85956, which rounds to 6.86.  (It could be truncated to 6.85 if you really weren't paying attention, I suppose.)

For completeness, the Control mean of 6.59 is possible: 6.59 times 18 is 118.62, and 118.666 divided by 18 is 6.59259, which rounds and truncates to 6.59.

So this means that, given the dfs as they are reported in Cuddy et al.'s article, the two means corresponding to the experimentally manipulated conditions are necessarily incorrect.

A possible solution that allows the means to work is if the dfs of the second t test were misreported.  If you change t(35) to t(34), that implies nHigh = 19, nLow = 18, nControl = 18, and now the means can be computed correctly.  But one way or another, there's yet more uncertainty here.

To summarise, either:
/a/ Both of the t statistics, both of the p values, and one of the dfs in the sentence about paired comparisons is wrong;
or
/b/ "only" the t statistics and p values in that sentence are wrong, and the means on which they are based are wrong.

And yet, the sentence about paired comparisons is pretty much the only evidence for the authors' purported effect.  Try removing that sentence from the Results section and see if you're impressed by their findings, especially if you know that the means that went into the first ANOVA are possibly wrong too.

As of today, Cuddy et al.'s article has 523 citations, according to Google Scholar; yet, presumably, none of the people citing it, nor indeed the reviewers, can have actually read it very carefully.  So I guess some of the old stereotypes are true, at least when it comes to what people say about social psychology.

(*) Note that the study design arguably did not really test any efforts by the elderly person to gain competence; it tested how participants reacted to descriptions of the person's competence by a third party, which is not quite the same thing.

(**) I presume that the term "paired comparisons" refers to the fact that the comparison was between a pair of groups in each case, e.g., High/Low or High/Control.  The authors can't have performed a paired samples t test, since the samples were independent.

[Update 2016-07-04 13:32 UTC: Thanks to Simon Columbus for his comment, pointing out the PubPeer thread on this article.  Apparently a correction has been drafted (or maybe published already?) that fixed the t values, and then claims, utterly bizarrely, that this does not change the conclusion of the paper.  But even if we accept that for a nanosecond, it does not address the question of why the means were not correctly reported.  It looks like a second correction may be in order.  I wonder what Lady Bracknell would say?]

[Update 2016-07-09 22:17 UTC: Fixed an error; see comment by John Bullock.]

14 December 2015

My (current) position on the PACE trial

I have written this post principally for people who have started following me (formally on Twitter, or in some other way) because of my somewhat peripheral involvement on the PACE trial discussions.

First off, while I try to be reasonably politically correct, I don't always get all the details right.  I've tried to be respectful to all involved here.  In particular, someone told me that "CFS/ME" is not always an appropriate label to use.  I hope anyone who thinks that will allow me a pass on that, from my position of ignorance.

I've learned a lot about CFS/ME over the past few weeks.  Some of what I've been told --- but above all, what I've observed --- about how some of the science has been conducted, has disturbed me.  The people whose opinions I tend to trust on most issues, who usually put science ahead of their personal political position, seem to be pretty much unanimous that the PACE trial data need to be released so that disinterested parties can examine them.

But I want to make it clear that I have no specific interest in CFS/ME.  I don't personally know anyone who suffers from it, and it's not something I've really ever thought about much.  I don't especially want to become an advocate for patients, except to the extent that, having had my own health problems in the last couple of years, I wish every sick person a speedy recovery and access to the finest medical treatment they can get.  So I'm not sure I can even call myself an "ally"; allies have to take a non-trivial position, and I don't think my position here is much more than trivial.  If the PACE trial data emerge tomorrow, I will not personally be reanalysing them.  I don't know enough about this kind of study to do so.

What I do care about is the integrity of science.  You can see this, I hope, if you Google some of the stuff I've been doing in psychology.  Science, imperfect though it is, is about the only rational game in town when it comes to solving the problems facing society, and when scientists put their own interests above those of the wider community, it usually doesn't turn out well.

So, on to the PACE trial... I want to say that I can understand a lot of defensiveness on the part of the PACE researchers.  They have heard stories of others being harassed and even receiving death threats.  Maybe some of them have experienced this themselves.  For the purposes of this post (please bear with me!), I'm going to assume --- because I have no evidence to the contrary, and people generally don't make these accusations lightly --- that the stories of CFS/ME researchers being harassed in the past are true; arguably, for the purposes of this discussion, it doesn't make any difference whether they are true or not.  (Of course, in another context, such claims are very important, but let me try to put that aside for now.)  Apart from anything else, given the size of the CFS/ME community, it would be unreasonable not to expect there to be some fairly unpleasant people to have also developed the condition.  We all know people like that, whatever our and their health status.  CFS/ME strikes people from all walks of life, including some saints and some sinners.

Now, with that said, I am unconvinced --- actually, "bewildered" would be a better word --- by the argument that releasing the data would somehow expose the researchers to (further) harassment.  Indeed, it seems to me that withholding the data plays directly into the hands of those who claim that the PACE researchers have "something to hide", and they are presumably the most likely to escalate their anger into harassment.  I actually don't believe that the researchers have anything to hide, in the sense of feeling guilty because they did something bad in their analyses.  I've seen enough cases like this in my working life to know that incompetence --- generally in the form of a misplaced sense of loyalty to a group rather than to the wider truth and public interest --- is always to be preferred as an alternative explanation to malice, first because malice is harder to prove, and second because it just almost always turns out to be the case than incompetence was behind a screw-up.

About the only reason I can sort of imagine for the argument that releasing the data might lead to harassment of the researchers, is if the alternative were for the question to somehow go away.  That's perhaps a reasonable argument with some political issues; for example, there is (I think) a legitimate debate to be had over whether it's helpful to reproduce, say, cartoons that might cause people to get over-excited, when they could just be left to one side.  But that's simply not going to happen here.  People with a chronic, debilitating condition, and no cure in sight, are not going to suddenly forget tomorrow that they have that condition.  So far, none of the replies to people who have asked for the data, and been told it will lead to harassment, have explained the mechanism by which that is supposed to happen.

The researchers' argument also seems to conflate the presence in the CFS/ME activist community of some unpleasant people --- which, again, for the sake of this discussion, I will assume is probably true --- to the idea that "anyone from the CFS/ME activist community who asks about PACE is probably trying to harass us".  This is not good logic.  It's what leads airline passengers to demand that Muslim passengers be thrown off their plane.  It's called the base rate fallacy, and avoiding it is supposed to be what scientists --- particularly, for goodness sake, scientists involved in epidemiology --- are good at.

A further problem with the arguments that a request for the data --- whether it comes from patients with scientific training, or scientists such as Jim Coyne --- is designed to be "vexatious" or to "lack serious purpose" or that its intent is "polemical" (all terms used by King's in their reply to Coyne), is that such arguments are utterly unfalsifiable.  Given the public profile of this matter, essentially anyone who asks for the data is going to have their credentials examined, and unless they meet the unspecified high standards of the researchers, they won't get to see the data.  (Yes, Jim Coyne --- who, full disclosure, is my PhD supervisor --- can be a bit shouty at times.  But this is not kindergarten.  Scientists don't get to withhold data from other scientists just because they don't play nice.  Ask any scientist if science is about robust disagreement and you will get a "Yes", but if that idealism isn't maintained when actual robust disagreement takes place, then we might as well conduct the whole process through everything-is-fine press releases.)

Actually, in their reply to Coyne, King's College did seem to give a hint as to who might be allowed to see the data, in their statement "We would expect any replication of data to be carried out by a trained Health Economist", with an nice piece of innuendo carried over from the preceding sentence that this health economist had better have a lot of free time, because the original analysis took a year to complete.  This suggests that unless you declare your qualifications as an unemployed health economist, you aren't going to be judged worthy to see the data (and if you come up with conclusions after a week, it might well be suggested that you didn't look hard enough). But the idea that it will take a year, or indeed need specialised training in health economics, to determine whether the Fisher's exact tests from the contingency tables were calculated correctly, or whether the results really show that people got better over the course of the study, is absurd.  Apart from anything else, science is about communicating your results in a coherent manner to the rest of the scientific community.  If you submit an article and then claim that its principal conclusions cannot be verified except by a few dozen highly trained specialists with a year's effort, that's an admission right there that your article has failed.  Of course there will be questions of interpretation, over things like what "getting better" means, but nobody should have to accept the researcher's claims that their interpretation is the right one.  There needs to be a debate, so that a consensus, if one is possible, can emerge.  (Who knows?  Maybe the evidence for CBT is overwhelming.  There are plenty of neutral scientists who can reach a fair conclusion about that, but right now, they are being deprived of the opportunity to do so.)

A further point about the failure to share data is that the researchers agreed, when they published in PLoS ONE, to make their data available to anyone who asked for it.  This is a condition of publishing in that journal.  You can't have the cake of "we're transparent, we published in an open access journal" and then eat that cake too with "but you can't see the data".  PLoS ONE must insist that the authors release the data as they agreed to do as a condition of publication, or else retract the article because their conditions of publication have been breached.  See Klaas van Dijk's formal request in this regard.

These data are undoubtedly going to come out at some point anyway.  The UK's Information Commissioner will see to that, even if PLoS ONE doesn't persuade the authors to release the data.  As the risk management specialist Peter Sandman points out, openness and transparency at the earliest possible stage translate into reduced pain and costs further down the line.

I want to end with a small apology.  I wrote a post yesterday on an unrelated topic (OK, it was also critical of some poor science, but the relation with the subject of this post was peripheral).  Two people submitted comments on that post which drew a link with the PACE trial.  After some thought, I decided not to publish those comments, as I wanted to keep discussion on that other post on-topic.  I apologise to the authors of those comments that Blogger.com's moderation system did not let me explain the reasons why they were not published.  I would happily publish those same comments on this post; indeed, I will publish pretty much any reasonable comments on this post.

13 December 2015

Digging further into the Bos and Cuddy study

*** Post updated 2015-12-19 20:00 UTC
*** See end of post for a solution that matches the reported percentages and chi-squares.
 A few days ago, I blogged about Professor Amy Cuddy's op-ed piece in the New York Times, in which she cited a non-published, non-peer reviewed study about "iPosture" by Bos and Cuddy of how people allegedly deferred more to authority when they used smaller (versus larger) computing devices, because using smaller devices caused them to hunch (sorry, "iHunch") more, and then something something assertiveness something something testosterone and cortisol something.  (The authors apparently didn't do anything as radical at to actually measure, or even observe, how much people hunched, if at all; they took it for granted that "smaller device = bigger iHunch", so that the only possible explanation for the behaviours they observed was the one they hypothesized.  As I noted in that other post, things are so much easier if you bypass peer review.)

Just for fun, I thought I'd try and reconstruct the contingency tables for "people staying on until the experimenter came and asked them to leave the room" from the Bos and Cuddy article, mainly because I wanted to make my own estimate of the effect size.  Bos and Cuddy reported this as "[eta] = .374", but I wanted to experiment with other ways of measuring it.

In their Figure 1, which I have taken the liberty of reproducing below (I believe that this is fair use, according to Harvard's Open Access Policy, which is to be found here), Bos and Cuddy reported (using the dark grey bars) the percentage of participants who left the room to go and collect their pay, before the experimenter returned.  Those figures are 50%, 71%, 88%, and 94%.  The authors didn't specify how many participants were in each condition, but they had 75 people and 4 conditions (phone, tablet, laptop, desktop), and they stated that they randomised each participant to one condition.  So you would expect to find three groups of 19 participants and one of 18.



However, it all gets a bit complicated here.  It's not possible to obtain all four of the percentages that were reported (50%, 71%, 88%, and 94%), rounded conventionally, from a whole number of participants out of 18 or 19.  Specifically, you can take 9 out of 18 and get 50%, or you can take 17 out of 18 and get 94% (0.9444, rounded down), but you can't get 71% or 88%, with either 18 or 19 as the cell size.  So that suggests that the groups must have been of uneven size.  I enumerated all the possible combinations of four cell sizes from 13 to 25 which added up to 75 and also allowed for the percentages of participants who left the room, correctly rounded, to be one of the integers we're looking for.  Here they those possible combinations, with the total numbers of participants first and the percentage and number of leavers in parentheses:

14 (50%=7), 21 (71%=15), 24 (88%=21), 16 (94%=15)
18 (50%=9), 24 (71%=17), 16 (88%=14), 17 (94%=16)
20 (50%=10), 21 (71%=15), 16 (88%=14), 18 (94%=17)
20 (50%=10), 14 (71%=10), 24 (88%=21), 17 (94%=16)
22 (50%=7), 21 (71%=15), 16 (88%=14), 16 (94%=15)

Well, I guess that's also "randomised" in a sense.  But if your sample sizes are uneven like this, and you don't report it, you're not helping people to understand your experiment.

But maybe they still round their numbers by hand at Harvard for some reason, and sometimes they make mistakes.  So let's see if we can get to within one point of those percentages (49% or 51% instead of 50%, 70% or 72% instead of 71%, etc).  And it turns out that we can, just, as shown in the figure below, in which yellow cells are accurately-reported percentages, and orange cells are "off by one".  We can take 72% for N=18 instead of 71%, and 89% for N=19 instead of 88%.  But then, we only have a sample size of 73.  So we could allow another error, replacing 94% for N=18 with 95% for N=19, and get up to a sample of 74.  Still not right.  So, even allowing for three of their four percentages to be misreported, the per-cell sample sizes must have been unequal.



However, if I was going to succeed in my original aim of reconstructing plausible contingency tables, there would be too many combinations to enumerate if I included these "off-by-one" percentages.  So I went back to the five possible combinations of numbers that didn't involve a reporting error in the percentages, and computed the chi-square values for the contingency tables implied by those numbers, using the online calculator here.  They came out between 10.26 and 12.37, with p values from .016 to .006; this range brackets the numbers reported by Bos and Cuddy (chi-square 11.03, p = .012), but none of them matches those values exactly; the closest is the last set (22, 21, 16, 16) with a chi-square of 11.22 and a p of .011.

So, I'm going to tentatively presume that in fact the sample sizes were all equal (give or take one for not having a number of participants divisible by four), and it's in fact the percentages on the dark grey bars in Bos and Cuddy's Figure 1 that are wrong.  For example, if I build this contingency table:

Leavers
9 14 16 18
Stayers
9 5 3 1
% Leavers 50% 74% 84% 95%

then the sample size adds up to 75, the per-condition sample sizes are equal, and the chi-square is 11.086 and the p value is .0113.  That was the closest I could get to the values of 11.03 and .012 in the article, although of course I could have missed something.  These numbers are close enough, I guess, although I'm not sure if I'd want to get on an aircraft built with this degree of attention to detail; we still have inaccuracies in three of the four percentages as well as the approximate chi-square statistic and p value.

Normally in circumstances like this, I'd think about leaving a comment on the article on PubPeer.  But it seems that, in bypassing the normal academic publishing process, Professor Cuddy has found a brilliant way of avoiding, not just regular peer review, but post-publication peer review as well.  In fact, unless the New York Times directs its readers to my blog (or another critical review) for some reason, Bos and Cuddy's study is impregnable by virtue of not existing in the literature.


PS:  This tweet, about the NY Times article, makes an excellent point:
Presumably we should all adopt the wide, expansive pose of the broadsheet newspaper reader. Come to think of it, in much of the English-speaking world at least, broadsheets are typically associated with higher status than tabloids.  Psychologists! I've got a study for you...

PPS: The implications of the light grey bars, showing the mean time taken to leave the room by those who didn't stay for the full 10 minutes, are left as an exercise for the reader.  In the absence of standard deviations (unless someone wants to reconstruct possible values for those from the ANOVA), perhaps we can't say very much, but it's interesting to try and construct numbers that match those means.


*** Update 2015-12-19 20:00 UTC: An alert reader has pointed out that there is another possible assignment of subjects to the conditions:
16 (50%=8), 24 (71%=17), 17 (88%=15), 18 (94%=17)
This gives the Chi-square of 11.03 and p of .012 reported in the article.
So I guess my only remaining complaint (apart from the fact that the article is being used to sell a book without having undergone peer review) is that the uneven cell sizes per condition was not reported.  This is actually a surprisingly common problem, even in the published literature.


A cute story to be told, and self-help books to be sold - so who needs fuddy-duddy peer review?

Daniel Kahneman's warning of a looming train wreck in social psychology took another step closer towards realisation today with the publication of this opinion piece in the New York Times.

In the article, entitled "Your iPhone Is Ruining Your Posture — and Your Mood", Professor Amy Cuddy of Harvard Business School reports on "preliminary research" (available here) that she performed with her colleague, Maarten Bos.  Basically, they gave some students some Apple gadgets to play with, ranging in size from an iPhone up to a full-size desktop computer.  The experimenter gave the participants some filler tasks, and then left, telling them that s/he would be back in five minutes to debrief and pay them, but that they could also come and get him/her at the desk outside.  S/he then didn't come back after five minutes as announced, but instead waited ten minutes.  The main outcome variable was whether the participants came to get their money, and if they did how long they waited before doing so, as a function of the size of the device that they had.  This was portrayed as a measure of their assertiveness, or lack thereof.

It turned out that, the smaller the device, the longer they waited, thus showing reduced assertiveness.  The authors' conclusion was that this was caused by the fact that, to use a smaller device, participants had to slouch over more.  The authors even have a cute name for this: the "iHunch".  And — drumroll please, here's the social priming bit — the fact that the participants with smaller devices were hunched over more made them more submissive to authority, which made them more reluctant to go and tell the researcher that they were ready to get paid their $10 participation fee and go home.

It's hard to know where to begin with this.  There are other plausible explanations, starting with the fact that a lot of people don't have an iPhone and might well enjoy playing with one compared to their Android phone, whereas a desktop computer is still just a desktop computer, even if it is a Mac.  And the effect size was pretty large: the partial eta-squared of the headline result is .177, which should be compared to Cohen's (1988) description of a partial eta-squared of .14 as a "large" effect.  Oh, and there were 75 participants in four conditions, making a princely 19 per cell.  In other words, all the usual suspect things about priming studies.

But what I find really annoying here is that we've gone straight from "preliminary research" to the New York Times without any of those awkward little academic niceties such as "peer review".  The article, in "working paper" form (1,000 words) is here; check out the date (May 2013) and ask yourself why this is suddenly front-page news when, after 30 months, the authors don't seem to have had time to write a proper article and send it to a journal, although one of them did have time to write 845 words for an editorial in the New York Times.  But perhaps those 845 words didn't all have to be written from scratch, because — oh my, surprise surprise — Professor Cuddy is "the author of the forthcoming book 'Presence: Bringing Your Boldest Self to Your Biggest Challenges.'"  Anyone care to take a guess as to whether this research will appear in that book, and whether its status as an unreviewed working paper will be prominently flagged up?

If this is the future — writing up your study pro forma and getting it into what is arguably the world's leading newspaper, complete with cute message that will appeal to anyone who thinks that everybody else uses their smartphone too much — then maybe we should just bring on the train wreck now.




*** Update 2015-12-17 09:50 UTC: I added a follow-up post here. ***


Reference
Cohen, J. (1988). Statistical power analysis for the behavioral sciences (2nd ed.). Hillsdale, NJ: Lawrence Erlbaum Associates.