Showing posts with label publish or perish. Show all posts
Showing posts with label publish or perish. Show all posts

Friday, January 5, 2018

Teamwork and the Reproducibility Problem

It has been known for some time that psychology has a reproducibility problem, though we may not always agree on how to handle or discuss these issues. I remember chatting with another researcher at a conference shortly after I finished my masters thesis on stereotype threat and its impact on math performance in women. I had failed to replicate stereotype threat effects in my study. She, on the other hand, said her effects were incredibly strong; she described a participant experiencing a panic attack when she was told she had to do math problems, and had even noticed her female participants' math performance was negatively affected when her research assistant had been knitting during the session. (I also remember a reviewer telling me I must have performed the study poorly, not because the reviewer found any flaws in my methods, but because I had failed to reproduce the stereotype threat effects in my research.)

Efforts to handle this crisis thus far have included making psychological research more transparent and large-scale meta-analyses. And a new effort is already underway to harness the power of multiple research labs across the world: the Psychological Science Accelerator. Christie Aschwanden of FiveThirtyEight has more:
[Psychologist Christopher] Chartier, a researcher at Ashland University, doesn’t think massively scaled group projects should only be the domain of physicists. So he’s starting the “Psychological Science Accelerator,” which has a simple idea behind it: Psychological studies will take place simultaneously at multiple labs around the globe. Through these collaborations, the research will produce much bigger data sets with a far more diverse pool of study subjects than if it were done in just one place.

The accelerator approach eliminates two problems that can contribute to psychology’s much-discussed reproducibility problem, the finding that some studies aren’t replicated in subsequent studies. It removes both small sample sizes and the so-called weird samples problem, which is what happens when studies rely on a very particular population — like relatively wealthy college students from Western countries — that may not represent the world at large.

So far, the project has enlisted 183 labs on six continents. The idea is to create a standing network of researchers who are available to consider and potentially take part in study proposals, Chartier said. Not every lab has to participate in any given study, but having so many teams in the network ensures that approved studies will have multiple labs conducting their research.
According to the blog, the Psychological Science Accelerator is taking on its second study, this one on gendered social category representation. And if you're attending the Association for Psychological Science meeting in May, you can check out a symposium on "Large Scale Research Collaborations: Applications in Crowd-Sourcing and Undergraduate Research Experience, Replications, and Cross-Cultural Research." (Day and time TBD - APS is still finalizing the program, and is still accepting poster submissions through the end of this month.)

Friday, December 22, 2017

Travel Day Links

I finally saw The Last Jedi last night and loved it. I'll try to have more reactions soon. For now, I'll say I loved and I'm so happy to no longer have to dodge spoilers.

I'm heading out of town for the holidays later on this morning/afternoon. I have a few articles up to read:


Happy holidays, everyone! I'm driving into cold temperatures and lots of snow, so I'm packing a ton of books and my laptop (and lots of sweaters and yoga pants), and planning to spend much of my time reading and writing. 

Wednesday, November 29, 2017

Statistical Sins: When the Data are Too Perfect

Yesterday, Ars Technica published an article about an investigation into the research of Nicolas Guéguen, a psychologist who has received a great deal of media attention for his shocking findings in gender effects. His research includes findings that men prefer women in heels or wearing red, and that men are more likely to help a woman wearing her hair down instead of up. But, according to James Heathers and Nick Brown, Guéguen's data and high publication rate are suspect:
What they've found raises a litany of questions about statistical and ethical problems. In some cases, the data is too perfectly regular or full of oddities, making it difficult to understand how it could have been generated by the experiment described by Guéguen.

Social media is where it all kicked off, when Nick Brown saw a tweet about a paper claiming that men were less likely to help a woman who had her hair tied up in a ponytail or a bun. When they looked more closely at the paper, something odd jumped out at them: the numbers in the paper looked strangely regular.

When you’re dividing by three, the decimal points will always follow this pattern: either .000, .333, or .666. If you divide by 30, the pattern just moves up a decimal place: the second decimal will always be 3 or 6.

In this study, every average score was divided by 30, because each group (male-ponytail, male-loose, female-bun, and so on) had 30 people in it. But every average number was perfectly round: 1.80, 2.80, 1.60. That’s … unlikely. “The chance of all six means ending in zero this way is 0.0014,” write Heathers and Brown in their critique.
Many of Guéguen's studies involve elaborate situations using confederates - research assistants who pretend to be a participant or random person on the street. But Guéguen publishes many single author papers without acknowledgements. When Heathers and Brown reached out to Guéguen for more information on how he could publish so many elaborate studies on his own, he explained that he supervises many student projects. But why aren't the students at least thanked in the papers? Or listed as a coauthor, as they should be if they're doing a great deal of the work?

Heathers and Brown have repeatedly reached out to Guéguen for some documentation to substantiate that these studies occurred as described, but email correspondence, ethics review committee reports, and original datasets have not been shared in many cases.

Brown will be publishing the results of his and Heathers's examination of Guéguen's work on his blog. The first critique can be found here.

As has happened before, this particular instance of alleged academic dishonesty is liable to lead to a discussion about the problems of the "publish or perish" mentality in academia and research. But, as in previous cases, it's unlikely that such a discussion will result in any real improvements to dissuade such honesty. When the benefit of publishing a great deal is high and the probability of being caught is low, these things will continue to happen. Completely fabricated research is rare and likely to remain so, but tiny slips into academic dishonesty - massaging numbers or dropping cases - will happen, even by the most honest of researchers.

Saturday, July 22, 2017

A Long Time Ago in a Journal Far Away

Predatory journals have been around for a while, but thanks to the new availability of open access options online, they're becoming a lot harder to spot. They were once known as vanity journals - you basically pay to have your article published. With new open access options, many journals - predatory and non-predatory - routinely charge fees to offset the cost of publishing. But with predatory journals, you'll start to notice other added costs, such as fees for the review process and even sometimes a slight hint that if you pay more at this stage, your paper is more likely to get a favorable review.

Needless to say predatory journals are a huge problem. While publication bias - the tendency to only publish studies with significant results - hurts the field, a journal that doesn't even bother going through peer review can also hurt the field, by allowing garbage research to proliferate.

So one researcher decided to brilliantly strike back. I mean, we all know the odds of successfully navigating the research field are... you know what, never tell me the odds. There's no need to fear - fear is the path to the Dark Side. Do or do not, there is no try. And the force is strong with this one.

That's right - this researcher wrote a Star Wars-themed research paper about midichloria, filled with plagiarized material from Wikipedia and copied and pasted movie quotes, and it's bloody brilliant:


The paper references Force sensitivity and name drops Star Wars characters, including the "Kyloren cycle" and "midichloria DNA (mtDNRey)" and "ReyTP." At one point in the article, it switches rather abruptly to the monologue about the Tragedy of Darth Plagueis the Wise:


And here's how the article fared:
Four journals fell for the sting. The American Journal of Medical and Biological Research (SciEP) accepted the paper, but asked for a $360 fee, which I didn’t pay. Amazingly, three other journals not only accepted but actually published the spoof. Here’s the paper from the International Journal of Molecular Biology: Open Access (MedCrave), Austin Journal of Pharmacology and Therapeutics (Austin) and American Research Journal of Biosciences (ARJ) I hadn’t expected this, as all those journals charge publication fees, but I never paid them a penny.

Credit where credit’s due, a number of journals rejected the paper: Journal of Translational Science (OAText); Advances in Medicine (Hindawi); Biochemistry & Physiology: Open Access (OMICS).

Two journals requested me to revise and resubmit the manuscript. At JSM Biochemistry and Molecular Biology (JSciMedCentral) both of the two peer reviewers spotted and seemingly enjoyed the Star Wars spoof, with one commenting that “The authors have neglected to add the following references: Lucas et al., 1977, Palpatine et al., 1980, and Calrissian et al., 1983”. Despite this, the journal asked me to revise and resubmit.

At the Journal of Molecular Biology and Techniques (Elyns Group), the two peer reviewers didn’t seem to get the joke, but recommended some changes such as reverting “midichlorians” back to “mitochondria.”

Finally, I should note that as a bonus, “Dr Lucas McGeorge” was sent an unsolicited invitation to serve on the editorial board of this journal.

All of the nine publishers I stung are known to send spam to academics, urging them to submit papers to their journals. I’ve personally been spammed by almost all of them. All I did, as Lucas McGeorge, was test the quality of the products being advertised.

Tuesday, May 30, 2017

Gender Bias in Political Science

This morning, the Washington Post published a summary (written by the study authors) of a study examining gender bias in publications in the top 10 political science journals.
Our data collection efforts began by acquiring the meta-data on all articles published in these 10 journals from 2000 to 2015. Web-scraping techniques allowed us to gather information on nearly 8,000 articles (7,915), including approximately 6,000 research articles (5,970). The journals vary in terms of the level of information they provide about the nature of each article, but we were generally able to determine the type of article (whether a research article, book review, or symposium contribution), the names of all authors—from which we could calculate the number of authors—and often the institutional rank of each author (for example, assistant professor, full professor, etc.). In what follows, we describe the variable generation process for all types of articles in the dataset, but note that the findings we report stem from an analysis of authorship for research articles only, and not reviews or symposia.

Using an intelligent guessing technique (compared against a hand-coding method) we used authors’ first names to code author gender for all articles in the database. We also hand-coded the dominant research method employed by each research article. We were further able to generate women among authors (%) which is the share of women among all authors published in each journal, as well as other variables related to the gender composition for each article, which include information about whether each article was written by a man working alone, a woman working alone, an all-male team, an all-female team, or a co-ed team of authors. Because the convention in political science is generally to display author names alphabetically, we have not coded categories like “first author” or “last author” which are important in the natural sciences.
As you can see from the table below, there were low percentages of women among authors across all 10 journals:


One explanation people offer for underrepresentation of women is that there are simply fewer women in the field. But that's not the case here:
Women make up 31 percent of the membership of the American Political Science Association and 40 percent of newly minted doctorates. Within the 20 largest political science PhD programs in the United States, women make up 39 percent of assistant professors and 27 percent of tenure track faculty.
Instead, they offer 2 explanations:

1) Women aren't being offered as many opportunities for coauthorship:
The most common byline across all the journals we surveyed remains a single male author (41.1 percent); the second most common form of publication is an all-male “team” of more than one author (24 percent). Cross-gender collaborations account for only 15.4 percent of publications. Women working alone byline about 17.1 percent of publications, and all-female teams take a mere 2.4 percent of all journal articles.
2) The research methods most often used by women political sciences (qualitative methods) are less likely to be published in these top journals than studies using quantitative methods. As a mixed methods researcher, I frequently use qualitative methods - this was especially true in my work for the Department of Veterans Affairs, where we studied topics that were not only complex and nuanced, but poorly studied and sometimes occurring in a small subset of the population. These are the perfect conditions for a well-done qualitative study to establish some concepts that can be studied quantitatively. But it's difficult to write a survey or create a measure without that basic knowledge. (That doesn't stop people from doing it, leading to bad research. But hey, it uses numbers, so it must be good, right? </sarcasm>) I frequently received snide remarks from other researchers and consumers of research, who didn't believe qualitative methods were rigorous or even scientific. And, as I've blogged about before, I received similar comments in some of my peer reviews.

The authors recognize that perhaps the reason for low representation of women may be because they simply aren't submitting to these journals. But:
[I]f women are not submitting to certain journals in numbers that represent the profession, this is the beginning and not the end of the story. Why not?

Political scientists have helped forge crucial insights into the “second” and “third faces” of power — ideas that help explain that the effects of power can be largely invisible.

The second face of power refers to a conscious decision not to contest an outcome in light of limited prospects for success, as when congressional seats go uncontested in districts that are solidly red or blue.

The third face of power is more subtle and refers to the internalization of biases that operate at a subconscious level, as when many people assume, without thinking, that wives — and not husbands — will adjust their careers and even their expectations to accommodate family and spouse.

Let’s apply those insights to the findings from our study. If women aren’t submitting in proportional numbers to prestigious journals, that may result from conscious decisions based on the second face of power: They don’t expect their work to be accepted because they don’t see their type of scholarship being published by those journals. Or they may refrain from submitting because of a more internalized, third-face logic, taking it for granted that scholars like “me” don’t submit to journals like that.

Either way, publication patterns are self-enforcing over time, as authors come to see it as a waste of time to submit to venues whose past publications do not include the kind of work they do or work by scholars like them.

Saturday, April 15, 2017

M is for Meta-Analysis

I've blogged before about meta-analysis (some examples here, here, and here), but haven't really gone into detail about what exactly it is. It actually straddles the line between method and analysis. Meta-analysis is a set of procedures and analyses that allow you to take multiple studies on the same topic, and aggregate their results (using different statistical techniques to combine results).

Meta-analysis draws upon many of the different concepts I've covered so far this month. Aggregating across studies increases your sample size, maximizing power and providing a better estimate of the true effect (or set of effects). It's an incredibly time-intensive process, but it is incredibly rewarding and the results are very valuable for helping to understand (and come to a consensus on) an area of research and guide future research on the topic.

First of all, you gather every study you can find on a topic, including studies you ultimately might not include. And when I say every study, I mean every study. Not just journal articles but conference presentations, doctoral dissertations, unpublished studies, etc. Some of it you can find in article databases, but some of it you have to find by reaching out to people who are knowledgeable about an area or who have research published on that topic. You'd be surprised how many of them have another study on a topic they've been unable to publish (what we call the "file drawer problem" and relatedly, "publication bias"). The search then weeding through is a pretty intensive process. It helps to have a really clear idea of what you're looking for, and what aspects of a study might result in it being dropped from the meta-analysis.

Next, you would "code" the studies on different characteristics you think might be important. That is, even if you have very narrow criteria for including a study in your meta-analysis, there are going to be differences in how the study was conducted. Maybe the intervention used was slightly different across studies. Maybe the samples were drawn from college freshmen for some studies and community-dwelling adults in others. You decide which of these characteristics are important to examine, then create a coding scheme to pull that information from the articles. To make sure your coding scheme is clear, you'd want to have another person code independently with the same scheme and see if you get the same results. (Yes, this is one of the times I used Cohen's kappa in my research.)

You would use the results of the study (the means/standard deviation, statistical analyses, etc.) to generate an effect size (or effect sizes) for the study. I'll talk more about this later, but basically an effect size allows you to take the results of the study and convert it to a standard metric. Even if the different studies you included in the meta-analysis examined the data in different ways, you can find a common metric so you can compare across studies. At this point, you might average these effect sizes together (using a weighted average - so studies with more people have more impact on the average than studies with fewer people), or you might use some of the characteristics you coded for to see if they have any impact on the effect size.

This is just an overview, of course. I could probably teach a full semester course on meta-analysis. (In fact, that's something I would love to do, since meta-analysis is one of my areas of expertise.) They're a lot of work, but also lots of fun: you get to read and code studies (don't ask me why but this is something I really enjoy doing), and you end up with tons of data to analyze (ditto). If you're interested in learning more about meta-analysis, I recommend starting with this incredible book:


It's a really straightforward, step-by-step approach to conducting a meta-analysis (giving attention to the statistical aspect but mostly focusing on the methods). For a more thorough introduction to the different statistical analyses you can conduct for meta-analysis, I highly recommend the work of Michael Borenstein.

Thursday, February 2, 2017

Strong Psychological Science in an Age of Uncertainty

In our post-truth, alternative facts America, many things are uncertain - even things that really shouldn't be. But this increased uncertainty is also present in my field, not only because of politics, but recent efforts to replicate well-known research findings that have called many "established truths" into question.

With that in mind is a well-timed article in Perspectives on Psychological Science, which asks, "What Constitutes Strong Psychological Science?" The problem he brings up in the article is one many researchers know well: the trade off between doing "sexy" cutting edge research, which may lead to insignificant, or worse yet, incorrect, results, and "safer" established research topics, which lead to more accurate but less surprising results. He proposes a third option that falls somewhere in the middle of the two:
Science is a pluralistic endeavor that should not be forced into the corset of one specific format. If science is to flourish and to achieve progress, there must be room for competing theories, methods, and different conceptions of what science is about. Symbiotic collaboration must be possible between theory-driven and phenomenon-driven research. There is no reason to disqualify or downgrade properly conducted research of any particular type.

However, for science to grow and to unfold its potential in the future, it is essential to recognize the chances and limitations of distinct types of research and to deal with many challenges in theorizing and logic of science—beyond superficial issues of data analysis. No statistical analysis can be better than the design of a study, and no research design can be better than the rationale of the underlying theory.

The future growth of psychological science calls for a change in the value hierarchy from statistics to research design and theorizing. For research to flourish and to enable strong scientific inferences, in addition to surprising and inspiring discoveries and reputable methods and models, it is essential to take the diagnosticity of empirical hypothesis tests and the a priori likelihood of underlying theories into account.
Basically, we should continue exploring new topics of study, while also conducting research that is theoretically-driven. That is, use established theory and principles to generate hypotheses about more novel phenomena or test old principles/theories in new situations/applications. This gives the research a solid footing, by drawing on prior research about that theory or principle, while also giving room for exploration.

I agree completely that a lot of research has been conducted without a theoretical grounding. One of my favorite topics, pretrial publicity, has mostly been conducted atheoretically. But when researchers, including, myself have tried to apply a particular theory to understand pretrial publicity effects, the results don't conform to the theory, even though we still see a negative impact of pretrial publicity. This put me in a really uncomfortable position when I tried to publish a meta-analysis of my results; I was told by reviewers that I needed to do more with theory (like include some) and perhaps use the aggregated data to test a particular theory or set of theories. I understood their criticisms, because a theoretical basis is something this topic really needs. At the same time, when you do a meta-analysis, you're at the mercy of what previous researchers did and the type of data they collected, which differs across studies, sometimes dramatically. This makes it really difficult to test a theory with all (or even part) of your data.

This is the main reason my meta-analysis STILL isn't published, almost 7 years after I finished it.

Monday, January 9, 2017

Gender, Co-Authors, and Attributions about Contribution

It's not very often that a researcher who publishes an article solo decides to call that out with a footnote reading "This paper is intentionally solo authored." Why did Harvard University PhD student Heather Sarsons call this out? Because her study examines gender and co-authorship among economics faculty members seeking tenure. Women are less likely to be tenured in economics departments, and Sarsons wanted to find out why that might be. It turns out that having co-authored articles on one's CV has differential impacts on tenure decisions, depending on the gender of the author as well as the gender or his/her co-authors:
To determine the impact of co-authorship, Sarsons tracked all of economics professors who came up for tenure between 1985 and 2014 at 30 top universities, all places that stress tenure candidates' research credentials. She considered various factors to control for paper and journal quality through such measures as citation indexes.

Her findings:
  • Men and women who are solo authors of most of their papers have similar rates of tenure, when factoring in measures of paper quality.
  • When men co-author papers, each such paper is associated with an increase of 8 percent in the odds of the man earning tenure. But when women co-author papers, each such paper is associated only with a 2 percent increase in the odds of earning tenure.
Sarsons argues in her paper that there is additional evidence that women and men are judged differently when they co-author papers. When women co-author papers with women, the impact of co-authored papers is similar to that for male faculty members. But when papers are co-authored with men, there is more of an impact, suggesting that review committees assume that papers written by a man and a woman reflect the work of the man more than the woman.
Part of the issue is that the convention in economics is to list article names alphabetically. So it would be interesting to see if these effects hold true in fields where authors are listed in terms of contribution/effort. Sarsons herself says more research is needed on this topic, and was hesitant to offer advice based on her findings, though she mentioned women might want to try to work with other female co-authors to ensure their efforts are being weighted properly. Of course, let's not forget that reviewers have rejected articles written by women for failing to have a male co-author.

So which do we academic ladies prefer: rock or hard place?

Sunday, December 4, 2016

Academic Dishonesty, Post-Peer Review and Debunking Research

As a researcher, publishing is a very important part of my job and ongoing career options. Though most researchers engage in research honestly, and if their results are incorrect, it's more likely due to error than malice, there are still cases in which researchers have fabricated data and even entire studies (for more, see here and here). Recently, a friend brought to my attention yet another instance of research dishonesty - a case that came to light last year, but I only learned about today. What is surprising to me, in this case, is that both the dishonest researcher and the one who debunked the research are (or were at the time) graduate students:
The exposure of one of the biggest scientific frauds in recent memory didn’t start with concerns about normally distributed data, or the test-retest reliability of feelings thermometers, or anonymous Stata output on shady message boards, or any of the other statistically complex details that would make it such a bizarre and explosive scandal. Rather, it started in the most unremarkable way possible: with a graduate student trying to figure out a money issue.
Michael LaCour, a graduate student at UCLA, talked to David Broockman, grad student at UC Berkley, about a multiphase study he performed in which canvassers were able to change respondents attitudes about gay marriage by revealing their sexual orientation. Broockman, who was fascinated by the results, set out to replicate the study and encountered the first issue: Labour's survey had included 10,000 respondents paid $100 a piece, a rather large grant for a graduate student. So Broockman approached polling firms about the study idea - most said the study they couldn't carry out such a study, and if they could, it wouldn't be feasible on the usual grants grad students could obtain.

So Broockman started talking to people - carefully, because he was informed by many, and suspected himself, that exposing another researcher could get him labeled as a troublemaker or incapable of coming up with his own research ideas. And in fact, LaCour had written the paper on the results with a well-respected political scientist at Columbia, Donald Green. Broockman even said when he described the results to others, they were surprised that the results seemed to fly in the face of previous theory and research, but dropped those arguments when they heard Green was involved. In fact, when Jon Krosnick of Stanford was contacted about the study, he said, "I see Don Green is an author. I trust him completely, so I’m no longer doubtful."

Broockman hit many snags along the way, not just because he was a busy grad student working on his own research and finishing his degree - he was cautioned about exploring these issues by nearly everyone he spoke to. An anonymous post on the poliscirumors.com laying out his suspicions was deleted. And his analyses on the distribution of the data, which looked too clean to be real, failed to uncover major issues.

But still, there were hints that something was wrong. When he messaged LaCour with questions about methodology, the answers were vague and unhelpful. A similar study Broockman conducted with fellow grad student Josh Kalla showed response rates for the first wave around 1%, even though they were offering as much money as LaCour, who reported response rates of 12%. An email to the survey research firm LaCour said he had worked with on the study showed that, not only had LaCour never worked with the firm, the person he claimed to be in contact with (and had emails from) didn't exist. Then, they hit gold: a 2012 Cooperative Campaign Analysis Project that was a perfect match for LaCour's "first wave data."
By the end of the next day, Kalla, Broockman, and Aronow had compiled their report and sent it to Green, and Green had quickly replied that unless LaCour could explain everything in it, he’d reach out to Science and request a retraction. (Broockman had decided the best plan was to take their concerns to Green instead of LaCour in order to reduce the chance that LaCour could scramble to contrive an explanation.)

After Green spoke to Vavreck, LaCour’s adviser, LaCour confessed to Green and Vavreck that he hadn’t conducted the surveys the way he had described them, though the precise nature of that conversation is unknown. Green posted his retraction request publicly on May 19, the same day Broockman, Kalla, and Aronow posted their report. That was also the day Broockman graduated. “Between the morning brunch and commencement, Josh and I kept leaving the ceremonies to work on the report,” Broockman wrote in an email.
So what happened to the grad student who was repeatedly cautioned that debunking research could be a career killer? The response he received was "uniformly positive" and, oh, by the way, he's now tenure track at Stanford University. About this issue, he says: "I think my discipline needs to answer this question: How can concerns about dishonesty in published research be brought to light in a way that protects innocent researchers and the truth — especially when it’s less egregious?” he wrote. “I don’t think there’s an easy answer. But until we have one, all of us who have had such concerns remain liars by omission."

I think many of us in the research field have witnessed activities that were questionable, perhaps even clearly unethical. But rarely are we encouraged to bring our suspicions to light, and there are certainly no safe venues to bring up concerns that may or may not be accurate. While I've never been actively discouraged from reporting ethical issues, I'm sure there are many researchers who have, like Broockman. And for many grad students and post-docs, it's more likely they are working with more seasoned faculty than other grad students, so when ethical dilemmas come up, the power dynamic may discourage them from doing the right thing. While we certainly don't want witch hunts for data that looks "too good to be true," we need to find ways to protect fellow researchers and the public from bad science and false data. Because that hurts all of us.

Saturday, October 22, 2016

How Statisticians Solve Disagreements

I'm currently taking an online course on meta-analysis, which is a set of statistical and methodological techniques that allow you to combine multiple studies on a topic and generate an estimate (or set of estimates) about the true effect. It's almost like crowd-sourcing data - you're taking advantage of all the work others have done and capitalizing on the strength of having an increased number of participants, difference treatment methods, and so on. I did a candidacy exam in grad school on meta-analysis, and have conducted one before (on pretrial publicity effects), so I know a bit about the topic. This course is devoted to using the R Statistical Package, an open-source program with powerful analysis and graphing capabilities, to conduct a meta-analysis.

For the first week, we were assigned to read up on the R packages we'll be using, as well as an article from the creator of meta-analysis, Gene Glass. I've read some of Glass's work before, but for some reason, didn't encounter this article until now, which tells the reason meta-analysis was created. In addition to wanting to contribute to the field, and have a good topic to introduce in his Presidential Address to the American Educational Research Association, it was really developed to solve a disagreement.

Glass, like many grad students, left grad school with a brand new PhD and a case of depression. He found his way into psychotherapy and was so pleased with his progress, he began studying clinical psychology and became psychotherapy's biggest fan. However, another researcher, Hans Eysenck, became psychotherapy's biggest critic, constantly arguing that any effects were merely placebo:
I found this conclusion personally threatening—it called into question not only the preoccupation of about a decade of my life but my scholarly judgment (and the wisdom of having dropped a fair chunk of change) as well. I read Eysenck's literature reviews and was impressed primarily with their arbitrariness, idiosyncrasy and high-handed dismissiveness. I wanted to take on Eysenck and show that he was wrong: psychotherapy does change lives and make them better.
Glass goes through the decisions Eysenck made in conducting his literature review on the subject, and it's easy to see why, based on these decisions, Eysenck concluded psychotherapy was ineffective - or rather, it easy to see that because Eysenck strongly believed going in that psychotherapy was ineffective, he looked for evidence that supported and ignored evidence that refuted his conclusion. First, he refused to include any research that was not published in a peer reviewed journal, even studies that have to undergo another form of peer review, such as dissertations, theses, or conference presentations. But there is much reason to believe that peer reviewed articles could be biased.

Next, he eliminated any study that didn't have a control group (a group that received no treatment). So if a study compared two forms of therapy, it was tossed out. This left only 11 studies. He then did a vote count, which involves tallying up the number of studies finding a significant difference and the number finding no significant difference. "All that Eysenck considered worth noting about an experiment was whether the differences reached significance at the .05 level. If it reached significance at only the .07 level, Eysenck classified it as showing 'no effect for psychotherapy.'"

And finally, here's the real gem: if he didn't like the outcome they used (that is, he considered it subjective), he discounted the finding, and if a study found differences for one outcome but not for a second one, he also discounted it, calling it "inconsistent." This was the case even if one of the outcomes was something that might be only show a small change due to therapy, such as GPA, versus an outcome that would show a big difference, such as a measure of symptom severity. Eysenck's review didn't even take into account effect sizes: what outcomes would show big differences after psychotherapy and what would show small difference.

And that's where meta-analysis comes in:
Looking back on it, I can almost credit Eysenck with the invention of meta-analysis by anti-thesis. By doing everything in the opposite way that he did, one would have been led straight to meta-analysis. Adopt an a posteriori attitude toward including studies in a synthesis, replace statistical significance by measures of strength of relationship or effect, and view the entire task of integration as a problem in data analysis where "studies" are quantified and the resulting data-base subjected to statistical analysis, and meta-analysis assumes its first formulation. (Thank you, Professor Eysenck.)
So the TL;DR is, how to statisticians solve disagreements? They create new statistics, and then publish pithy articles where they thank the person they disagreed with. Love. It.

Tuesday, October 4, 2016

The Five-Year Legal Battle Against Bad Science

Chronic fatigue syndrome (CFS), also known as myalgic encephalomyelitis (ME), is a neuroimmune disease that affects 1-2.5 million Americans (17 million worldwide). The symptoms are widespread, ranging from memory issues and poor sleep to pain and swollen lymph nodes. The main symptom, of course, is severe fatigue that results from any kind of exertion, which is caused by an abnormal immune response to exertion that makes it difficult for people with CFS/ME to recover. Though the cause of CFS/ME is unknown, some research suggests it can occur after a bacterial or viral infection.

The best treatment for CFS/ME? According to one study (the so-called PACE trial) published in the prestigious Lancet medical journal, it's cognitive behavioral therapy and exercise. But wait, wouldn't exercise be a really bad idea for people whose immune systems freak out at any kind of exertion (leaving some sufferers bedbound)? That's what a lot of people with CFS/ME said after the article came out and especially after the article influenced treatment recommendations from such places as the Centers for Disease Control and Prevention, Mayo Clinic, and Kaiser. And it turns out, those skeptical patients were right:
If your doctor diagnoses you with chronic fatigue syndrome, you’ll probably get two pieces of advice: Go to a psychotherapist and get some exercise. Your doctor might tell you that either of those treatments will give you a 60 percent chance of getting better and a 20 percent chance of recovering outright. After all, that’s what researchers concluded in a 2011 study published in the prestigious medical journal the Lancet, along with later analyses.

Problem is, the study was bad science. And we’re now finding out exactly how bad.

Under court order, the study’s authors for the first time released their raw data earlier this month. Patients and independent scientists collaborated to analyze it and posted their findings Wednesday on Virology Blog, a site hosted by Columbia microbiology professor Vincent Racaniello. The analysis shows that if you’re already getting standard medical care, your chances of being helped by the treatments are, at best, 10 percent. And your chances of recovery? Nearly nil.
In fact, that 10 percent number is based on a reanalysis by the original authors. The analysis by independent scientists found far worse results: 4.4 percent of exercise patients and 6.8 percent of cognitive therapy patients met the criteria for "recovered," compared to 3.1 percent of people who received neither treatment. None of these differences were statistically significant.

The issues with the study are widespread, ranging from lack of proper blinding, shifting definitions of recovery and improvement between the original protocol and the final analysis, and potentially invalid thresholds for physical functioning. Some critics even suggest the the inclusion criteria are so poorly written, there may be participants in the study who don't even have CFS/ME. As Jonathan Edwards, a professor emeritus of medicine quoted in the article, put it, "They’ve set this trial up to give the strongest possible chance of there being a placebo effect that you can imagine."

And yet, this article passed a rigorous peer review process and was published, in one of the top medical journals. The question that Lancet should be asking at the moment is "How?" Once they figure that out, they need to fix whatever problem they uncover with their system, because this seriously damages their credibility.

This is an interesting counterpoint to the arguments around Susan Fiske's attack on "methodological terrorists," which some (but not all) perceived as being an attack on anyone who dares to criticize published research. In fact, the researchers in the PACE trial claimed they had received death threats - claims that appear to have been false - and the naysayers were referred to as a "vocal minority." However, I see no issue with what the lawsuit set out to do: get the researchers to release their deidentified raw data so that independent statisticians could reanalyze them. This is what good science is all about - replicability, not only in replicating a study but replicating results from the same dataset. And when the analysis approached might be invalid, this includes using different approaches, to see how sensitive the results are to, say, the cut-offs the researchers adopted. (In fact, we refer to this as "sensitivity analysis" - does a different approach make a difference in the results?) Julie Rehmeyer, the author of the article linked above, agrees:
Watching the PACE trial saga has left me both more wary of science and more in love with it. Its misuse has inflicted damage on millions of ME/CFS patients around the world, by promoting ineffectual and possibly harmful treatments and by feeding the idea that the illness is largely psychological. At the same time, science has been the essential tool to repair the problem.

Saturday, October 1, 2016

Yelp for Academic Journals

This such a great idea, I can't believe no one has thought it before! This website lets you leave reviews for journals you've worked with. So (to name a couple different experiences I've had) whether the editor was prompt in responding to questions, or the review process took over a year and the paper kept getting lost in the shuffle, you can finally share that information with the masses. Huzzah!

Though all reviews are anonymous, they ask you to set up an account - this allows them to make sure no one misuses the system. They also provide a few basic rules:
  1. Accurately report your own experiences.
  2. Do not use anyone’s name.
  3. No links.
  4. Stay on topic.
Basically, don't be a troll and stay on target:

Monday, September 26, 2016

Now I Am Become Blog

On Thursday, I wrote about Susan Fiske's upcoming article in the APS Observer. Today, Neuroskeptic, published its own response to Fiske's article. Unlike other bloggers (such as Andrew Gelman), Neuroskeptic seems to agree with my interpretation that Fiske was not talking about just anyone who criticizes scientific research, but people who do so in an unethical manner. However, he takes things one step farther than me by demanding that Fiske name names:
We should hold the offenders accountable with reference to specific examples of their attacks. After all, these people (Fiske says) are vicious bullies who are behaving in seriously unethical ways. If so, they deserve to be exposed.

Yet Fiske doesn’t do this. She says, “I am not naming names because ad hominem smear tactics are already damaging our field.” But it’s not an ad hominem smear to point to a case of bullying or harassment and say ‘this is wrong’. On the contrary, that would be standing up for decency.

Thursday, September 22, 2016

Tastes Like the Real Thing

And in the category of "peer review is so screwed", some researchers machine-generated reviews and presented them along with actual reviews to participants, who generally could not notice a difference:
Peer review is widely viewed as an essential step for ensuring scientific quality of a work and is a cornerstone of scholarly publishing. On the other hand, the actors involved in the publishing process are often driven by incentives which may, and increasingly do, undermine the quality of published work, especially in the presence of unethical conduits.

We presented to [16] subjects a mix of genuine and machine generated reviews and we measured the ability of our proposal to actually deceive subjects judgment. The results highlight the ability of our method to produce reviews that often look credible and may subvert the decision.
God help us all.

What Has Happened Down Here is a Miscommunication

Has this ever happened to you? You go to see a movie with a friend. You love every minute of the movie, laughing at the jokes, crying when something sad happens, and cheering when the hero saves the day. The credits roll and you turn to your friend and say, "What did you think?" Your friend proceeds to trash-talk the whole movie and you find yourself thinking, "Did we watch the same film?"

That's kind of what I thought when I read a blog post a grad school classmate shared. The post was a response to a forthcoming article for the APS Observer, the magazine of the Association for Psychological Science. The article, by social psychologist Susan Fiske, deals with the new(ish) trend of criticizing psychological research in social media settings (the text of her article is provided in the blog post linked above). While her insistence that criticism of psychological research should be done either in private (i.e., peer review) or in moderated settings (e.g., letters to the editor/invited responses or discussions during conference presentations) is a bit short-sighted in my opinion, she does make a point that because it has gotten easier for people to a) get their message out there and b) get contact information for researchers, some criticisms have been little more than attacks. Attacks that are not necessarily because of issues with the validity of the research or the soundness of the methods, but because of vehement disagreement with the conclusions of the research. Although her article is short and she doesn't call anyone out by name or topic area ("because ad hominem smear tactics are already damaging our field"), she's probably talking about this:
Some [researchers] have weathered frightening vitriol and threats to their reputations. Back in 1975, US Sen. William Proxmire bestowed the first of his infamous “Golden Fleece” awards on a small federal grant given to APS William James Fellows Elaine C. Hatfield of the University of Hawaii and Ellen S. Berscheid of the University of Minnesota. Proxmire denounced their study on social justice and equity in romantic relationships as a waste of taxpayers dollars. The publicity generated threatening letters and phone calls to both scientists, and their federal funding dried up because of the stigma.

In the 1990s, renowned memory researcher and APS Past President Elizabeth F. Loftus, at the University of California, Irvine, drew considerably hostile reactions when her studies challenged people’s claims that they had uncovered — often with the help of therapists — repressed memories of abuse, molestation, and even alien abduction. Loftus even had to have armed guards accompany her to lectures after she received death threats.
Fiske talks (once again, in the general sense) about attacks that share some common elements of these extreme cases:
The destructo-critics are ignoring ethical rules of conduct because they circumvent constructive peer review: They attack the person, not just the work; they attack publicly, without quality controls; they have sent their unsolicited, unvetted attacks to tenure-review committees and public-speaking sponsors; they have implicated targets' family members and advisors.

Which is why I was completely dumbfounded when, after sharing the article in its entirety, the author of the blog post, Andrew Gelman, summed it up as follows:
In short, Fiske doesn’t like when people use social media to publish negative comments on published research. She’s implicitly following what I’ve sometimes called the research incumbency rule: that, once an article is published in some approved venue, it should be taken as truth.
Did we really just read the same article?

But it gets even weirder. Gelman begins talking about the new movement in psychological science to encourage replication of past studies, a movement that has at least created some serious doubts about the validity of past studies. He aims a lot of criticism at Proceedings of the National Academy of Sciences, and a set of articles edited by Fiske. In fact, he's done this before. To be totally honest, I agree with many of his criticisms of these papers, his concerns about the validity of studies that current researchers have failed to replicate, and even the potential errors he highlights in one of Fiske's own papers. So yes, perhaps Fiske does deserve some criticism.

Except that's not what her article is about. She isn't saying there should be no criticism; she's saying that, just as there are ethical guidelines for the proper conduct of research, there are (or should be) ethical guidelines about how to offer criticism of research. But Gelman refers to Fiske as attacking "science reformers" - the people doing replication research - when I think she's referring to ad hominem attacks. I think she would have far less issue with Gelman going through Fiske's work and picking it apart, discussing methodological and analytical errors, than she would with someone writing a blog post about how much Fiske sucks and that she should lose her position at Princeton, and hey, here's the contact information of her department chair and dean of the school, why don't you, dear reader, call them up and tell them how much you hate Fiske.

So I agree with Gelman on his criticism of some of the key studies in psychological science, and his desire for more transparency and replication - something Fiske also references in her short article. And I agree with his final conclusion:
Let me conclude with a key disagreement I have with Fiske. She prefers moderated forums where criticism is done in private. I prefer open discussion. Personally I am not a fan of Twitter, where the space limitation seems to encourge snappy, often adversarial exchanges. I like blogs, and blog comments, because we have enough space to fully explain ourselves and to give full references to what we are discussing.
I frequently do the same thing on my blog. But I take issue with his insistence that Fiske's "destructo-critics" are Gelman's "science reformers." The winds may have changed in psychological science research, but Fiske and Gelman are sailing different seas.

Monday, August 15, 2016

Open Source Publishing, Peer Review, and a What-If Scenario

On the Reviewer 2 Must Be Stopped Facebook page, someone posted an interesting scenario:


Some of the responses included references to other types of social media: e.g., "It would be like Facebook" or "It would be like blogging". But what this user is proposing is somewhat different than that. True, I could publish my research on my blog - which would in many cases preclude me from publishing it elsewhere - but you either have to know about my blog to find it or it would have to come up during a web search. Same thing with Facebook: you'd have to know me to see my posts. This scenario, on the other hand, involves putting articles from different authors in one place. So a user would just have to know about the journal to access the articles.

Of course, that doesn't make this a good scenario. Let's, for the sake of argument, say there are multiple such sites for different subjects - that deals with the problem of having articles on so many different topics that it fails to be readable. After all, to make this like a regular journal, it would need to have aims and scope: a description of what the journal is about so that authors know whether their article would be a good fit. But we already have a potential failure point - authors may be very bad at determining whether their article fits, or they may be such poor writers that they fail to show the article fits. Part of what happens during review is the editor and reviewers determine if the article is a good fit. So now you have articles on, say the strength of different concrete mixtures next to an article about college students' social media behaviors.

Obviously, if this is an online journal, people can search the articles. People who search for articles on concrete shouldn't find articles on social media behavior. But once again, we have a failure point: who checks those keywords, to make sure they accurately reflect the subject of the article? In the current state of publishing, authors do generate their own keywords, often using specific standards. Certain keywords are more likely to be searched for than others, so authors might be tempted to pad their keywords with more common headings, even headings that are only somewhat relevant, to increase the chances their article is found. Fortunately, editors can double-check those keywords and drop ones that don't fit. But with the proposed system, there is no quality control.

Two major issues, and we haven't even gotten to the posting or user comments yet. And before you say, "The post said to drop reviewers, not the editor," remember that the proposal was a publishing source where instead of review, users up/down-voted articles and left comments. If editors can decide what does and does not get published, and can control (to some extent) the content, you still have the same system as you do now, where a small number of people control the flow of information. For this to work as the poster intended, you can't really have editors.

Now for the key portion of the proposal: users get to rate and comment on articles. This is where the similarities to Facebook are strongest. What posts do you "like" on Facebook? Often, ones you agree with. You would have the same danger here: that people would up-vote the articles whose results they agree with, even if the study is methodologically flawed. When you evaluate scientific research, you have to evaluate the methods. If the methods are sound, the results are presumed to be valid, even if you disagree with them. That is, the rating system would be driven by opinion rather than scientific validity. Sure, some people would evaluate the strength of the methods to generate their rating, but their voices would probably be drowned out by ratings based purely on opinion. And if you have many non-scientists visiting the articles and giving ratings, the difference between the two would be even stronger. Once again, there is no quality control of who does the rating and whether they have the necessary knowledge, as there would be if an article is peer-reviewed.

This system would really only work if you assume that everyone using it - authors and readers - do so honestly and with the best of intentions. I try to see the best in people, individually (because the way one person will behave is an unknown), but here, we're talking about patterns of groups, which are far more predictable. As much as I want to like this idea - because peer review can be unfair and problematic in its own ways - it would likely be chaos.

What do you think, readers?

Thursday, July 14, 2016

Obama Being Awesome Again

Back in May, a work colleague published a paper in JAMA (the Journal of the American Medical Association) - this is a pretty big deal, as it is one of the most prestigious journals and probably the most prestigious in my field. In honor of this accomplishment, we threw a paJAMA party. Here's some photographic evidence from the day:


Yesterday, President Obama also had an article published in JAMA, becoming the first US President to publish a scholarly article while in office.
Basically, Obama is laying out how the next president could continue to improve health care. Obama recommends things like lowering the cost of prescription drugs and making a "public option" available for people buying health care coverage as a cheaper alternative to buying coverage from private companies.

Keep in mind the article isn't marked as peer-reviewed, though it did go through extensive fact checking and editing, Forbes reported.

"While we of course recognized the author is the president of the United States, JAMA has enormously high standards and we certainly expected the president to meet those standards," Howard Bauchner, JAMA's editor-in-chief, told Bloomberg in an interview.
I propose a nationwide paJAMA party to celebrate this accomplishment.

Monday, May 30, 2016

Follow-Up on Peer Review

I've blogged many times about the peer review process (here and here, especially) - one of the first steps in publishing scientific research in scholarly journals. While the purpose of peer review is, in part, to improve the manuscript, there can certainly be a "too many chefs" component:

Monday, May 9, 2016

John Oliver on Science in the Media

I've blogged about media representation of science before - but for the tl;dr, here's John Oliver's take:



P-hacking, which he discusses in the story, is definitely a thing. My grad school statistics professor called it "fishing." Basically, it's what happens when you run multiple statistical analyses on results, looking for something significant. My dissertation director joked about doing this (not publishing) with some data on Alcatraz inmates; the only significant relationship they found was that Alcatraz inmates were significantly more likely to be Capricorns. She then looked at me very seriously, and asked, "You're not a Capricorn, right?"

Yes, I am.

Statistical results are probabilistic; we look for results that have a low chance of happening if no real relationship exists. We usually set that value at 5%. What that means is, if I run 20 tests, one those will probably be significant by chance alone. That's less of a concern if I have pre-existing (a priori) hypotheses, based on past research and/or theory, I'm testing but even if I am testing a priori hypotheses, I should apply a correction to account for the number of tests I'm running.

The problem with p-hacking is that, not only does it involve running many tests, it also usually involves only reporting the significant results. So a reader would have no idea that a person ran potentially dozens of tests based on reading the article. Unfortunately, this is one of the negative consequences of the "publish or perish" mentality. Scientists feel so much pressure to come up with results, that they'll do things they know are questionable in order to meet their publication quotas for tenure and/or funding. And that problem compounds when journals reject articles that replicate past studies. As John Oliver says in the story, "There's no Nobel prize for fact-checking."