Tuesday, September 5, 2017

Show Your Library Some Love

September is Library Card Sign-Up Month!
This September, crimefighting DC Super Heroes, the Teen Titans, will team up with the American Library Association (ALA) to promote the value of a library card. As honorary chairs, DC’s Teen Titans will remind parents, caregivers and students that signing up for a library card is the first step towards academic achievement and lifelong learning.

The focus is really on getting teens to use their library, and help build important skills for academics and life, but anyone can join in on the fun. Recently, I picked up an alumni card from Loyola, where I attended grad school, that lets me check out books and also access electronic resources while I'm on campus. And Chicago's Harold Washington Library is one of my favorite places to write. Though I already have all the library cards I need/am eligible to get, I'll definitely be showing my libraries some love this month.

How are you planning to celebrate?

Numbers in Your Life

Last week, I took to Facebook and LinkedIn to make a small request from my friends and followers in order to get some inspiration for a writing project: What are the numbers you encounter on a regular basis? Not the numbers themselves; rather, the types of numbers they encounter regularly. I gave the example of credit score and calories, and told people a great piece of advice I learned from a friend in my dance class: There are no wrong answers in brainstorming.

I feared I'd only get a few responses and that most would consist of, "Yeah, credit score, calories." I ended up getting responses from almost 60 friends, each with great lists of numbers from their personal lives, profession, child-raising, health, and so on.

Some examples by category:
  • Finance: Credit score, account balance, interest rate, stock prices
  • Health and well-being: Weight, blood glucose, cholesterol, dosages, hours slept, number of minutes of meditation
  • Library: Dewey decimal system, number of stacks, linear feet of storage space, barcodes galore
  • Alcohol: Beverages consumed, alcohol by volume, IBUs, gravity degrees plato, blood alcohol content
  • Fitness tracking: Steps, miles, heart rate, pace, laps, reps
  • Healthcare: Number of meds passed, patient satisfaction score, Medicare diagnosis-related group
  • Childcare: Number of minutes on iPad, hours slept, diapers, formula measurements
The descriptions ranged from numbers that count things, numbers that categorize things, and numbers that fall on a continuum. More than one person commented that this was a fun brainstorming exercise. 

So now, readers, I look to you. What numbers do you encounter regularly? It doesn't have to be every day, just the types of numbers that are part of your life - whether at home, work, the gym, or the bar. And remember: There are no wrong answers in brain storming!

(Note: I've checked my comment settings so hopefully they should work for you to add your numbers below. You shouldn't have to have a Google login to comment. You're also welcome to provide your numbers as a reply to this post in Facebook groups.)


Sunday, September 3, 2017

Statistics Sunday: Everyone Loves a Log (Odds Ratio)

A couple weeks ago, I introduced the concept of the odds ratio, the odds of one outcome relative to another. Odds ratios are often used to present and understand dichotomous outcome data, and researchers using logistic regression - which like linear regression, uses one or more variables to predict an outcome, but unlike linear regression, predicts a dichotomous (not continuous) outcome - will often present results in terms of odds ratios. And odds ratios are used a lot in news stories because they're a bit easier for us to understand: e.g., people who do X are twice as likely to have this outcome than people who do Z. We're naive statisticians, with a rudimentary understanding of gambling, so we have some understanding of odds.

The thing about odds ratios - and this issue becomes more pronounced when you're working with a bunch of odds ratios - is that the distribution is not symmetrical, which creates some very interesting results when we look at the inverse odds. For instance, X being twice as likely as Y (odds ratio = 2.0) makes sense. Y being half as likely as X (odds ratio = 0.5) might not make as much sense. But they're the same. And it gets more tricky with other odds ratios. Because something with 50/50 odds will have an odds ratio of 1.0, a more likely outcome A will be greater than 1 and a less likely outcome A will be the inverse, a fraction between 0 and 1. The upper bound of an odds ratio > 1.0 is infinite (∞). And the lower bound of an odds ratio < 1.0 is also infinite, but as a fraction (1/∞) moving asymptotically toward 0.

Most people I know who work with odds ratios regularly will - instead of presenting a fraction odds ratio - simply switch the order of the variables in the analysis. But what if you're working with a bunch of odds ratios around an outcome and you know that some will be greater than 1.0 and some will be less than 1.0?

It's not an unusual situation. A logistic regression may have multiple predictors, some of which will have a negative coefficient, meaning a less likely outcome A. And if you wanted to do a meta-analysis on something with a binary outcome, your effect size will be odds ratio. Some of the analyses you would run in a meta-analysis - such as a special type of regression frequently called meta-regression - won't work so well with variables that have an asymmetric distribution. Meta-regression, which is similar to linear regression but adds an additional weight to outcomes, assumes a continuous linear outcome.

But have no fear! There's a solution: the log odds ratio. Here are the basic equations you need:


It's a really easy correction. You're simply doing a log-transform of your odds ratio - more specifically, you're taking the natural log of your odds ratio.

You can do this in Excel with =LN(oddsratio). And many statistics programs are capable of log-transformations.

In SPSS, the syntax is very similar to Excel: COMPUTE log_oddsratio = LN(oddratio). (Or, if you prefer to use the GUI, go to the Transform menu, and click Compute Variable. You can type that text directly in the box, or find the LN function in the Arithmetic function group.)

The syntax in R is simply: dataframe$log_oddsratio <- LOG(dataframe$oddsratio)

And if you're working in SQL to interact with a relational database, natural log is a mathematical function, usually LOG or LN, depending on which vendor you're using. (For instance, I just wrapped an online course where I learned PostgreSQL, for which the syntax is LN.)

Because this is probably better seen than described, I've done the following. I referred back to the 2x2 contingency table from the odds ratio post and decided to test out some different frequencies. This table has 4 cells, so I simplified things a bit. I made it so the placebo group had 50/50 odds of being in remission, so those two cells are both 250. The only thing I changed is the drug group, where I tested values of 1 in the 'in remission' group and 499 in the 'not in remission' group all the way to 499 in remission and 1 not. The odds ratios for those combinations ranged from 0.002 to 499.0 (Note, these extremes would be exceedingly rare - it's very unlikely you'll see an odds ratio over 10, let alone in the 100s. This is purely for demonstration purposes.) When you graph it, it looks like this:


When I took the natural log of that array, I had a perfectly symmetrical (though not completely linear - but close enough for the typical range of log odds ratios) -6.21 to +6.21:


The natural log gets rid of the inverse property of odds ratios less than 1.0, and sets the bounds from -∞ to +∞. Now you can analyze those values and, if you want to present summary statistics as odds ratios instead of log odds ratios, you can just convert them back. To do that, you raise the natural number, e, to the power of the log odds ratio.

The syntax is EXP and the number you're converting:

Excel: =EXP(log_oddsratio)

SPSS: COMPUTE oddsratio = EXP(log_oddsratio) or select Exp from the Arithmetic functions in the Transform->Compute Variable dialog box

R: dataframe$oddsratio<-exp(dataframe$log_oddsratio)

SQL: EXP(log_oddsratio)

So you can play too, if you'd like, here's the Excel file containing the raw data - I've left the functions in, as well as the two charts, so you can change numbers if you'd like to play around.

Meta-analysis isn't the only analysis that uses (log) odds ratios. The Rasch measurement model (the psychometric approach I use) is built on log odds ratios. That's part of the magic behind its ability to turn ordinal scales into interval scales of measurement. (More on that later.)

BTW, for anyone wondering why I named this post as I did: Hopefully I'm not the only one who remembers the great Slinky parody seen on Ren & Stimpy. (In fact, when I typed "Ren and Stimpy" into Google, it auto-completed with "log," so clearly I'm not.)

For review:

Friday, September 1, 2017

Great Minds in Statistics: Jerzy Neyman's Confidence Intervals

Wednesday, August 30th was the 80-year anniversary of the publication of Jerzy Neyman's article, Outline of a Theory of Statistical Estimation Based on the Classical Theory of Probability. The classical theory refers, in this case, to Neyman's work with E. Pearson on null hypothesis significance testing and the concepts of Type I and Type II error. But this paper was groundbreaking, not just in connecting to these concepts, but in telling how we, as statisticians and scientists, should be dealing with uncertainty in the presentation of our results.

Jerzy Neyman, photographed while he was at UC-Berkley.
By Konrad Jacobs, Erlangen, Copyright is MFO - Mathematisches Forschungsinstitut Oberwolfach, http://owpdb.mfo.de/detail?photo_id=3044, CC BY-SA 2.0 de, Link
All of this starts with the basic assumption that we are trying to estimate a population parameter, which is unknown. In theoretical work, such as Neyman's paper, this value is often represented as θ (theta). We use a sample to attempt to estimate theta; we can call that estimate T. If we're estimating a mean, we have one measure of precision already included in our analysis - our standard deviation can be thought as a measure of precision of the estimate, in that it expresses the typical variation we see in scores. And in fact, as Neyman notes in his article, prior to his introduction of confidence intervals, people would often present estimates as + or - standard deviation. But, Neyman states, it probably makes more sense conceptually to use a multiple of standard deviation, if you want to express an interval with a high likelihood of containing the actual population value.

Why? Because the standard deviation tells us the typical spread of scores (the individual units that make up the mean) but that doesn't tell us the typical (expected) spread of means. Confidence intervals allow you to do that, not just for means but for a variety of aggregate statistics.

But I'm perhaps getting ahead of myself. I dug into Neyman's paper to try to summarize it for you. I've only read Neyman's work summarized in the past. In Fisher, Neyman, and the Creation of Classical Statistics, I've read some of his early correspondence with E. Pearson, when Neyman was still learning English. So I wasn't completely sure what to expect when I read his confidence interval article from 1937. I highly recommend reading it, as Neyman excellently summarizes complicated mathematical concepts in plain language. His work is highly approachable and he uses lots of examples to help drive home his points.

Neyman argued that, unless we have access to the full population we are studying, and are capable of measuring each individual in that population, there will be probabilities associated with our work; both the estimation process and the estimate itself should be expressed in probability. In fact, his classical approach to statistics includes probability in the estimation process, through the use of significance testing. He acknowledges that there are many different approaches to estimating values, and that while some are more right than others, none are likely to get you the exact population value, θ. They will all be estimates within a certain margin of error. Confidence intervals communicate that margin of error.

That is, he essentially says there is disagreement on the process of estimating population values from samples, and disagreement on the use of different estimation techniques (such as maximum likelihood, developed by R.A. Fisher). Though some approaches may be superior - and some of Neyman's footnotes feel very directed at Fisher - we are still trying to estimate an unknown parameter, so there is really no way to prove one is superior. But we can perhaps identify an interval surrounding the true population value.

That interval - the confidence interval - will be based on probability - the confidence coefficient - which is greater than 0 and less 1. The usual convention is 0.95 (95%), though he uses a variety of confidence coefficients in his paper and doesn't really settle on one as the gold standard. The 95% convention came later, probably because of it's connection to the probability we use in significance testing, where the convention for alpha is 0.05.

Without knowing the precise shape of the distribution of population values, we would instead use values with "intuitive" (his word) appeal, such as, for instance, the normal distribution (either the standard normal distribution or the t-distribution). He offers a variety of equations for different scenarios, but this one - the one that works with the normal distribution, which came at the end of the paper - is probably the approach most statistics students are familiar with. That is, we can use the values associated with different proportions of the curve around the mean to generate our confidence interval. We use these values from the z or t distribution (Neyman recommends t) as the multiple for the standard deviation. The exact procedure for confidence intervals varies depending on what type of estimate you're working with. For instance, some confidence intervals use standard error instead. But the basic procedure of choosing a probability and combining it with results of your analysis and values from a known distribution remains.

As the size of the sample used to estimate the population value increases, the bias (difference between the estimate and actual value) reduces toward 0. So estimates based on larger samples are more likely to be close to the true population value, and confidence intervals generated from that estimate will be narrower while still being likely to contain the actual value.

How likely? We don't actually know. As Neyman points out continuously in his paper, the probability that a range actually contains the population value is 0 or 1. There is no in between when it comes to real probability; it's either there or it isn't. But when we generate a confidence interval around our estimate, we don't know if it truly contains the actual value or not. So we draw upon the law of large numbers, that over time, with repeated estimates of the population value, we'll have the real population value in our ranges a certain proportion of the time, with that proportion equal to the confidence coefficient we choose.

Say we always use 95% confidence intervals in a certain area of study. (That's the convention, anyway.) With repeated research (conducted in an unbiased way), we'll have the real population value in our range much of the time, with the actual percentage of the time approaching 95% as the number of studies approaches infinity. As has been shown repeatedly, chance is lumpy, and a 95% change of something doesn't mean it will happen exactly 95% of the time, just like you won't have a perfect 50% heads in your coin flips.

A glance at Neyman's reference section shows many of the greats of statistics: Fisher, Hotelling, Kolmogorov, Lévy, Markoff... Many names you'll hear again and again in these GMIS posts.

Thursday, August 31, 2017

Climate Change and the Behavior of a Storm

Hurricane Harvey has been more devastating than most of us expected. As I stopped to grab breakfast on my way to work this morning, I saw an infographic on the front page of USA Today detailing just how bad things are in Texas in terms of the cost of the damage (to say nothing of the loss of human life):


Part of the reason Harvey has been so devastating is because its behavior has been different from many previous hurricanes, and climate change may be to blame:
In the case of Harvey, which is dumping rivers of rain in and around Houston and threatening millions of people with catastrophic flooding (see photos), at least three troubling factors converged. The storm intensified rapidly, it has stalled out over one area, and it is expected to continue dumping record rains for days and days.

Hurricanes tend to weaken as they approach land because they are losing access to the hot, wet ocean air that gives the storms their energy. Harvey's wind speeds, on the other hand, intensified by about 45 miles per hour in the last 24 hours before landfall, according to National Hurricane Center data.

[Kerry Emanuel, an atmospheric sciences professor at the Massachusetts Institute of Technology,] analyzed the evolution of 6,000 simulated storms, comparing how they evolved under historical conditions of the 20th century, with how they could evolve at the end of the 21st century if greenhouse gas emissions keep rising. The result: A storm that increases its intensity by 60 knots in the 24 hours before landfall may have been likely to occur once a century in the 1900s. By late in this century, they could come every five to 10 years.
As the article points out, the big reason for all the damage is the amount of rainfall, resulting in flooding. That too is likely due to climate change. In fact:
Every scientist contacted by National Geographic was in agreement that the volume of rain from Harvey was almost certainly driven up by temperature increases from human carbon-dioxide emissions.
This is of course exacerbated by the fact that Harvey has stalled over land. Most hurricanes break apart or move off. Interestingly enough, the article notes that most climate scientists don't think this particular stall can be attributed to climate change, just bad luck; more research is needed, though, because some say climate change could result in changes in pressure fronts, which would impact how long a storm stalls in one place.

Wednesday, August 30, 2017

Statistical Sins in History: Handling and Understanding Criticism

Today's Statistical Sins will be a little bit different, using an example from history of statistics to talk about an aspect of research publication. I'm currently reading Fisher, Neyman, and the Creation of Classical Statistics by Erich L. Lehmann. I've talked before about Egon Pearson, who was Jerzy Neyman's long-time collaborator. The feud between Neyman and Ronald Fisher is legendary in statistical history, as is the feud between Karl Pearson and Fisher. But not as much attention has been given to the feud between E. Pearson and Fisher. The start of that feud - though arguably mostly caused by Fisher and K. Pearson's ongoing competition of who could be more petty - can probably be traced to a review, authored by E. Pearson, about Fisher's book, Statistical Methods for Research Workers.

The title page from the first edition
The review in question was regarding Fisher's second edition of the book. It was positive overall, except for this (note: this is quoted from Fisher, Neyman, and the Creation of Classical Statistics; I didn't track down the original):
There is one criticism, however, which must be made from the statistical point of view. A large number of the tests developed are based... on the assumption that the population sampled is of the "normal" form. That this is the case may be gathered from a careful reading of the text, but the point is not sufficiently emphasized. It does not appear reasonable to lay stress on the "exactness" of the tests when no means whatever are given of appreciating how rapidly they become inexact as the population sampled diverges from normality... [N]o clear indication of the need for caution in their application is given.
The issues E. Pearson is addressing here are 1) the robustness of a test and 2) determining how far a dataset needs to diverge from normal before it no longer satisfies the requirements of the test. These are legitimate questions, and further, there is a very good reason E. Pearson raised them. But first, the fallout.

Fisher was pissed. He was so pissed he wrote a response to the journal that originally published the review (Nature). We don't know exactly what this letter said, but based on later correspondence, it appears Fisher believed the question of normality was irrelevant to the content of the book (and I'm sure there was some name-calling as well). As often occurs with a letter to the editor regarding a published paper, the editor sent it to E. Pearson and asked if he would like to respond. He wrote his response but before sending it off, showed it to William Sealy Gosset.

Gosset, who had a good working relationship with Fisher, decided to serve as mediator, and wrote a letter to Fisher to try to settle the dispute. Apparently that approach worked, because Fisher decided to withdraw his letter to Nature (which is why we don't know what it said) and suggested Gosset should instead write a letter (on Fisher's behalf) responding to E. Pearson's review. Of course, Fisher did end up writing a response... to Gosset's letter, because Gosset agreed with E. Pearson's comment about normality, saying that, though he believed the Student distribution (which he created) could withstand "small departures from normality," we needed more research into this topic, and in the meantime, experts in statistical distributions (like Fisher) could help guide us on how to respond when our data aren't normal. Gosset knew Fisher was a better mathematician, and likely saw this as a way of asking Fisher for help in answering these questions.

Fisher, instead, brought up the possibility of distribution-free tests.

The thing Fisher never really considered is why E. Pearson was so fixated on this issue of robustness and normality. Do you know what are two of E. Pearson's contributions to the field of statistics? Exploration into determining the best goodness of fit test (that is, the best way to determine if a set of data matches a theoretical distribution, like the normal distribution - part of his collaboration with Neyman) and the concept of robustness. In fact, he was already working on much of this when he wrote that review in 1929.

E. Pearson was not trying to make Fisher look bad or call him dumb. On the contrary: E. Pearson was trying to connect what he was working on to Fisher's work and set the stage for his own contributions. In fact, this is often the reason researchers will criticize another researcher's work in a paper or letter to the editor: they're setting the stage for the contribution they're about to make. They're taking the opportunity to say "we need X," only to turn around and deliver X soon after.

This is done all the time. People even do it in their own papers, when they highlight a certain shortcoming of their research in the discussion section; they're probably highlighting a flaw that they've already figured out how to fix and may already be testing in a new study. (Or they added it to make a reviewer happy.)

Fisher's response was because he couldn't see the reason E. Pearson was criticizing him. He just saw the criticism and went into rage mode. It's easy to do. Hearing criticism sucks. And while, as researchers we frequently have to deal with criticism of our work in dissertation defenses and peer reviews, they are rarely so public as they are with a published book review or letter to the editor.

I'll admit, when I received an email from a journal that someone had written a letter to the editor in response to one of my articles (and asking if I'd like to write a response), I made that sound kids make when they have a skinned knee:



It took some courage to open the file and read the letter. I was amazed to see it was incredibly positive. I can only imagine what my reaction would be if it hadn't been positive.


But if we can take a step back and realize why this researcher might be waging a particular criticism, it might make it a bit easier to handle the hurt feelings. Who knows how different things would have been for the field of statistics if - instead of throwing a tantrum and writing a pissed off letter to the editor - Fisher had written E. Pearson a letter directly saying, "I think this issue of normality is irrelevant to what I was trying to do. Why do you think it's important?" Maybe we would be talking today about the amazing collaboration between E. Pearson and Fisher. (Probably not, but a girl can dream, right?)

Tuesday, August 29, 2017

The Unpopularity of Presidential Pardons

Following up on yesterday's post about the Arpaio pardon, here's an article from FiveThirtyEight examining Presidential pardons over the year, highlighting not only the unpopularity of these pardons, but what makes Trump's pardon of Arpaio so unconventional:
Several political allies and foes immediately condemned the move as inappropriate and an insult to the justice system. But most of the criticized characteristics of Arpaio’s pardon have at least some parallels to previous ones.

The number of controversial characteristics of the Arpaio pardon, however, is unusual and raises questions about the political fallout that Trump will face. The Arpaio pardon, in other words, does have historical precedents (as Trump said on Monday) — just not good ones.
“A pardon is a judgment call that the president makes, and we get to police that through the political process,” [Michigan State University law professor Brian] Kalt said. Noah Feldman, a professor at Harvard Law School, said that the fact that Arpaio was convicted for deliberately ignoring a court’s order to stop violating individuals’ constitutional rights places him in a category of his own. The only recourse for such a dramatic abuse of presidential power, according to Feldman, is impeachment. Or, short of impeachment, Kalt pointed to Ford’s pardon of Nixon: “Ford decided it was the right thing to do, and he lost the election as a result.”