Showing posts with label books. Show all posts
Showing posts with label books. Show all posts

Monday, January 4, 2021

My Dark Vanessa: Book Review

Content Warning: sexual abuse, child abuse, and rape

Earlier today, I finished My Dark Vanessa, a debut novel by Kate Elizabeth Russell. The book is told by Vanessa Wye, a young woman who was abused by her boarding school English teacher starting when she was 15. The book spans 17 years, jumping between Vanessa's youth and adulthood. Before I get into my (slightly spoiler-y) review, I want to say: I loved this book, and I also have no desire to ever read it again.

As you can imagine from the title and brief synopsis, this a difficult book, as we hear everything that happened in the mind of a young woman who was gaslighted into believing she had all the power in situations when she had almost none and at the same time, that she had no power in situations where she could do something to stop the abuse. It's a deep and disturbing dive into the way an older man selects and grooms his victim, changing her thinking and behavior for decades, and convincing her that she's a willing participant, even when she describes very clear dissociation (being outside of herself) during the episodes of abuse - a reaction often seen in victims of child abuse.

The book also digs into two really key issues, that I haven't often seen explored or explored this well: 1. The narrative that women have "feminine gifts" that allow them to have power over a man, and make these men do things they wouldn't otherwise do. And 2. That coming forward is the only responsible thing a victim of such abuse can do.

The first issue (the "power" of femininity to take away men's agency) is such a pervasive part of rape culture. But this book also explores how this narrative has been romanticized to even apply to situations of a very young girl and a much older man, in stories like Lolita, American Beauty, and Pretty Baby. It is this romanticization and narrative that makes Vanessa continue her relationship with her abuser, Jacob Strane, even when it actively hurts her. He convinces her that he has so much more to lose than she does if their "love" becomes public, that he cannot help himself, that she has the power in the situation to consent or decline (even when he ignores her requests for him to stop and/or fails to ask for consent for very extreme sexual acts), and that, most of all, she is special because of this power she has. For a lonely girl, away from home for the first time, it's so easy to see how he selects and grooms her. But perhaps one of the most frustrating things is, even as I was reading and feeling what Vanessa feels, the descriptions and behaviors were so clear, I would shout at her as I read that there's some textbook-level gaslighting going on. It's why this is such a good book - that the author can give us those really clear cues while still telling the book in first-person, and avoid the "unreliable narrator" trope - and also one I hope to never read again.

This narrative of feminine wiles is perhaps ones of biggest issues we need to contend with if we want to do away with rape culture for good. It's a narrative that, on its surface, appears to assign all the power to the woman and none to her rapist or abuser, when at its core, it instead makes the woman powerless to stop (and deserving of) whatever harm is done to her. It's also a narrative that can be so easily spun as a positive thing when it is actually toxic and harmful. 

The second issue is a bit more ambiguous, at least for me, because before I read this book, I would have agreed with this second statement, that victims must come forward so that the abuser can be brought to justice and that others can be protected. I believed this even as a person who did not bring my own rapist to justice, something I was very ashamed of about myself. But this book made me realize just how tricky this issue is. 

At a surface level, it seems like a conflict between the needs of the individual and the needs of many, and from a philosophical standpoint, the needs of the many should outweigh the needs of the individual. But framing it in such a way takes away the individual's autonomy, a major issue considering that the abuse/rape was all about taking away one individual's autonomy. And victims already feel a great deal of guilt and self-blame for the event; they don't need the guilt of believing they failed others, or that they are in some way responsible for the reprehensible actions of another. 

Framing it as needs of the individual vs. needs of the many oversimplifies exactly what the needs of the individuals are (privacy, self-care, fear of reprisals, and so on), while also making that individual an accomplice in how another person's actions affects the many. In the case of adults in positions of power abusing the people they should be protecting, no victim should ever be to blame; this is on the system that put (and often helps to keep) that person in power, and on all of us, for the ways (big and small) that we may contribute to these power dynamics and rape culture.

This book was very triggering for me (even though my personal experiences do not resemble Vanessa), and I'm still working through the emotions it's brought up. I was reminded of a book I read in college, Bastard Out of Carolina, which also details years of sexual abuse of a child. When I finished that book, I threw it against the wall. Fortunately this book didn't elicit that reaction, but I didn't have a super-positive reaction to the ending either. 

I'm still glad I read it, though I probably wouldn't recommend it to anyone who might also be triggered, especially if they haven't been able to work through their own trauma through therapy or treatment. And I'll definitely keep an eye out for future books by Kate Elizabeth Russell.

Saturday, September 14, 2019

Movie Review: It Chapter 2

Speaking of Stephen King, I recently went to see It Chapter 2 with a friend. Here's what I thought (while I try to keep spoilers to a minimum).


The story starts off in Derry, Maine, present day, when Pennywise the Clown is seen again. Mike Hanlon finds a message at the site of a murder that says, "Come home." He proceeds to call his fellow members of the Losers' Club, telling them it's time to keep the promise they made 27 years before: to come back and kill Pennywise if he ever returns. Unfortunately, the remaining "losers" don't remember their time in Derry, and have to be reminded of many of the events from the first movie in order to effectively fight Pennywise.

First, what I liked about the movie. While most of the movie takes place in the present day, there are a few scenes that go back to the losers as they were 27 years before. The clubhouse they built, which was very important in the book, finally makes an appearance. We also get to meet Bill's bike, Silver, at last. The movie, while scary, is also incredibly funny. The characters impart their dark, dry humor with each other often, even during tense scenes, which feels completely real and believable. And my favorite part was a great reference to this scene in my all-time favorite horror movie, John Carpenter's The Thing:


Oh yeah, and in addition to The Thing reference, the movie features some fun fan service for people who love horror movies, and great Easter eggs for anyone who loves horror movie trivia.

I also liked seeing Mike get a much more important role in this movie, since he was relegated to the sidelines in the first one and much of his contribution to that story was given instead to Ben. And some of the subplots from the book, while interesting, were cut from the movie, making it a much more straightforward story.

At the same time, the things I didn't like as much about the movie were also related to departures from the book. The clubhouse, while finally appearing, was given little to no importance in terms of the ritual to fight Pennywise. I also didn't like some of the changes they made to Mike. In the book, Mike often didn't tell the others things he remembered but they didn't because they needed to find them out in their own time. But movie Mike also lied to his friends and purposefully put them in danger, not something book Mike would have done. In the book, the danger was always Pennywise. Really, my biggest complaint about both movies has to do with their changes to Mike's character.

The Ritual to kill Pennywise was also much more interesting in the book, though I suppose it would have been difficult to film coherently, because the book version was much more about emotions and thoughts, as opposed to clear actions. Honestly, I didn't really like the way they defeated Pennywise in the movie, but I still enjoyed the rest of the movie, so I'll let it go.

Overall, It Chapter 2 was a fun, entertaining movie that neatly wrapped up the Pennywise and Losers' Club storylines. While I'm a little sad about some of the elements from the book that were cut or changed, I'm glad that they did this as a single film instead of a two-parter, as they would have had to in order to include some of the scenes Stephen King requested they keep in the movie. The tone of this movie is certainly different than the first, but it works. As I said, you could see these characters growing up into the sarcastic, wry, somewhat dark humor they impart throughout their scenes, based on what they went through in the first movie. Despite fighting a shape-shifting, pan-dimensional fear and flesh devourer, the characters felt real.

Tuesday, September 3, 2019

Totally Superfluous Book Review: Stephen King's Cujo

I've been a fan of Stephen King's since I was a child, and have recently created a personal goal for myself to read all of his books. The most recent entry into that read list was Cujo, the story of a rabid St. Bernard who terrorizes two families in Castle Rock, Maine.


This book was an incredibly difficult read and I finished it last night feeling gross all over. I have to say, I hated this book. I sincerely hope people who want to get into Stephen King don't choose this as their first read, because it will leave them with a completely inaccurate view of King's writing. His style is brutally honest and often darkly funny, but not mean-spirited and sexist as this book is. The monster wins, and I know that message of this book was a reflection of what was going on in his own life at the time. He had a severe substance abuse problem at the time, and as such, says he has no memory of writing this book. The monster in his closet was winning, and he poured all of that strife and darkness into this book. If anything, this book is a reflection of how the history of the writer influences how one of their projects is viewed and interpreted. This book is filled with hopelessness and anger.

If you, like me, want to read all of King's books, you should probably read this one. But otherwise, this is one to be avoided, unless you'd like a demonstration of how personal demons seep into one's writing.

Thursday, September 20, 2018

On Books, Bad Marketing, and #MeToo

I recently read a book I absolutely loved. Luckiest Girl Alive by Jessica Knoll really got to me and got under my skin in a way few books have. When I went to Goodreads to give the book a 5-star rating, I was shocked at so many negative reviews of 1 and 2 stars. What did I see in this book that others didn't? And what did others see in the book that I didn't?

Nearly all of the negative reviews reference Gone Girl, a fantastic book by one of my new favorite authors, Gillian Flynn. And in fact, Knoll's book had an unfortunate marketing team that sold her book as the "next Gone Girl." Gone Girl this book is not, and that's okay. In fact, it's a bad comparative title, which can break a book. The only thing Ani, the main character in Luckiest Girl Alive has in common with Amy is that they both tried to reinvent themselves to be the type of girl guys like: the cool girl. Oh, and both names begin with A (kind of - Ani's full name is Tifani). That's about it.

While Gone Girl is a twisted tale about the girl Amy pretended to be, manipulating everyone along the way, Luckiest Girl Alive is the case a girl who was drugged and victimized by 4 boys at her school, then revictimized when they turned her classmates against her, calling her all the names we throw so easily at women: slut, skank, an so on. No one believed her, and she was ostracized. She responded by putting as much distance as she could from the events of the past and the person she was, hoping that with the great job, the handsome and rich fiance, designer clothes, and a great body (courtesy of a wedding diet that's little more than an eating disorder), she can move past what happened to her. She becomes bitter and compartmentalized (and also shows many symptoms of PTSD), which may be why some readers didn't like the book and had a hard time dealing with Ani.

I know what's it like to be Ani. Like her, and so many women, I have my own #MeToo story. And I know what it's like to have people I love react like Ani's fiance and mother, who'd rather pretend the events in her past never happened, who change the subject when it's brought up. It's not that we fixate on these things. But those events change us, and when a survivor needs to talk, the people she cares about need to listen, understand, and withhold judgment.

That marketing team failed Jessica Knoll for what should have been a harsh wakeup call in the age of #MeToo, to stop victim blaming. Just like Ani's classmates failed her. And just like nearly everyone close to her failed Amber Wyatt, a survivor of sexual assault who was recently profiled in a Washington Post article (via The Daily Parker).

Given these reactions, it's no wonder many sexual assaults go unreported. Some statistics put that proportion at 80% or higher, with more conservative estimates at 50%. And unfortunately, the proportion of unreported sexual assaults is higher when the victim knows the attacker(s). They're seen as causing trouble, rocking the boat, or trying to save face after regretting a consensual encounter. It's time we start listening to and believing women.

Tuesday, September 18, 2018

I've Got a Bad Feeling About This

In the 1993 film Jurassic Park, scientist Ian Malcolm expressed serious concern about John Hammond's decision to breed hybrid dinosaurs for his theme park. As Malcolm says in the movie, "No, hold on. This isn't some species that was obliterated by deforestation, or the building of a dam. Dinosaurs had their shot, and nature selected them for extinction."

This movie was, and still is (so far), science fiction. But a team of Russian scientists are working to make something similar into scientific fact:
Long extinct cave lions may be about to rise from their icy graves and prowl once more alongside woolly mammoths and ancient horses in a real life Jurassic Park.

In less than 10 years it is hoped the fearsome big cats will be released from an underground lab as part of a remarkable plan to populate a remote spot in Russia with Ice Age animals cloned from preserved DNA.

Experiments are already underway to create the lions and also extinct ancient horses found in Yakutia, Siberia, seen as a prelude to restoring the mammoth.

Regional leader Aisen Nikolaev forecast that co-operation between Russian, South Korean and Japanese scientists will see the “miracle” return of woolly mammoths inside ten years.
Jurassic Park is certainly not the only example of fiction exploring the implications of man "playing god." Many works of literature, like Frankenstein, The Island of Doctor Moreau, and more recent examples like Lullaby (one of my favorites), have examined this very topic. It never ends well.

By Mauricio Antón - from Caitlin Sedwick (1 April 2008). "What Killed the Woolly Mammoth?". PLoS Biology 6 (4): e99. DOI:10.1371/journal.pbio.0060099., CC BY 2.5, Link

Sunday, September 16, 2018

Statistics Sunday: What Should I Read Next?

When You Need a New Book to Read I log all of my books on Goodreads. On top of that, whenever I hear about a new book I have to read, I add it on Goodreads, so I remember it. Of course, this means my Goodreads bookshelves are a little out of control. Fortunately, I can use R to dig through my Goodreads to-read shelf and figure out the next book to read and/or buy.

If you're on Goodreads, you can easily download your entire bookshelf, including your to-read books, by going to "My Books" then clicking "Import and Export". On the right side of the screen will be a link for "Export Library". Click that and give it a minute (or several). Soon, a link will appear to download your entire library in a CSV file. You can then bring that into R.

If I'm ever stuck for the next book to read, I can use this file to randomly select a book from my to-read list to check out next. Because I own a lot of books on my to-read list, I'd like to filter that dataset to only include books I own. (Note: You can add books to your "owned" list by clicking on "My Books" then "Owned Books" to select which books you already have in your library. Otherwise, you can keep running the sample function until you get a book you already own or have ready access to. You'd just want to skip the first part of the filter in "reading_list" below.)

setwd("~/Dropbox")
library(tidyverse)
books <- read_csv("goodreads_library_export.csv", col_names = TRUE)
reading_list <- books %>%
  filter(`Owned Copies` == 1, `Exclusive Shelf` == "to-read")
head(reading_list)
## # A tibble: 6 x 31
##   `Book Id` Title    Author  `Author l-f` `Additional Auth… ISBN    ISBN13
##       <int> <chr>    <chr>   <chr>        <chr>             <chr>    <dbl>
## 1  27877138 It       Stephe… King, Steph… <NA>              1501…  9.78e12
## 2     10611 The Eye… Stephe… King, Steph… <NA>              0751…  9.78e12
## 3     11570 Dreamca… Stephe… King, Steph… William Olivier … 2226…  9.78e12
## 4  36452674 The Squ… Kevin … Hearne, Kev… <NA>              <NA>  NA      
## 5  38193271 Bickeri… Mildre… Abbott, Mil… <NA>              <NA>  NA      
## 6  20873740 Sapiens… Yuval … Harari, Yuv… <NA>              <NA>  NA      
## # ... with 24 more variables: `My Rating` <int>, `Average Rating` <dbl>,
## #   Publisher <chr>, Binding <chr>, `Number of Pages` <int>, `Year
## #   Published` <int>, `Original Publication Year` <int>, `Date
## #   Read` <date>, `Date Added` <date>, Bookshelves <chr>, `Bookshelves
## #   with positions` <chr>, `Exclusive Shelf` <chr>, `My Review` <chr>,
## #   Spoiler <chr>, `Private Notes` <chr>, `Read Count` <int>, `Recommended
## #   For` <chr>, `Recommended By` <chr>, `Owned Copies` <int>, `Original
## #   Purchase Date` <chr>, `Original Purchase Location` <chr>,
## #   Condition <chr>, `Condition Description` <chr>, BCID <chr>

Now I have a data frame of books that I own and have not read. This data frame contains 55 books. Drawing a random sample of 1 book is quite easy.

reading_list[sample(1:nrow(reading_list), 1),]
## # A tibble: 1 x 31
##   `Book Id` Title    Author  `Author l-f`  `Additional Auth… ISBN   ISBN13
##       <int> <chr>    <chr>   <chr>         <chr>             <chr>   <dbl>
## 1     14201 Jonatha… Susann… Clarke, Susa… <NA>              0765… 9.78e12
## # ... with 24 more variables: `My Rating` <int>, `Average Rating` <dbl>,
## #   Publisher <chr>, Binding <chr>, `Number of Pages` <int>, `Year
## #   Published` <int>, `Original Publication Year` <int>, `Date
## #   Read` <date>, `Date Added` <date>, Bookshelves <chr>, `Bookshelves
## #   with positions` <chr>, `Exclusive Shelf` <chr>, `My Review` <chr>,
## #   Spoiler <chr>, `Private Notes` <chr>, `Read Count` <int>, `Recommended
## #   For` <chr>, `Recommended By` <chr>, `Owned Copies` <int>, `Original
## #   Purchase Date` <chr>, `Original Purchase Location` <chr>,
## #   Condition <chr>, `Condition Description` <chr>, BCID <chr>

According to this random sample, the next book I should read is Jonathan Strange & Mr Norrell. Now if I'm ever stuck for a book to read, I can use this code to find one. And if I'm in a bookstore, picking up something new - as is often the case, since bookstores are one of my happy places - I can update the code to tell me which book I should buy next.

to_buy <- books %>%
  filter(`Owned Copies` == 0, `Exclusive Shelf` == "to-read")
to_buy[sample(1:nrow(to_buy), 1),]
## # A tibble: 1 x 31
##   `Book Id` Title    Author   `Author l-f` `Additional Aut… ISBN    ISBN13
##       <int> <chr>    <chr>    <chr>        <chr>            <chr>    <dbl>
## 1   2906039 Just Af… Stephen… King, Steph… <NA>             1416…  9.78e12
## # ... with 24 more variables: `My Rating` <int>, `Average Rating` <dbl>,
## #   Publisher <chr>, Binding <chr>, `Number of Pages` <int>, `Year
## #   Published` <int>, `Original Publication Year` <int>, `Date
## #   Read` <date>, `Date Added` <date>, Bookshelves <chr>, `Bookshelves
## #   with positions` <chr>, `Exclusive Shelf` <chr>, `My Review` <chr>,
## #   Spoiler <chr>, `Private Notes` <chr>, `Read Count` <int>, `Recommended
## #   For` <chr>, `Recommended By` <chr>, `Owned Copies` <int>, `Original
## #   Purchase Date` <chr>, `Original Purchase Location` <chr>,
## #   Condition <chr>, `Condition Description` <chr>, BCID <chr>

So next time I'm at a bookstore, which will be tomorrow (since I'll be hanging out in Evanston for a class at my dance studio and plan to hit up the local Barnes & Noble), I should pick up a copy of Just After Sunset.

If you're on Goodreads, feel free to add me!

Monday, August 27, 2018

Statistics Sunday: Visualizing Regression

Statistics Sunday: Visualizing Regression I had some much needed downtime this weekend, after an exhausting week, along with some self-care - Saturday I had a one-hour deep tissue massage, which left me a little bruised but much more relaxed, and Sunday I spent a few hours in the salon chair having my color touched up, which left me much blonder. Which is why I'm a little late with my Statistics Sunday post, but today, I'm introducing another recently discovered r package: rpart. Short for "recursive partitioning," this package creates decision trees for classification, regression, and survival analyses. Today, I'm going to demonstrate using the rpart package for visualizing regression.

To demonstrate this technique, I'm using my 2017 reading dataset. A reader requested I make this dataset available, which I have done - you can download it here. This post describes the data in more detail, but the short description is that this dataset contains the 53 books I read last year, with information on book genre, page length, how long it took me to read it, and two ratings: my own rating and the average rating on Goodreads. Look for another dataset soon, containing my 2018 reading data; I made a goal of 60 books and I've already read 50, meaning lots more data than last year. I'm thinking of going for broke and bumping my reading goal up to 100, because apparently 60 books is not enough of a challenge for me now that I spend so much downtime reading.

First, I'll load my dataset, then conduct the basic linear model I demonstrated in the post linked above.

setwd("~/R")
library(tidyverse)
## Warning: Duplicated column names deduplicated: 'Author' => 'Author_1' [13]
## Parsed with column specification:
## cols(
##   .default = col_integer(),
##   Title = col_character(),
##   Author = col_character(),
##   G_Rating = col_double(),
##   Started = col_character(),
##   Finished = col_character()
## )
## See spec(...) for full column specifications.
colnames(books)[13] <- "Author_Gender"
myrating<-lm(My_Rating ~ Pages + Read_Time + Author_Gender + Fiction + Fantasy + Math_Stats + YA_Fic, data=books)
summary(myrating)
## 
## Call:
## lm(formula = My_Rating ~ Pages + Read_Time + Author_Gender + 
##     Fiction + Fantasy + Math_Stats + YA_Fic, data = books)
## 
## Residuals:
##      Min       1Q   Median       3Q      Max 
## -0.73120 -0.34382 -0.00461  0.24665  1.49932 
## 
## Coefficients:
##                 Estimate Std. Error t value Pr(>|t|)    
## (Intercept)    3.5861211  0.2683464  13.364   <2e-16 ***
## Pages          0.0019578  0.0007435   2.633   0.0116 *  
## Read_Time     -0.0244168  0.0204186  -1.196   0.2380    
## Author_Gender -0.1285178  0.1666207  -0.771   0.4445    
## Fiction        0.1052319  0.2202581   0.478   0.6351    
## Fantasy        0.5234710  0.2097386   2.496   0.0163 *  
## Math_Stats    -0.2558926  0.2122238  -1.206   0.2342    
## YA_Fic        -0.7330553  0.2684623  -2.731   0.0090 ** 
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
## 
## Residual standard error: 0.4952 on 45 degrees of freedom
## Multiple R-squared:  0.4624, Adjusted R-squared:  0.3788 
## F-statistic:  5.53 on 7 and 45 DF,  p-value: 0.0001233

These analyses show that I give higher ratings for books that are longer and Fantasy genre, and lower ratings to books that are Young Adult Fiction. Now let's see what happens if I run this same linear model through rpart. Note that this is a slightly different technique, looking for cuts that differentiate outcomes, so it will find slightly different results.

library(rpart)
tree1 <- rpart(My_Rating ~ Pages + Read_Time + Author_Gender + Fiction + Fantasy + Math_Stats + YA_Fic, method = "anova", data=books)
printcp(tree1)
## 
## Regression tree:
## rpart(formula = My_Rating ~ Pages + Read_Time + Author_Gender + 
##     Fiction + Fantasy + Math_Stats + YA_Fic, data = books, method = "anova")
## 
## Variables actually used in tree construction:
## [1] Fantasy    Math_Stats Pages     
## 
## Root node error: 20.528/53 = 0.38733
## 
## n= 53 
## 
##         CP nsplit rel error  xerror    xstd
## 1 0.305836      0   1.00000 1.03609 0.17531
## 2 0.092743      1   0.69416 0.76907 0.12258
## 3 0.022539      2   0.60142 0.71698 0.11053
## 4 0.010000      3   0.57888 0.74908 0.11644

These results differ somewhat. Pages is still a significant variable, as is Fantasy. But now Math_Stats (indicating books that are about mathematics or statistics, one of my top genres from last year) also is. These are the variables used by the analysis to construct my regression tree. If we look at the full summary -

summary(tree1)
## Call:
## rpart(formula = My_Rating ~ Pages + Read_Time + Author_Gender + 
##     Fiction + Fantasy + Math_Stats + YA_Fic, data = books, method = "anova")
##   n= 53 
## 
##           CP nsplit rel error    xerror      xstd
## 1 0.30583640      0 1.0000000 1.0360856 0.1753070
## 2 0.09274251      1 0.6941636 0.7690729 0.1225752
## 3 0.02253938      2 0.6014211 0.7169813 0.1105294
## 4 0.01000000      3 0.5788817 0.7490758 0.1164386
## 
## Variable importance
##      Pages    Fantasy    Fiction     YA_Fic Math_Stats  Read_Time 
##         62         18          8          6          4          3 
## 
## Node number 1: 53 observations,    complexity param=0.3058364
##   mean=4.09434, MSE=0.3873265 
##   left son=2 (9 obs) right son=3 (44 obs)
##   Primary splits:
##       Pages         < 185 to the left,  improve=0.30583640, (0 missing)
##       Fiction       < 0.5 to the left,  improve=0.24974560, (0 missing)
##       Fantasy       < 0.5 to the left,  improve=0.20761810, (0 missing)
##       Math_Stats    < 0.5 to the right, improve=0.20371790, (0 missing)
##       Author_Gender < 0.5 to the right, improve=0.02705187, (0 missing)
## 
## Node number 2: 9 observations
##   mean=3.333333, MSE=0.2222222 
## 
## Node number 3: 44 observations,    complexity param=0.09274251
##   mean=4.25, MSE=0.2784091 
##   left son=6 (26 obs) right son=7 (18 obs)
##   Primary splits:
##       Fantasy    < 0.5 to the left,  improve=0.15541600, (0 missing)
##       Fiction    < 0.5 to the left,  improve=0.12827990, (0 missing)
##       Math_Stats < 0.5 to the right, improve=0.10487750, (0 missing)
##       Pages      < 391 to the left,  improve=0.05344995, (0 missing)
##       Read_Time  < 7.5 to the right, improve=0.04512078, (0 missing)
##   Surrogate splits:
##       Fiction   < 0.5 to the left,  agree=0.773, adj=0.444, (0 split)
##       YA_Fic    < 0.5 to the left,  agree=0.727, adj=0.333, (0 split)
##       Pages     < 370 to the left,  agree=0.682, adj=0.222, (0 split)
##       Read_Time < 3.5 to the right, agree=0.659, adj=0.167, (0 split)
## 
## Node number 6: 26 observations,    complexity param=0.02253938
##   mean=4.076923, MSE=0.2248521 
##   left son=12 (7 obs) right son=13 (19 obs)
##   Primary splits:
##       Math_Stats    < 0.5 to the right, improve=0.079145230, (0 missing)
##       Pages         < 364 to the left,  improve=0.042105260, (0 missing)
##       Fiction       < 0.5 to the left,  improve=0.042105260, (0 missing)
##       Read_Time     < 5.5 to the left,  improve=0.016447370, (0 missing)
##       Author_Gender < 0.5 to the right, improve=0.001480263, (0 missing)
## 
## Node number 7: 18 observations
##   mean=4.5, MSE=0.25 
## 
## Node number 12: 7 observations
##   mean=3.857143, MSE=0.4081633 
## 
## Node number 13: 19 observations
##   mean=4.157895, MSE=0.132964

we see that Fiction, YA_Fic, and Read_Time were also significant variables. The problem is that there is multicollinearity between Fiction, Fantasy, Math_Stats, and YA_Fic. All Fantasy and YA_Fic books are Fiction, while all Math_Stats books are not Fiction. And all YA_Fic books I read were Fantasy. This is probably why the tree didn't use Fiction or YA_Fic. I'm not completely clear on why Read_Time didn't end up in the regression tree, but it may be because my read time was pretty constant among the different splits and didn't add any new information to the tree. If I were presenting these results somewhere other than my blog, I'd probably want to do some follow-up analyses to confirm this fact.

Now the fun part: let's plot our regression tree:

plot(tree1, uniform = TRUE, main = "Regression Tree for My Goodreads Ratings")
text(tree1, use.n = TRUE, all = TRUE, cex = 0.8)
This tree shows that, before taking into account anything, my average book rating was 4.09. If a book is shorter than 185 pages (9 books in my dataset), it's average rating was 3.33. For longer books (44), the average rating was 4.25. But there's more to it than that. For non-Fantasy books (26), the average was 4.08, while the Fantasy (18 books) average was 4.5. If the book was Math_Stats (7), I gave it an average of 3.86, and if it was not Math_Stats (19), the average was 4.16. (Unfortunately, R cuts off the bottom part of the plot.)

While this plot is great for a quick visualization, I can make a nicer looking plot (which doesn't cut off the bottom text) as a PostScript file.

post(tree1, file = "mytree.ps",
        title = "Regression Tree for Rating")

I converted that to a PDF, which you can view here.

Hope you enjoyed this post! Have any readers used this technique before? Any thoughts or applications you'd like to share? (Self-promotion highly encouraged!)

Sunday, August 19, 2018

Statistics Sunday: Using Text Analysis to Become a Better Writer

Using Text Analysis to Become a Better Writer We all have words we love to use, and that we perhaps use too much. As an example: I have a tendency to use the same transitional statements, to the point that, before I submit a manuscript, I do a find all to see how many times I've used some of my favorites, e.g., additionally, though, and so on.

I'm sure we all have our own words we use way too often.


Text analysis can also be used to discover patterns in writing, and for a writer, may be helpful in discovering when we depend too much on certain words and phrases. For today's demonstration, I read in my (still in-progress) novel - a murder mystery called Killing Mr. Johnson - and did the same type of text analysis I've been demonstrating in recent posts.

To make things easier, I copied the document into a text file, and used the read_lines and tibble functions to prepare data for my analysis.

setwd("~/Dropbox/Writing/Killing Mr. Johnson")

library(tidyverse)
KMJ_text <- read_lines('KMJ_full.txt')

KMJ <- tibble(KMJ_text) %>%
  mutate(linenumber = row_number())

I kept my line numbers, which I could use in some future analysis. For now, I'm going to tokenize my data, drop stop words, and examine my most frequently used words.

library(tidytext)
KMJ_words <- KMJ %>%
  unnest_tokens(word, KMJ_text) %>%
  anti_join(stop_words)
## Joining, by = "word"
KMJ_words %>%
  count(word, sort = TRUE) %>%
  filter(n > 75) %>%
  mutate(word = reorder(word, n)) %>%
  ggplot(aes(word, n)) +
  geom_col() + xlab(NULL) + coord_flip()


Fortunately, my top 5 words are the names of the 5 main characters, with the star character at number 1: Emily is named almost 600 times in the book. It's a murder mystery, so I'm not too surprised that words like "body" and "death" are also common. But I know that, in my fiction writing, I often depend on a word type that draws a lot of disdain from authors I admire: adverbs. Not all adverbs, mind you, but specifically (pun intended) the "-ly adverbs."

ly_words <- KMJ_words %>%
  filter(str_detect(word, ".ly")) %>%
  count(word, sort = TRUE)

head(ly_words)
## # A tibble: 6 x 2
##   word         n
##   <chr>    <int>
## 1 emily      599
## 2 finally     80
## 3 quickly     60
## 4 emily’s     53
## 5 suddenly    39
## 6 quietly     38

Since my main character is named Emily, she was accidentally picked up by my string detect function. A few other top words also pop up in the list that aren't actually -ly adverbs. I'll filter those out then take a look at what I have left.

filter_out <- c("emily", "emily's", "emily’s","family", "reply", "holy")

ly_words <- ly_words %>%
  filter(!word %in% filter_out)

ly_words %>%
  filter(n > 10) %>%
  mutate(word = reorder(word, n)) %>%
  ggplot(aes(word, n)) +
  geom_col() + xlab(NULL) + coord_flip()


I use "finally", "quickly", and "suddenly" far too often. "Quietly" is also up there. I think the reason so many writers hate on adverbs is because it can encourage lazy writing. You might write that someone said something quietly or softly, but is there a better word? Did they whisper? Mutter? Murmur? Hiss? Did someone "move quickly" or did they do something else - run, sprint, dash?

At the same time, sometimes adverbs are necessary. I mean, can I think of a complete sentence that only includes an adverb? Definitely. Still, it might become tedious if I keep depending on the same words multiple times, and when a fiction book (or really any kind of writing) is tedious, we often give up. These results give me some things to think about as I edit.

Still have some big plans on the horizon, including some new statistics videos, a redesigned blog, and more surprises later! Thanks for reading!

Thursday, August 16, 2018

Finding Strengths

As part of my job and newly reorganized departments, my boss had many us take the Clifton StrengthsFinder, a measure developed by Donald O. Clifton and Gallup. This measure, developed through semi-structured interviews and subsequent psychometric research, identifies an individual's top 5 strengths from a list of 34. Here are my results:


The book that comes along with the assessment describes the 34 themes in detail and gives very basic information on the measure's development. But for the psychometricially inclined, you can read a detailed technical report of the measure's evidence for reliability and validity here. In general, the measure shows acceptable reliability and construct validity. There are moderate to strong correlations with the Big Five Personality Traits. My themes, specifically, relate to my high Agreeableness and Openness to Experience on the Big Five. (For comparison, here are my Myers-Briggs results.)

The report also talks about how the themes relate to leadership potential. What I'm best at, according to these results, are Relationship Building and Strategic Thinking.

And, of course, I always enjoy taking tests and measures, especially if I think they'll tell me something about myself.

Tuesday, August 7, 2018

Statistics Sunday: Highlighting a Subset of Data in ggplot2

Highlighting Specific Cases in ggplot2 Here's my belated Statistics Sunday post, using a cool technique I just learned about: gghighlight. This R package works with ggplot2 to highlight a subset of data. To demonstrate, I'll use a dataset I analyzed for a previous post about my 2017 reading habits. [Side note: My reading goal for this year is 60 books, and I'm already at 43! I may have to increase my goal at some point.]

setwd("~/R")
library(tidyverse)
books<-read_csv("2017_books.csv", col_names = TRUE)
## Warning: Duplicated column names deduplicated: 'Author' => 'Author_1' [13]
## Parsed with column specification:
## cols(
##   .default = col_integer(),
##   Title = col_character(),
##   Author = col_character(),
##   G_Rating = col_double(),
##   Started = col_character(),
##   Finished = col_character()
## )
## See spec(...) for full column specifications.

One analysis I conducted with this dataset was to look at the correlation between book length (number of pages) and read time (number of days it took to read the book). We can also generate a scatterplot to visualize this relationship.

cor.test(books$Pages, books$Read_Time)
## 
## 	Pearson's product-moment correlation
## 
## data:  books$Pages and books$Read_Time
## t = 3.1396, df = 51, p-value = 0.002812
## alternative hypothesis: true correlation is not equal to 0
## 95 percent confidence interval:
##  0.1482981 0.6067498
## sample estimates:
##       cor 
## 0.4024597
scatter <- ggplot(books, aes(Pages, Read_Time)) +
  geom_point(size = 3) +
  theme_classic() +
  labs(title = "Relationship Between Reading Time and Page Length") +
  ylab("Read Time (in days)") +
  xlab("Number of Pages") +
  theme(legend.position="none",plot.title=element_text(hjust=0.5))

There's a significant positive correlation here, meaning the longer books take more days to read. It's a moderate correlation, and there are certainly other variables that may explain why a book took longer to read. For instance, nonfiction books may take longer. Books read in October or November (while I was gearing up for and participating in NaNoWriMo, respectively) may also take longer, since I had less spare time to read. I can conduct regressions and other analyses to examine which variables impact read time, but one of the most important parts of sharing results is creating good data visualizations. How can I show the impact these other variables have on read time in an understandable and visually appealing way?

gghighlight will let me draw attention to different parts of the plot. For example, I can ask gghighlight to draw attention to books that took longer than a certain amount of time to read, and I can even ask it to label those books.

library(gghighlight)
scatter + gghighlight(Read_Time > 14) +
  geom_label(aes(label = Title),
             hjust = 1,
             vjust = 1,
             fill = "blue",
             color = "white",
             alpha = 0.5)


Here, the gghighlight function identifies the subset (books that took more than 2 weeks to read) and labels those books with the Title variable. Three of the four books with long read time values are non-fiction, and one was read for a course I took, so reading followed a set schedule. But the fourth is a fiction book, which took over 20 days to read. Let's see how month impacts reading time, by highlighting books read in November. To do that, I'll need to alter my dataset somewhat. The dataset contains a starting date and finish date, which were read in as characters. I need to convert those to dates and pull out the month variable to create my indicator.

library(lubridate)
## 
## Attaching package: 'lubridate'
## The following object is masked from 'package:base':
## 
##     date
books$Started <- mdy(books$Started)
books$Start_Month <- month(books$Started)
books$Month <- ifelse(books$Start_Month > 10 & books$Start_Month < 12, books$Month <- 1,
                      books$Month <- 0)
scatter + gghighlight(books$Month == 1) +
  geom_label(aes(label = Title), hjust = 1, vjust = 1, fill = "blue", color = "white", alpha = 0.5)


The book with the longest read time was, in fact, read during November, when I was spending most of my time writing.

Friday, August 3, 2018

Stats Note: Making Sense of Open-Ended Responses with Text Analysis

Using Text Mining on Open Ended Items Good survey design is both art and science. You have to think about how people will read and process your questions, and what sorts of responses might result from different question forms and wording. One of the big rules I follow in survey design is that you don't assess any of your most important topics with an open-ended item. Most people skip them, because they're more work than selecting options from a list, and people who do complete may give you terse, unhelpful, or gibberish answers.

When I was working on a large survey among people with spinal cord injuries/disorders, the survey designers decided to assess the exact details of the respondent's spinal injury or disorder with an open-ended question so that people could describe it "in their own words." As you might have guessed, the data were mostly meaningless and as such unusable. But many hypotheses and research questions dealt with looking at differences between people who sustained an injury and those with a disorder affecting their spinal cord. We couldn't even begin to test or examine any of those in the course of our study. It was unbelievably frustrating, because we could have gotten the information we needed with a single question and some categorical responses. We could have then asked people to supplement their answer with the open-ended question. Most would skip it, but we'd have the data we needed to answer our questions.

In my job, I inherited the results of a large survey involving a variety of dental professional populations. Once again, certain items that could have been addressed with a few close-ended questions were instead open-ended questions and not many of the responses are useful. The item that inspired this blog post assessed the types of products dental assistants are involved in purchasing, which can include anything from office supplies to personal protective equipment to large equipment (X-ray machine, etc.). Everyone had a different way of articulating what they were involved with purchasing, some simply saying "all dental supplies" or "everything," while others gave more specific details. In total, about 400 people responded to this item, which is a lot of data to dig through. But thanks to my new experience with text mining in R, I was able to try to make some sense of responses. Mind you, a lot of the responses can't be understood, but it's at least something.

I decided I could probably begin to categorize responses by slicing the responses up into the individual words. I can generate counts overall as well as by respondent, and use this to begin examining and categorizing the data. Finally, if I need some additional context, I can ask R to give me all responses that contain a certain word.

Because I don't own these data, I can't share them on my blog. But I can demonstrate with different text data to show what I did. To do this, I'll use one of my all-time favorite books, The Wonderful Wizard of Oz by L. Frank Baum, which is available through Project Gutenberg. The gutenbergr package will let me download the full-text.

library(gutenbergr)
gutenberg_works(title == "The Wonderful Wizard of Oz")
## # A tibble: 1 x 8
##   gutenberg_id title   author   gutenberg_autho~ language gutenberg_books~
##          <int> <chr>   <chr>               <int> <chr>    <chr>           
## 1           55 The Wo~ Baum, L~               42 en       Children's Lite~
## # ... with 2 more variables: rights <chr>, has_text <lgl>

Now we have the gutenberg_id, which will be used to download the fulltext into a data frame.

WOz <- gutenberg_download(gutenberg_id = 55)
## Determining mirror for Project Gutenberg from http://www.gutenberg.org/robot/harvest
## Using mirror http://aleph.gutenberg.org
WOz$line <- row.names(WOz)

To start exploring the data, I'll need to use the tidytext package to unnest tokens (words) and remove stop words that don't tell us much.

library(tidyverse)
library(tidytext)
tidy_oz <- WOz %>%
  unnest_tokens(word, text) %>%
  anti_join(stop_words)
## Joining, by = "word"

Now I can begin to explore my "responses" by generating overall counts and even counts by respondent. (For the present example, I'll use my line number variable instead of the respondent id number I used in the actual dataset.)

word_counts <- tidy_oz %>%
  count(word, sort = TRUE)
resp_counts <- tidy_oz %>%
  count(line, word, sort = TRUE) %>%
  ungroup()

When I look at my overall counts, I see that the most frequently used words are the main characters of the book. So from this, I could generate a category "characters." Words like "Kansas", "Emerald" and "City" are also common, and I could create a category called "places." Finally, "heart" and "brains" are common - a category could be created to encompass what the characters are seeking. Obviously, this might not be true for every instance. It could be that someone "didn't have the heart" to tell someone something. I can try to separate out those instances by looking at the original text.

heart <- WOz[grep("heart",WOz$text),]
head(heart)
## # A tibble: 6 x 3
##   gutenberg_id text                                                  line 
##          <int> <chr>                                                 <chr>
## 1           55 happiness to childish hearts than all other human cr~ 48   
## 2           55 the heartaches and nightmares are left out.           62   
## 3           55 and press her hand upon her heart whenever Dorothy's~ 110  
## 4           55 strange people.  Her tears seemed to grieve the kind~ 375  
## 5           55 Dorothy ate a hearty supper and was waited upon by t~ 517  
## 6           55 She ate a hearty breakfast, and watched a wee Munchk~ 545

Unfortunately, this got me any line containing a word starting with "heart" like "hearty" and "heartache." Let's rewrite that command to request an exact match.

heart <- WOz[grep("\\bheart\\b",WOz$text),]
head(heart)
## # A tibble: 6 x 3
##   gutenberg_id text                                                  line 
##          <int> <chr>                                                 <chr>
## 1           55 and press her hand upon her heart whenever Dorothy's~ 110  
## 2           55 "\"Do you suppose Oz could give me a heart?\""        935  
## 3           55 brains, and a heart also; so, having tried them both~ 975  
## 4           55 "rather have a heart.\""                              976  
## 5           55 grew to love her with all my heart.  She, on her par~ 992  
## 6           55 now no heart, so that I lost all my love for the Mun~ 1024

Now I can use this additional context to determine which instances of "heart" are about the Tin Man's quest.

It seems like these tools would be most useful for open-ended responses that are still very categorical, but it could help with determining themes for more narrative data.

Saturday, July 28, 2018

New Twitter to Follow

If you enjoy Flash Fiction as much as me, you definitely want to check out the Micro Flash Fiction Twitter account. Here's a taste:

Someone liked this one so much, they illustrated it:

Wednesday, July 25, 2018

Blogging Break

You may have noticed I haven't been blogging as much recently. Though in some aspects, I'm busier than I've been in a while, I still have had a lot of downtime, but sadly not as much inspiration to write on my blog. I've got a few stats side projects I'm working on, but nothing to a point I can blog about, and I'm having difficulty with writing some of the code on projects I've been plan on writing about. Hopefully I'll have something soon, and will get to back to posting weekly Statistics Sunday posts.

Here's what's going on with me currently:

  • I had my first conference call with my company's Research Advisory Committee last night, a committee I imagine I'll inherit as my own now that I'm Director of Research
  • I submitted my first novel query to an agent earlier today and received a confirmation email that she got it
  • I've been reading a ton and apparently am 5 books ahead of schedule on my Goodreads reading challenging: 38 books so far this year
  • The research center I used to work for was not renewed, so they'll be shutting their doors in 14 months; I'm sad for my colleagues
  • Today is my work anniversary: I've been at my current job 1 year! My boss emailed me about it this morning, along with this picture:

Friday, July 6, 2018

Lots Going On

So much going on right now that I haven't had much time for blogging.
  • I'm almost completely transitioned from my old department, Exam Development, to my new one, Research, of which I am Director and currently the only employee. But my first direct report will be coming on soon! We also have a newly hired Director of Exam Development, and we've already started chatting about some Research-Exam Development collaborative projects.
  • I'm preparing for multiple content validation studies, including one for a brand new certification we'll be offering. So I've been reviewing blueprints for similar exams, curriculum for related programs, and job descriptions from across the US to help build potential topics for the exam, to be reviewed by our expert panel.
  • I've been participating in Camp NaNoWriMo, which happens in April and July, and allows you to set whatever word/page/time/etc. goals you'd like for your manuscript or project. My goal is to finish the novel I wrote for 2016 NaNoWriMo, so I'm spending most of the month editing as well as writing toward a goal of 8,000 additional words.
  • Also related to my book, I got some feedback from an agent that I need to play up the mystery aspect of my book, and try to think of some comparative titles and/or authors - that is, "If you liked X, you'll like my book." So in addition to doing a lot writing, I'm doing a lot of reading, looking for some good comparative works. I've asked a few friends to read some chapters and see if they come up with something as well.
  • I started recording the promised Mixed-Effects Meta-Analysis video earlier this week, but when I listened to what I recorded, you can clearly hear my neighbors shooting off fireworks in the background. So I need to re-record that part and record the rest. Hopefully this weekend.
Bonus writing-related pic, just for fun: I found a book title generator online, and this is the result I got for Mystery:


That's all for now! Hopefully back to regular blogging soon!

Thursday, June 14, 2018

Working with Your Facebook Data in R

How to Read in and Clean Your Facebook Data - I recently learned that you can download all of your Facebook data, so I decided to check it out and bring it into R. To access your data, go to Facebook, and click on the white down arrow in the upper-right corner. From there, select Settings, then, from the column on the left, "Your Facebook Information." When you get the Facebook Information screen, select "View" next to "Download Your Information." On this screen, you'll be able to select the kind of data you want, a date range, and format. I only wanted my posts, so under "Your Information," I deselected everything but the first item on the list, "Posts." (Note that this will still download all photos and videos you posted, so it will be a large file.) To make it easy to bring into R, I selected JSON under Format (the other option is HTML).


After you click "Create File," it will take a while to compile - you'll get an email when it's ready. You'll need to reenter your password when you go to download the file.

The result is a Zip file, which contains folders for Posts, Photos, and Videos. Posts includes your own posts (on your and others' timelines) as well as posts from others on your timeline. And, of course, the file needed a bit of cleaning. Here's what I did.

Since the post data is a JSON file, I need the jsonlite package to read it.

setwd("C:/Users/slocatelli/Downloads/facebook-saralocatelli35/posts")
library(jsonlite)

FBposts <- fromJSON("your_posts.json")

This creates a large list object, with my data in a data frame. So as I did with the Taylor Swift albums, I can pull out that data frame.

myposts <- FBposts$status_updates

The resulting data frame has 5 columns: timestamp, which is in UNIX format; attachments, any photos, videos, URLs, or Facebook events attached to the post; title, which always starts with the author of the post (you or your friend who posted on your timeline) followed by the type of post; data, the text of the post; and tags, the people you tagged in the post.

First, I converted the timestamp to datetime, using the anytime package.

library(anytime)

myposts$timestamp <- anytime(myposts$timestamp)

Next, I wanted to pull out post author, so that I could easily filter the data frame to only use my own posts.

library(tidyverse)
myposts$author <- word(string = myposts$title, start = 1, end = 2, sep = fixed(" "))

Finally, I was interested in extracting URLs I shared (mostly from YouTube or my own blog) and the text of my posts, which I did with some regular expression functions and some help from Stack Overflow (here and here).

url_pattern <- "http[s]?://(?:[a-zA-Z]|[0-9]|[$-_@.&+]|[!*\\(\\),]|(?:%[0-9a-fA-F][0-9a-fA-F]))+"

myposts$links <- str_extract(myposts$attachments, url_pattern)

library(qdapRegex)
myposts$posttext <- myposts$data %>%
  rm_between('"','"',extract = TRUE)

There's more cleaning I could do, but this gets me a data frame I could use for some text analysis. Let's look at my most frequent words.

myposts$posttext <- as.character(myposts$posttext)
library(tidytext)
mypost_text <- myposts %>%
  unnest_tokens(word, posttext) %>%
  anti_join(stop_words)
## Joining, by = "word"
counts <- mypost_text %>%
  filter(author == "Sara Locatelli") %>%
  drop_na(word) %>%
  count(word, sort = TRUE)

counts
## # A tibble: 9,753 x 2
##    word         n
##    <chr>    <int>
##  1 happy     4702
##  2 birthday  4643
##  3 today's    666
##  4 song       648
##  5 head       636
##  6 day        337
##  7 post       321
##  8 009f       287
##  9 ð          287
## 10 008e       266
## # ... with 9,743 more rows

These data include all my posts, including writing "Happy birthday" on other's timelines. I also frequently post the song in my head when I wake up in the morning (over 600 times, it seems). If I wanted to remove those, and only include times I said happy or song outside of those posts, I'd need to apply the filter in a previous step. There are also some strange characters that I want to clean from the data before I do anything else with them. I can easily remove these characters and numbers with string detect, but cells that contain numbers and letters, such as "008e" won't be cut out with that function. So I'll just filter them out separately.

drop_nums <- c("008a","008e","009a","009c","009f")

counts <- counts %>%
  filter(str_detect(word, "[a-z]+"),
         !word %in% str_detect(word, "[0-9]"),
         !word %in% drop_nums)

Now I could, for instance, create a word cloud.

library(wordcloud)
counts %>%
  with(wordcloud(word, n, max.words = 50))

In addition to posting for birthdays and head songs, I talk a lot about statistics, data, analysis, and my blog. I also post about beer, concerts, friends, books, and Chicago. Let's see what happens if I mix in some sentiment analysis to my word cloud.

library(reshape2)
## 
## Attaching package: 'reshape2'
counts %>%
  inner_join(get_sentiments("bing")) %>%
  acast(word ~ sentiment, value.var = "n", fill = 0) %>%
  comparison.cloud(colors = c("red","blue"), max.words = 100)
## Joining, by = "word"

Once again, a few words are likely being misclassified - regression and plot are both negatively-valenced, but I imagine I'm using them in the statistical sense instead of the negative sense. I also apparently use "died" or "die" but I suspect in the context of, "I died laughing at this." And "happy" is huge, because it includes birthday wishes as well as instances where I talk about happiness. Some additional cleaning and exploration of the data is certainly needed. But that's enough to get started with this huge example of "me-search."