Chapter 1 Introduction

This book contains materials for Psychology 522/524, the first quarter graduate statistics course in the Department of Psychology at the University of Washington. It’s very much a work in progress.

A pdf version of this book can be found here: http://courses.washington.edu/psy524a/_book/Psych_524A_statistics_textbook.pdf This book was written using R’s ‘bookdown’, and its Pdf format is finicky, so there may be some formatting issues with the pdf version. I’m still working on it.

I’ve been teaching statistics at the undergraduate and graduate level for a couple of decades now. To be honest, I took on the undergraduate stats course, Psychology 315, because teaching it was easy. I had TA’d undergraduate statistics in psychology back as a graduate student in the 90’s and when I took on undergrad stats course at UW in 2010, and then graduate stats in 2013 the course material hadn’t changed in 20 years. All textbooks were pretty much the same, covering hypothesis tests, binomial distributions, correlations, simple ANOVA, all using tables in the back of the book to get p-values from known standard distributions. I used to joke that teaching 315 year after year was easy because it didn’t require me to keep up with the literature.

But then things started to change. Although it was initiated back in 1993, the statistical programming language R started to gain popularity in the 2010’s probably due to the availability of cheap, fast laptops and the push toward open source languages and free data sets. When I took over the grad stats course in 2013 everything was done in SPSS. Back then I surveyed the faculty about whether they thought it’d be useful for me to teach using R there was a resounding vote of ‘no’. So for the first few years I taught a hybrid class, each year emphasizing R more and SPSS less. By 2019 I had dropped SPSS entirely. The sounds of the students in lab course (522) has gradually switched from mouse clicking to keyboard typing.

Switching to R has lead to two major changes in the way my students learn and use statistics. The first is the elimination of tables for looking up probabilities. A significant portion of my old notes and lectures involved explaining which table to use and where to find the answer in the table. With R, this has been replaced with a single command like ‘pnorm’,‘pt’, or ‘pf’. The second major change is the transition from teaching ANOVA using sums-of-squares to using regression. R’s ‘lm’ and ‘lmer’ functions provide a natural way to conduct ANOVA tests (along with a bunch of other tests) using regression and the linear model. My lectures are now less filled with SS’s all over the board.

By the way, statisticians have always emphasized (often snidely) that ‘ANOVA is just regression’, but I’ve never found a good resource that really explains why. Hopefully my chapters here on how ANOVA is just regression will help make this link more clear.

1.1 Why can’t I just use ChatGPT?

I’m going to have to rewrite this section every year. Three years ago a student came up to me with a question on a homework problem and showed me what ChatGPT had generated - and it was nonsense. Two years ago AI could make a very good attempt at most homework problems, but did some stupid things and couldn’t interpret or generate graphs. Last year AI could easily pass this course, though I’ve found that its solutions are often more complex than necessary. Next year? It could probably teach this course and I’ll be out of a job.

Universities are way behind in understanding what AI can do, and what role AI should play in education. Here’s my take:

In the past I often used the analogy for learning statistics to learning a sport. You can watch people play a sport, like baseball, all day and you’ll learn a lot about the sport, but you’ll never be able to play it yourself. The analogy is that sitting in my lectures and reading this book is like watching baseball. But to actually learn how to conduct statistical analyses you need to play the sport, which in our case is doing the homework.

We can extend this analogy to AI, where AI puts you in a different role; not watching or playing, but perhaps managing the team. Is being the manager enough? Interestingly, of the 21 managers of the Seattle Mariners, every one of them played some sort of professional baseball (thanks ChatGPT for that statistic). I feel like this analogy holds. Ultimately you want to be an expert at conducting statistical analyses. AI increasingly lets you act as the manager rather than the player. You can ask it to fit the model, write the R code, make the graph, or even suggest an analysis. But managing well requires knowing what good play looks like. If you’ve never fitted a regression, struggled with a messy data set, or tried to work out why an analysis went wrong, it’s much harder to recognize when AI has made a poor choice. The point of doing the homework yourself isn’t that you’ll always need to do statistics this way. It’s that doing it yourself is part of learning enough statistics to use AI well.

Last year was interesting. One hundred percent of the grade was based on homework and the final project, and the students did as well as, or better than, students in previous years. But partway through the quarter, when I asked some fairly general questions that required an understanding of the material, I was surprised by how little they seemed to know. Interestingly, I think they were surprised too.

This points to a danger in using AI to learn. Getting an answer from AI can be deceptive. You may read its solution and understand every step. But understanding a solution when it is in front of you is not the same thing as being able to generate that solution yourself.

Psychologists have known versions of this for a long time. One influential idea is levels of processing: Craik & Lockhart (1972) argued that information is remembered better when it is processed more deeply and meaningfully. Related work has shown a generation effect: we remember things better when we have to generate them ourselves rather than simply read them. And research on retrieval practice shows that having to retrieve and use knowledge strengthens later learning. Some of the effort involved in solving a problem—the false starts, figuring out what approach to take, and realizing why something didn’t work—isn’t just an inconvenience. It’s part of what causes the learning.

AI can remove much of that effort. That’s often exactly what we want from a useful tool, but during learning it can be a problem. It’s a little like the old temptation to turn to the answers in the back of the book. Once you see the answer, it can seem perfectly obvious, and you may feel that you understand it. The real test is to close the book—or close ChatGPT—and try the next problem yourself. This is why, sadly, I’ve brought in-class exams back into the course.

By the end of this course, I don’t care whether you can do every statistical calculation without assistance. I do care whether you can recognize the problem, choose an appropriate analysis, understand what the analysis is doing, evaluate whether the result makes sense, and explain what it means. AI can help with all of those things, but it can’t substitute for learning them.

1.2 Course policy for AI

Early in the course, I recommend using AI mainly as a tutor rather than as a problem-solving machine. Start by trying the problems yourself. Decide what kind of problem you think it is, what analysis might be appropriate, and make a serious attempt. If you get stuck, that’s a good time to use AI. Ask it for a hint. Ask it to explain a concept in a different way. Show it your reasoning and ask where you went wrong. Ask it why a particular statistical test is appropriate, or what an error message in R means. These are all good uses of AI because they keep you actively involved in solving the problem. You can even have it generate similar problems for you to try.

What I don’t recommend, especially when you are first learning a topic, is copying a homework problem into AI and asking it for the complete solution before you’ve tried it yourself. You may understand the answer perfectly once you see it, but that is not the same as having learned how to produce the answer.

Later, once you understand the statistical ideas, I want you to become more comfortable using AI as a tool for doing statistics. Let it help write routine R code, diagnose programming errors, reshape data, suggest ways to visualize a result, or remind you of the syntax for a method you already understand. You can even ask it to suggest analyses you may not have considered. But at that stage your role changes: you are the manager, responsible for judging what AI gives you. Does the proposed analysis answer the question you care about? Are its assumptions reasonable? Is the graph misleading? Does the result make sense? Is there a simpler way to do it?

That is ultimately the skill I want you to develop. You should not leave this course believing that good statisticians are people who can do everything without AI. Increasingly, good statisticians will be people who know enough statistics to use AI well. The goal is to learn enough statistics that AI becomes a tool you control rather than a source of answers you depend on.