The SAT bias finding that appears only when you stop averaging
Averaged over 200 institutions, the SAT predicts first-year grades the same way for everyone. Taken college by college, it does not.
8 min read
An admissions test is not usually accused of being wrong about a student. It is accused of being wrong about a group, which is a different and harder claim, and the technical term for it is differential prediction: the test predicts the outcome differently depending on which group the candidate belongs to.
That is not the same thing as a score gap. Two groups can average different scores and still have their eventual grades predicted with equal accuracy. Conversely a test can produce identical averages and still systematically over-predict for one group and under-predict for another. Only the second is a prediction problem, and it is the one that decides whether an admissions or scholarship process treats candidates fairly.
The argument before this paper
In 2010, Herman Aguinis, Steven A. Culpepper and Charles A. Pierce published a paper in the Journal of Applied Psychology arguing that the methods used by the College Board and others to test for bias in admissions and employment testing may be flawed.
The distinction they drew then is worth preserving, because it was lost in most reporting. They did not conclude that the tests were biased. They concluded that the tests had the potential to be biased and that the standard methods for revealing bias were deficient. That is a methodological claim, and it is falsifiable in a specific way: apply better methods to real data and see whether anything appears.
In 2013 two research scientists then at the College Board, Krista Mattern and Brian F. Patterson, published a response in the same journal. Their objection was that the 2010 paper rested on a simulation. They used actual data — more than 475,000 students at more than 200 colleges between 2006 and 2008 — examined the relationship between SAT scores and first-year grade point average, and found that on average the relationship was the same across groups.
The part that made the reanalysis possible
The Journal of Applied Psychology required Mattern and Patterson to make the College Board’s data available. It appeared, for the first time, as a 400-page PDF.
Aguinis, Culpepper and Pierce extracted the data from that document. Their 2016 paper in the Journal of Educational Psychology, Differential Prediction Generalization in College Admissions Testing, is built on it.
The sequence deserves noting on its own. A methodological dispute produced a data-release requirement, the data-release requirement produced an awkward but usable artefact, and the artefact enabled the reanalysis. None of that was anyone’s plan.
What they found
They began by replicating. “The first thing we did was to do what they did, exactly what they did,” Aguinis said, and the result held: averaged across roughly 200 colleges, their findings matched the College Board’s.
Then they stopped averaging. Examined institution by institution, the picture broke apart. In the majority of colleges they found differential prediction, running in one direction at some and the opposite direction at others. Hundreds of thousands of students’ predicted grade point averages, based on SAT scores, were over- or under-estimated.
The comparison set was large: 257,336 female and 220,433 male students across 339 samples, and 29,734 African American and 304,372 white students across 264 samples, from 176 institutions.
Aguinis was direct about the direction, and this is the part most summaries omit. The effect is not uniformly against any one group. “Not against blacks or against women necessarily. It goes both ways. The paper is about predicting performance for all people, and the bias we found sometimes benefits one group and some other times the other.”
Why averaging destroyed the signal
The explanation is not about the test paper. It is about the thing being predicted.
A first-year grade point average is produced by an institution: its admissions policy determines who is in the room, its grading conventions determine what a given piece of work is worth, and its academic support resources determine who recovers from a difficult start. All three differ sharply between institutions, and often within them between departments.
If the outcome variable is institution-specific, then the relationship between test score and outcome is institution-specific too. An average across 200 colleges is an average across 200 different relationships, and opposing deviations cancel. The College Board’s result was not wrong; it answered a question about the aggregate, and the aggregate is not where an admissions decision happens.
What follows from it
Aguinis has been careful about the conclusion, and it is not the one the original coverage implied. He does not argue that the SAT is irrelevant; standardised tests remain among the better available predictors of academic success. He argues that the results must be interpreted within the local context of the institution using them, and that an admissions office cannot assume a validity claim published elsewhere applies to its own intake.
The scope is also wider than the SAT. The same reasoning applies to the GRE, the GMAT, civil service examinations and pre-employment tests measuring similar abilities: wherever a test is validated against an outcome that a local organisation defines, the validation is local.
His practical recommendation is unglamorous and hard to argue with. Find out how the test behaves in your own context, because it may be over-predicting for some groups and under-predicting for others, and the aggregate literature will not tell you which.
Questions
4 answeredWhat is differential prediction?
A test shows differential prediction when it predicts an outcome differently for different groups. It is not the same as a score gap. Two groups can score differently and still have their grades predicted equally well, or score identically and have their grades predicted unequally.
Did the researchers say the SAT is biased?
Not as a blanket statement. Their 2010 paper argued that the methods used to detect bias were deficient and that tests therefore had the potential to be biased. The 2016 paper found actual differential prediction in the majority of the colleges examined, running in both directions depending on the institution.
Do they think the SAT should be abandoned?
No. Aguinis has been explicit that standardised tests are among the best available measures of future academic success. His argument is that the results have to be interpreted in the local context of a specific institution rather than assumed to transfer.
How much data was involved?
They compared 257,336 female and 220,433 male students across 339 samples, and 29,734 African American and 304,372 white students across 264 samples, drawn from 176 colleges and universities between 2006 and 2008.
Sources
3 referenced
The dragon ants of New Guinea, and what X-rays found inside
Two ant species from Papua New Guinea got Game of Thrones names. The reason the paper matters is what microtomography revealed under the spines.
Ocean currents explained: the map, the types and the climate
An ocean currents map is really two maps printed on top of each other. Wind drives the top few hundred metres; density drives everything below.
New species discovered: how many each year, explained
How many new species are discovered each year is a question with a surprisingly firm answer, and it is much larger than most people guess.