Episode 20 · Season II · S2.E9 · May 20, 2026 · 18:46

When Assessment Becomes Gatekeeping: An Instrument That Was Never Calibrated Against You

Two students take the same standardized reading test. Question fourteen is about a regatta, a sailing race. The first student has been to the harbor every summer of her life. The second has never seen a regatta. The test reports the first student as a stronger reader. What the test measured was not reading comprehension. It was access to a particular cultural setting. This episode names the standardized test as the closing instrument of the legitimacy machine, names curriculum and assessment as a pair, and asks what an accountable assessment system would actually look like.

Share

Listen to The Cultural Context of Knowledge on one of your favorite podcast players.

In this episode
  1. 00:00Cold open: two students, the regatta
  2. 02:00The reveal: what the test actually measured
  3. 03:00Where this episode sits in Season 2
  4. 04:15Curriculum and assessment, paired
  5. 05:45What standardized assessment actually does
  6. 09:00Assessment as verdict, not measurement
  7. 10:30Cultural mismatch and stereotype threat
  8. 12:00Who pays for the mismeasurement
  9. 15:00What accountability could look like
  10. 17:00The deeper accountability move: the instrument, not the children
  11. 18:00Do this this week
  12. 19:30Landing line
A number issued by an instrument that was never calibrated against you is not a verdict. It is the instrument telling on itself.
Transcript
Read the full transcript

Two students sit down to take the same standardized reading test.

They have the same teacher. They are the same age. They have the same hours of homework behind them. They have, by the kinds of measures the test is designed to control for, the same readiness.

Question fourteen on the test is about a regatta, a sailing race. The passage describes a young person watching the boats come in at the end of the race. The vocabulary is specific. Starboard. Jib. Luff. Tacking.

The questions ask the reader to identify the writer's purpose, to infer the speaker's emotional state, to choose the best title for the passage.

end chunk.

The first student has been to the harbor with her family every summer of her life. She knows these words from the dock. The passage is a story she has heard before, a younger sibling watching the older ones race. She answers the questions in two minutes.

The second student has never seen a regatta. She has read about boats. She has not read about boats like these. She works through the vocabulary by context. She gets some of the questions right. She gets some wrong. She finishes in nine minutes.

The test reports the first student as a stronger reader than the second student.

What the test measured was not reading comprehension. What the test measured was access to a particular cultural setting.

But the score that gets entered into the record does not say that. The score says reading comprehension. And the score will follow the second student into every conversation about her academic potential for years to come.

end chunk.

Episode 9. When Assessment Becomes Gatekeeping.

Across this season, we have traced one relationship. Knowledge and power.

Last episode, we walked into the meeting where state standards get written and named that meeting as the place where decisions about other people's children get made by people who do not have to live with the consequences.

Today, we follow the standards document one step further into the system that it produces. To the closing instrument of the legitimacy machine. The standardized test.

Most public conversation about testing is about whether tests are too high-stakes, or too narrow, or too frequent. Today's episode is about something else. Today we ask what testing actually does. Not whether it measures well. But what it defines as worth measuring in the first place. And what it costs the children whose cultural setting was not the one the test was designed around.

end chunk.

Before we go further, I want to name a relationship that often goes unstated. Curriculum and assessment are not separate stages of a system. They are a pair.

The curriculum says what should be taught. The test says what gets rewarded. The two travel together. When the test gets weight, the curriculum bends to match it.

In practice, in U.S. public education, the test almost always wins. What gets tested gets taught. What does not get tested gets taught in the margins, or it does not get taught at all.

That is why this episode follows last episode without a seam. Episode 8 walked into the meeting where the curriculum gets written. Episode 9 walks into the meeting where the test gets calibrated. The same hands are not always in both meetings. But the same logic is. Whose cultural setting is at the center, and whose is in the footnotes.

The student who was missing from the curriculum is then mismeasured by the assessment. The mismeasurement gets entered into the record as her academic identity. And the next year's curriculum gets adjusted, not to serve her, but to raise her score on the same instrument that misread her in the first place.

That is the loop. Curriculum decides what counts. Assessment certifies who learned it. And the children whose cultural setting was excluded from the first stage pay the cost at the second.

end chunk.

Let me back up and explain what a standardized assessment is, in the most precise terms we can use.

A standardized test is a measurement instrument that has been calibrated against a population. The questions have been pre-tested on a sample of students. The items have been adjusted so that the test produces a consistent distribution of scores. The score a student receives is interpreted not in absolute terms but relative to the population the test was calibrated against.

That sounds technical. Here is what it means in practice.

The test is not measuring reading. It is measuring how the student's reading compares to the reading of the population the test was built around. If the population the test was built around shares a particular cultural setting, then the test rewards that setting. Not by mistake. By design.

The standardized assessment is one of the most powerful instruments education has invented. It governs admissions. It governs placement. It governs which children get tracked into which classrooms in fourth grade. It governs which schools get labeled successful and which schools get labeled failing. It governs which teachers keep their jobs. It governs which districts get funded.

When you change the standardized test, you change all of that downstream of it. When you don't change the standardized test, none of that changes either.

end chunk.

Take a breath.

Think back to a standardized test you took as a student. The SAT, the ACT, an end-of-course state exam, a placement test in college, a certification exam at work.

Did anyone tell you, before you took it, that the test had been calibrated against a particular population? Did anyone tell you that getting a question wrong might have less to do with what you knew and more to do with whether the question was written for someone with your specific cultural inheritance?

Most of us were not told. We were told the test measured ability. And we received the score as a verdict on what we were capable of.

Here is the frame I want to introduce.

An assessment is not a measurement of learning. It is a definition of what counts as learning. And what the test defines as learning becomes who is allowed to be called a learner.

end chunk.

When the test defines reading comprehension as the ability to answer questions about a regatta, the test has just decided that the cultural setting of the regatta is part of what reading comprehension means. The student who lacks that setting is not just at a disadvantage on this question. She has been told, by a system she has no power to talk back to, that her form of reading does not count as the form the test recognizes.

Researchers who study standardized testing have documented this for decades. The technical name for one piece of it is cultural mismatch in test items, questions whose vocabulary, references, or implicit framing assume a particular cultural setting that the test takers do not all share.

The decades of research on stereotype threat, beginning with Steele and Aronson, have documented another piece, the cognitive cost of taking a high-stakes test while aware that one's group is expected to perform poorly on it. That awareness alone, controlled for ability, lowers performance. The test is not measuring the student. It is measuring the student plus the weight of the room.

And then the score becomes the record. The record becomes the placement. The placement becomes the trajectory. The trajectory becomes who is read as a strong learner and who is read as needing remediation, by the very system that built the question that mismeasured the student in the first place.

end chunk.

We have named this pattern in different forms across the season.

When the model confabulates about a tradition the model has barely heard, the cost falls on the learner whose tradition was barely in the training data. When the standards committee writes a thin compromise about a community's history, the cost falls on the children of that community. When the test is calibrated against a particular cultural setting, the cost falls on the students whose cultural setting was not the one the test was built around.

In every case, the people designing the instrument do not pay the cost of the design. The people on the receiving end of the instrument pay it.

The standardized test is the place where the cost gets converted into a number. The number gets attached to the student. The number becomes part of what teachers, college admissions officers, and employers see when they meet her later. She does not get to walk into those conversations with the explanation. She gets to walk in with the number.

That is what assessment as gatekeeping means. It does not just measure. It certifies. It hands the student a verdict, written in the language of the people who built the test, about who she is allowed to be called as a learner.

end chunk.

I want to bring this back, one more time, to the relationship at the center of the season. Knowledge and power.

The standardized test is the most concentrated point in U.S. education where the question of whose knowledge counts produces a measurable verdict on a specific child. Not in theory. With a number. Filed in a record. Used to make decisions about her placement and her opportunities for the next decade.

And the legitimacy of that verdict comes from a single, powerful idea. The test is objective. The test is the same for everyone. The test removes the bias of human judgment.

Each of those claims is partially true. Standardized tests do remove some kinds of bias, the favoritism, the snap judgment, the teacher who likes one student more than another. Those are real harms, and the test does control for them.

The harm we have just named is a different harm. It is the harm of a measurement instrument calibrated against a population that does not include all the children taking it.

Removing one kind of bias by introducing another, more credible kind of bias is not progress. It is a transfer of who pays.

end chunk.

We have named the harm. The harder question is what accountable assessment could look like.

The same framework that has carried us through the last two episodes is useful here. Culturally Responsive Practices insist that good teaching is built from the knowledge non-dominant communities already carry. Their elders, their families, their traditions are not the audience. They are the source.

Apply that principle to assessment. An accountable assessment system would not just try to remove obviously biased items in the year-end review. It would be co-designed with the communities whose children take it. It would offer multiple modes of demonstrating knowledge, written, oral, performance, project, so that the form of the demonstration does not become a hidden test of cultural fluency. It would be honest about what it can and cannot measure, and it would refuse to issue verdicts about what it has not actually measured.

Some of this work already exists. Performance-based assessment networks have been building portfolio-based alternatives in some U.S. school districts for over twenty years. Indigenous nations have asserted assessment sovereignty alongside curriculum sovereignty in some tribal-affiliated public schools. These are early moves. They are not yet the standard. But they are the model for what an accountable assessment system could look like.

And it cannot stand alone. Accountable assessment cannot exist without accountable curriculum. The two have to be redesigned together. If you fix the curriculum and leave the test alone, the test will pull the curriculum back. If you fix the test and leave the curriculum alone, the curriculum will keep teaching to the old form. The pair has to move together. That is the work.

end chunk.

And here is the deeper accountability move, the one that brings the season's argument to its closing point in this episode.

Right now, the cost of mismeasurement falls on the children mismeasured. An accountable system would put the cost on the test.

It would treat a score that systematically misreads a particular group of children as a defect of the instrument, not a property of the children.

We do that for thermometers. We do that for blood-pressure cuffs. We do that for any other measurement instrument in a serious field. We have not yet done it for the standardized tests we use to certify the academic potential of children.

end chunk.

Take a moment.

Think of a child you know whose academic record contains a number. Imagine someone telling you that the number was generated by an instrument that was never calibrated against children like her, and that the number is now the thing the system trusts most about her.

What would you want done about the instrument? What would you want done with the number?

For educators, here is the practice. The next time a student's test score does not match what you have seen them do in the classroom, write down both. Note the score. Note the specific evidence you have of what the student actually understands. Then, in your evaluation of that student, weight the evidence over the score.

The score is one data point produced by an instrument built without that student in mind. Your direct observation is another. The latter is often more accurate.

end chunk.

For parents and community members, here is yours. Find out what standardized assessments your state or district uses, what they are calibrated against, and what the published cultural-bias reviews of those tests have found. The information is more accessible than people realize. School board meetings are public. State assessment vendor selection processes have public-comment periods. The same lever that opens the standards-writing meeting opens the assessment-vendor decision.

For those who design, build, sell, or commission these tests, here is yours. The accountability frame is not optional. Whose cultural setting is the test calibrated against? Whose performance is being mismeasured? And who pays the cost when the score becomes the verdict?

For learners listening, here is a smaller move. The next time you receive a score that feels lower than what you actually understood, write down what you actually understood. Keep that record. Your sense of what you know is more reliable than the score that knew less than you did.

end chunk.

We have been tracing, all season, a single relationship. Knowledge and power. The standardized test is where that relationship hands a child a number and tells her the number is who she is allowed to be called.

It is the closing instrument of the legitimacy machine, the place where the institution converts its judgment into a measurement, and the measurement into a record, and the record into a trajectory. It is also the place where the architecture is most visible, because it produces a number you can look at.

Next time, we close the season. We return to the classroom we walked into in Episode 1, the demographic pivot, the children already in the building, the question of whether education will pivot with them. We have spent nine episodes describing the architecture. The finale asks what it would take to redesign it. And it points at the lever the next season will take up.

A number issued by an instrument that was never calibrated against you is not a verdict. It is the instrument telling on itself.

Until next time.

This transcript comes from the production script. Wording may differ slightly from the episode as aired.

Tags
MethodologyEquityPowerSchools
Stay close to the show

Get the notes, the references, and the reading list.

Join the email list and get the written companion and full reference list for every episode, plus the series reading list to start.

Continue the season

All episodes →
021 · S2.E10

Will Education Pivot With It?: Designing for the World That Already Exists

We opened this season with a question. The demographic pivot has already happened. Will education pivot with it? After nine episodes describ…

15:39
019 · S2.E8

How State Standards Get Written: Curriculum as Compromise

State standards are the most concentrated place in U.S. public education where decisions about other people's children get made by people wh…

17:51
018 · S2.E7

AI as the New Gatekeeper: Whose Knowledge the Model Was Built to See

The newest gatekeeper between learners and what they are trying to know is a model that fills silence with fluent invention. This episode na…

17:55
S3 · E10Beyond the Schoolhouse Door: Ethnic Matching in Higher Education (S3 E10)
0:00 · 15:01