Home/Episode Companions/AI as the New Gatekeeper: Whose Voice the Models Already Heard

Episode companionS2 · E7May 18, 2026

AI as the New Gatekeeper: Whose Voice the Models Already Heard.

A companion essay to Season 2, Episode 7 of The Cultural Context of Knowledge: “AI as the New Gatekeeper.”

The fourth gate

Episode 2 of this season named three gates the field has long recognized as the places where knowledge becomes legitimate. Peer review. Citation. Accreditation. The gates are imperfect and they have histories, but everyone in higher education at least knows where to stand and argue with them.

A fourth gate has arrived without a hearing. Generative language models are quietly being asked to draft lesson plans, tutor undergraduates, evaluate essays, summarize the research base of a field, recommend what a learner should study next, and, in some early experiments, score work. None of these are small uses. Each one is a place where the system is doing what gatekeepers have always done: deciding what counts.

The episode opens with the line that the models learned to summarize a world they never visited. That is the move the rest of this essay is built on. A language model does not see classrooms. It has not sat with a learner. The model has absorbed the residue of how the dominant voice on the internet has talked about classrooms, learners, knowledge, and worth. When the model answers a question, it answers in the register of whoever it heard most.

Where the training data came from

The honest summary of how large language models are trained, told plainly: companies scrape enormous slices of the public web, license additional text from publishers, add curated books and academic papers when available, and tune the result on human-labeled examples of “good” and “bad” answers. The web layer carries the most weight by volume. Within that layer, English dominates. Within English, a handful of high-traffic domains — encyclopedias, Q-and-A sites, news, and well-indexed academic and technical sources — contribute outsized signal.

This is not a conspiracy. It is a sampling story. And researchers in cognitive science have a name for what happens when a sample looks like this: WEIRD. Western, educated, industrialized, rich, democratic. In 2022, Gutchess and Rajaram published a warning to memory researchers that the cognition literature had over-relied on WEIRD samples for decades and was generalizing findings as if all minds work the way college sophomores at a few research universities work. That warning applies almost exactly to the corpus from which today’s language models were trained.

What lands, then, is a model that has heard one slice of the world deeply and other slices faintly. Not silenced, faintly. There is a difference. Silence would at least be honest. Faint signal is more dangerous because it lets the system answer with confidence about communities it barely knows.

What the research says about how this lands in classrooms

Studies on representational harm in language models are not speculative. Audits have found that, when prompted neutrally, the models more often associate names from some communities with negative attributes than names from other communities, produce thinner and more stereotyped descriptions of non-dominant cultures, struggle to render dialects of English that are not the academic-prestige register, and treat the rhetorical conventions of those dialects as errors rather than as conventions. None of this is news to people who work in equity. The news is that the same patterns are now being installed inside the tools that draft, tutor, and score.

A first-generation college student who writes a brilliant essay drawing on lived experience and her grandmother’s voice now has two readers who can mark her work “not academic.” The professor from Episode 6, and the model. The professor at least leaves a margin note she can argue with. The model leaves a rubric score and a rewrite suggestion that quietly converts her voice into the register the model heard most.

This is the misrecognition tax from the show’s theoretical anchor, automated. The extra cognitive labor that marginalized learners pay to translate their knowledge into dominant norms used to be paid one classroom at a time. The translation is now being asked for by a piece of software that runs at the speed of a keystroke.

Cultural context check

There is a version of this conversation that says: the model is neutral, the user just needs to prompt it better. Set that version aside for a moment and notice what it is doing. It is taking a system whose defaults were trained on a particular voice and asking the people who do not match that voice to do extra work to correct for it. The defaults are not neutral. The defaults are a position. A learner who has to add three sentences of context to every prompt to get the model to recognize that her community has a name is paying a tax. The same tax the show has been naming for two seasons.

There is also a version that says: the model is a tool, no different from a calculator. A calculator computes within a closed mathematical system whose answers do not depend on whose history you grew up inside. A language model does the opposite. Every answer is a function of what voices the corpus heard. Calling it a calculator is a category error, and the category error is not innocent. The error makes it easier for institutions to adopt the tool without asking the legitimacy question Episode 2 of this season was built around.

What this means for the legitimacy machine

Season 2 has been tracing the machine that decides what counts as knowledge. The pieces named so far are universities, journals, peer review, citation, accreditation, who gets to teach, what gets called violent, and the hidden curriculum that scores some voices as academic and others as not. The pattern of the season is that each piece looks neutral from the inside and turns out to carry a position from the outside.

Generative AI is the newest piece. The honest thing to say about it is that it inherits all of the prior positions and adds one more. Universities decided what counted. Journals decided what got cited. Accreditors decided what counted as legitimate study. The next system in the chain has already decided whose voice sounds like knowledge. The decision was made before any educator sat down with a syllabus and asked whether to use the tool.

This is the through line from Season 1’s opening AI episodes. The show began by asking whether AI could be a learning companion. The question is not whether the companion is helpful. The question is whose companion the companion is.

Pause and reflect

Three prompts. Take a moment with whichever one finds you.

One. Think of a piece of knowledge from your community, your family, or your professional practice that has never been written down in the form an academic publisher would accept. Now ask whether a language model trained on the public internet would have heard it. If the answer is no, what happens when that model is the first reader a learner from your community will meet?

Two. The next time you use a generative tool to draft, summarize, or tutor, notice the register it returns. Whose voice does that register sound like? What would the same paragraph sound like if a teacher you learned from were writing it?

Three. If a fourth gate has arrived in the legitimacy machine, who in your setting is allowed to argue with it? Faculty senate has rules for arguing with peer review and accreditation. What is the equivalent rule for arguing with the model?

Do this this week

One concrete action, this week, depending on which seat you sit in.

If you are a learner. Save one prompt and one response, and bring them to a person whose voice you trust. Read the response aloud together. Ask whether the response sounds like your voice or like a borrowed one. The point is not to refuse the tool. The point is to keep ownership of your voice while you use it.

If you are an educator. Pick one assignment you have considered running through a generative tool, either to draft or to score. Run it. Then read the output next to a student’s actual work. Notice what the tool flagged as weak that you would have read as voice. Decide, before the next class, what you will do with the gap.

If you are a person in a general audience. The next time a news story tells you a U.S. school district has adopted an AI tutor, ask which corpus the tutor was trained on. The question will be hard to answer. The fact that it is hard to answer is the point.

Landing line

The first three gates of legitimacy took a century to build and another century to argue with. The fourth gate took five years. The voice on the other side of it is the voice the model heard most, and the model is not finished listening.

Next episode looks at the room where state standards are written. Picture the table. Picture who is at it. The question is whose absence has already been written into the document by the time the room is called to order.

Dr. Donald Easton-Brooks

About the author

Dr. Donald Easton-Brooks

Scholar, author of Ethnic Matching (Rowman & Littlefield, 2019), and host of The Cultural Context of Knowledge. Research on representation, the teacher workforce, and whose knowledge counts as knowledge.

S3 · E10Beyond the Schoolhouse Door: Ethnic Matching in Higher Education (S3 E10)
0:00 · 15:01