
AI as the New Gatekeeper: Whose Knowledge the Model Was Built to See
Listen to The Cultural Context of Knowledge on one of your favorite podcast players.
- 00:00Cold open — two students, same model, two outcomes
- 02:00The reveal — what the model actually returned
- 03:00Where this episode sits — stepping into the new space
- 04:00The choice of a question
- 05:30Pause and reflect — your own deep knowledge
- 06:15What confabulation actually is
- 08:30Who pays — the asymmetric harm
- 10:00Cultural context check — confabulation as a power relationship
- 11:30What accountability could look like — culturally responsive AI
- 13:30Moving the cost — from the powerless to the powerful
- 14:15Do this this week
- 16:00Landing line
“A model that fills a silence with itself has not answered. It has spoken over you.
Read the full transcript
A high-school senior in Detroit asks an AI to help her write a paper. She is writing about jazz, about the city's musical inheritance, about the families who built it, about the relationship between Black migration and the sound that came out of it.
She types her question. The model responds in seconds. The answer is fluent. It is structured. It cites the right people. Her paper gets a solid grade.
That same afternoon, in another city, a different high-school senior asks the same model to help her write a paper. She is writing about her grandmother's tradition of healing, the herbs her grandmother used, what the herbs meant, the prayers her grandmother said in her own language while she gathered them.
She types her question. The model responds in seconds. The answer is fluent. It is structured. It cites the right-sounding people.
Her paper also gets a solid grade.
end chunk.
But what the model gave her about jazz, actually existed in the model's training data. The model had read a thousand books about jazz. It was, in some real sense, returning what is known.
What the model gave her about her grandmother's tradition, was mostly invented. The model had read very little about it. The names it cited were partly real and partly fabricated. The interpretations it offered were borrowed from a small handful of outside ethnographies, not from the tradition itself. The model did not say so.
The first student got a paper grounded in a tradition the world had recorded.
The second student got a paper that sounded equally confident, but had been quietly invented around a tradition the model had not been given to know.
Same query. Same model. Same teacher reading the result. Same grade.
But what these two students just learned about whose knowledge counts, is not the same.
end chunk.
Episode 7. AI as the New Gatekeeper.
Last episode, we went inside the student. We named what schools have been doing to many children of color for decades, the developmental harm researchers have called curriculum violence, and the daily mechanism through which it gets transmitted.
Today, we step into a new space. A space many learners now enter first when they want to know something. Not the library. Not the textbook. The model.
For longtime listeners, this is the AI thread we started in Season 1, picked back up at a different angle. Season 1 asked what AI does to the process of learning, what it shortcuts, what it skips. Today we ask a different question. Not what AI does to the process of learning. But whose knowledge AI was built to see in the first place.
Most public conversation about AI in classrooms is about whether it works, whether it cheats, whether it helps. Today's episode is about something else. Today we ask what AI actually does when a learner brings their own knowledge to it. And what the answer to that question reveals about who the technology was built to serve.
end chunk.
I want to focus on a single mechanism today. The mechanism that links what we are watching to what we have been talking about all season.
The mechanism is called confabulation. Researchers in machine learning sometimes call it hallucination, but I prefer the more honest word. Confabulation. The model produces an answer that sounds right because the right answer would otherwise be a silence the model is not designed to admit.
Here is the part most people miss. Confabulation is not random. It is patterned. The model confabulates the most reliably about exactly the kinds of knowledge that were already underrepresented in the written record, the oral traditions, the languages spoken at the kitchen table but not on the publishing platform, the practices that lived in communities but never made it into the textbook.
When you ask the model about something the dominant world has recorded, you get something close to the recorded answer. When you ask the model about something the dominant world has not recorded, or has recorded only through outside observers, you get a plausible-sounding sentence built out of fragments. And the model says it with the same confidence either way.
The choice of a question, then, becomes the difference between a learner who is being assisted and a learner who is being talked over.
end chunk.
Take a moment.
Think about something you know deeply. Not from a class. From your life. From your family, your work, your community.
Imagine asking a machine to summarize that thing for you. Imagine the machine answering with confidence, and being wrong in ways only you would catch.
Now imagine being a sixteen-year-old who has never been told that the machine could be wrong about her grandmother. What does she do with the answer the machine gives her?
end chunk.
It is worth being precise about what these systems are doing, because the harm we are about to name lives in the precision.
A large language model is not a database. It does not look things up. It generates the most statistically likely next word, given everything it has been trained on. When the training data on a topic is rich and consistent, the most likely next word is also usually true. When the training data on a topic is sparse, or biased toward outside observers, or filtered through one cultural lens, the most likely next word is no longer reliably true. It is just the most likely next word.
The model has no built-in mechanism for telling the difference. It does not know which of its answers are grounded in the world and which are stitched together out of fragments. It does not say I don't know, because it has not been trained to recognize the absence of knowledge as different from the presence of knowledge. To the model, both feel the same.
What the model does, when its training data thins out, is produce something that resembles the kind of answer it would have produced if its training data had been rich. Same shape. Same fluency. Same authoritative tone. Different relationship to truth.
From the inside of the model, there is no difference. From the inside of the learner reading the model's output, there is also no difference, unless the learner already knew the answer.
And the learners who tend to already know the answer about traditions the model has barely heard, are usually the learners whose communities the model has barely heard.
end chunk.
Here is where the harm lands.
When the model confabulates about jazz, the consequence is small. The student gets a slightly wrong sentence. The teacher reading it has enough background to catch it, or the student herself has enough access to other sources to verify. The system corrects for itself.
When the model confabulates about her grandmother's healing tradition, the consequence is something else. The student does not have a teacher who can fact-check it. She does not have an external source that documents it. The only person who could correct the model is her grandmother, and her grandmother is not in the room.
What the student takes away is the model's version. Confident. Polished. Cited. Wrong in ways she will not be able to detect for years, if ever.
And here is the deeper damage. After three or four such interactions, the student starts to reorganize her sense of what counts as knowable. The traditions the model can summarize cleanly start to feel real. The traditions the model fumbles start to feel like family stories, important, but unauthorized. She starts to discount what her grandmother gave her in favor of what the model can produce.
That is not a bug in the technology. That is the technology doing exactly what it was built to do. It is generating fluent text on demand. It just happens that the cost of its fluency falls on the learners whose traditions were never in the training data to begin with.
end chunk.
I want to bring this back to the relationship at the center of the season. Knowledge and power.
The model is not a neutral observer. It is a system that has already chosen who is the default. It made that choice not through malice but through composition, what was in the data it was trained on, who had the means and platform to publish in the languages and formats the model could read, whose knowledge had been institutionally recorded and whose had been kept in families.
The model inherited the power asymmetry of its training data. Then it scaled the asymmetry by being available, instantly, to every learner with a device. Then it dressed the asymmetry in the language of universal access.
Naming what we have just described, a system that systematically returns confident answers about some traditions and confidently invents about others, as a power relationship rather than a technical limitation is the move this episode is asking the listener to make.
The technical-limitation framing keeps the problem inside the engineering team. The power-relationship framing puts the problem where it actually lives. With the people on the receiving end of the model's confidence.
end chunk.
We have named the harm. The harder question is what an accountable response would look like. Not a perfect AI, that is not on offer. But an AI accountable to the people on the receiving end of its confidence.
There is already a framework for this kind of accountability, in education. Gloria Ladson-Billings named it. Geneva Gay built it. Django Paris extended it. The frame is called culturally responsive teaching, and culturally sustaining pedagogy.
The insight at the center is that good education does not just try to reach learners from non-dominant communities. It is built, in part, from the knowledge those communities already carry. Their elders, their families, their traditions are not the audience. They are the source.
Apply that frame to AI. A culturally responsive model would not be trying to be less biased after the fact. It would be being built with the active participation of the communities whose knowledge it claims to summarize. Elders in the design loop. Scholars from the represented communities sitting next to the engineers. Training data that leads with the recordings, the interviews, the writings, the practices that the dominant publishing record left out.
end chunk.
Some of this work already exists. Indigenous communities have led the way with what is called the CARE Principles for indigenous data governance, the principle that Indigenous communities hold authority over how data about them is gathered, stored, and used.
There are early projects in which communities themselves curate the training data for models that summarize their traditions. There are research teams calling for audits of AI systems that include community participation, not just engineering review.
These are early moves. They are not yet the standard. But they are the model for what it could mean to hold a system accountable to the people it claims to know.
And here is the deeper accountability move, the one that brings us back to the relationship we have been tracing all season.
Right now, when the model is wrong about a community, the cost falls on the learner from that community. An accountable system would put the cost on the system itself.
That is what accountability is. It is moving the cost from the powerless to the powerful.
end chunk.
The next time you ask a model a question and the answer comes back clean and sure, ask one more question.
Whose knowledge would have to have been in the training data for this answer to be more than a fluent guess?
Sometimes the answer is, most of the people whose voice would matter. Sometimes the answer is, almost none of them. Both answers are useful. The model rarely tells you which one it is.
For learners, here is the practice. Pick one thing you know deeply, from your family, your faith, your community, your work. Ask a model about it. Read the answer with the eye of someone who knows. Notice what it gets right. Notice what it gets confidently wrong. Notice what it never would have known to mention.
What you are calibrating, in that exercise, is not your trust in the model. It is your trust in your own knowing. The first one needed to come back. The second one was always going to.
end chunk.
For educators, here is the practice. The next time you ask students to use AI for research, build the assignment so that the first step is the student writing down what they already know about the topic, from family, from community, from earlier reading, before they ever query the model. Then have them compare.
The exercise teaches two things at once. The model is useful. The model is partial. Both lessons are essential, and both have to be taught explicitly, because the surface of the output does not teach them.
For those who build, deploy, or fund AI systems, here is yours. The accountability frame is not optional. If you are in a position to influence what training data gets used, who is in the design loop, or how a model is evaluated for cultural responsiveness, the questions this episode raised are the ones to bring into your work.
Whose knowledge is the model carrying? Whose is being invented around? Who pays when the model is wrong?
end chunk.
For everyone listening, here is the broader move. The next time you read something, watch something, or hear someone speak with the smooth confidence the model has trained us to expect, pause. Ask whether what you are receiving has been built from knowledge, or whether it has been generated from the shape of knowledge.
We have been tracing, all season, a single relationship. Knowledge and power. Last episode, we asked what that relationship costs the children inside our classrooms. Today, we have asked what that relationship costs the learner who turns to a machine for what her grandmother already knew.
The model is not the end of this thread. It is one stop along it. The mechanism that produces the harm we just named did not begin with the model. It began with the question, generations ago, of what would be written down and what would be left to remembering.
Next time, we go further upstream. To the place where the curriculum that shaped what the model was trained on actually got written.
A model that fills a silence with itself... has not answered. It has spoken over you.
Until next time.
This transcript comes from the production script. Wording may differ slightly from the episode as aired.
Companion essay — AI as the New Gatekeeper: Whose Voice the Models Already Heard
Continue the season
All episodes →
Will Education Pivot With It?: Designing for the World That Already Exists
We opened this season with a question. The demographic pivot has already happened. Will education pivot with it? After nine episodes describ…

When Assessment Becomes Gatekeeping: An Instrument That Was Never Calibrated Against You
Two students take the same standardized reading test. Question fourteen is about a regatta, a sailing race. The first student has been to th…

How State Standards Get Written: Curriculum as Compromise
State standards are the most concentrated place in U.S. public education where decisions about other people's children get made by people wh…
