By Isabel Andrade, Principal Product Manager, Khan Academy Kids

For publicly funded PreK programs serving dual language learners, it is important to give children opportunities to demonstrate what they know in the language or languages they know best. But language access is only part of the question. It is just as important to understand how the assessment was developed and validated. A translated assessment is developed and tested in one language, then translated into another. A bilingual assessment is developed in parallel across both languages and psychometrically validated in both.
If your program serves Spanish-speaking families—or any community of dual language learners — that distinction matters. An assessment can be available in Spanish and still not be a bilingual assessment. Program leaders need to know whether a tool was built and validated for bilingual children, or whether a Spanish version was added after the English assessment was already developed.
What’s the difference between a translated assessment and a bilingual assessment?
A translated assessment takes an instrument originally built for English-speaking children and renders the prompts in another language. A bilingual assessment is designed from the start to measure a child’s skills across both languages, with items calibrated for equivalent difficulty and validated with bilingual children.
The distinction matters because translation, on its own, introduces problems that look like child performance issues but are actually measurement issues.
Consider a simple counting task. If the assessment is delivered only in English to a child whose home language is Spanish, the score doesn’t tell you whether the child can count, it tells you how much English she’s picked up so far. She may be able to count to veinte at home every night and still be flagged as unable to count to five. Assessing in both languages with a bilingual assessment won’t perfectly isolate math skill, but it gets much closer: you learn what the child knows, and in which language he or she knows it.
On a recent Khan Academy Kids webinar, Dr. Sandra Barrueco, professor of psychology at the Catholic University of America and a leading researcher on bilingual preschool assessment, described the need for appropriate assessments. Inappropriate assessments can contribute to misdiagnosis in both directions: children who are developing typically as bilingual learners may be mistaken for having a delay, while children who do need language support may not be identified soon enough.
I have certainly seen my fair share of misdiagnoses of a child that was either bilingually developing in an adequate way or in other children, needing linguistic supports because they have a language delay, and people are waiting because they are multilingual.
Dr. Sandra Barrueco, Catholic University of America
Why does direct translation fail for PreK assessments?
Direct translation fails because words and grammatical structures rarely line up cleanly across languages, even when they appear to be exact equivalents.
Three examples that show up constantly in PreK assessment design:
- Words at different difficulty levels. Flip the direction for a moment: imagine a Spanish assessment translated into English. In Spanish, rápido is a word nearly every preschooler knows. Translate it literally and you get “rapid” — a word few English-speaking preschoolers use; they say “fast.” Same meaning, different difficulty. A word-for-word translation can change what an item measures. That’s why testing every item with bilingual children is so important.
- Grammatical complexity. A story translated directly from English to Spanish can shift the verb forms in ways that change the difficulty level entirely. A simple sentence in English such as “she hoped it would happen” requires the subjunctive in Spanish (esperaba que pasara)—a verb form many preschoolers don’t yet control. The oral story gets harder to follow, and the questions harder to answer.
- Dialectal variation. A strawberry is frutilla in parts of South America, fresa in Mexico and Central America, and fresón in parts of Spain. If your assessment uses only one of these, children from other Spanish-speaking communities may answer incorrectly not because they don’t know the concept, but because the word is unfamiliar. The same dynamic exists in English, where regional variation (pop versus soda, for example) can affect performance.
None of these problems are visible in the score. They show up as gaps that programs then try to close with instruction—time that could be better spent elsewhere.
Best practice…as we’re developing measures in science is to have representation of the different dialects at the table. We have such a mix of Latin American, Central American, and Caribbean Spanish speakers….So we need to have that represented in our measures.
Dr. Sandra Barrueco, Catholic University of America
What does a well-built bilingual assessment do differently?
A well-built bilingual assessment is developed simultaneously in both languages, calibrated against data from monolingual and bilingual children, and vetted by experts representing the dialectal range of the populations it will serve.
In practice, that means:
- Simultaneous authoring. Items are written in both languages from the start, not translated after the English version is finalized. Each item is reviewed for difficulty parity.
- Empirical calibration. When a story or task turns out to be measurably harder in one language, it’s adjusted — not left as a known measurement artifact.
- Dialectal representation. Content review includes speakers from multiple Spanish-speaking backgrounds, so a single regional word doesn’t disadvantage children from another community.
- Conceptual scoring where appropriate. When the question is “what does this child know?” rather than “what does this child know in English?”—giving credit for what a child demonstrates in either language. Few tools do this fully today; it’s the direction the research points, and the standard programs should hold vendors to.
Barrueco points to a field that holds both well-validated measures and much weaker ones. The barrier isn’t the absence of good tools. It’s that programs may be unaware of which tools were built for dual language learners and which were retrofitted less well with a translation.
Why bilingual assessment matters for program-level data
For administrators, the stakes go beyond any individual child. Assessment data drives funding conversations, program quality reviews, and resource decisions. When the underlying instrument underestimates bilingual children, entire programs can appear to be underperforming when they’re actually serving children well.
The pattern repeats at every level. Children look behind when they’re not. Classrooms look like they need intervention when their teachers are doing strong work. Programs serving large dual language learner populations carry a measurement penalty that has nothing to do with the quality of instruction.
Choosing a bilingual assessment that was designed for dual language learners, not adapted for them, is one of the most concrete steps program leaders can take to make their data trustworthy.
A bilingual PreK assessment, free to pilot
Khan Academy Kids PreK Assessments are bilingual (English/Spanish) by design, built simultaneously in both languages, calibrated for equivalent difficulty, and reviewed with dialectal variation in mind. Programs use them to generate consistent readiness data across math, literacy, receptive language, and executive function, without adding to teacher documentation load.
Thanks to grant funding, qualifying publicly funded PreK programs can pilot Khan Academy Kids for Schools, including the PreK Assessments, at no cost for the SY 26–27 and SY 27–28 school years.

