Skip to main content

Fair and Inclusive Language Assessment: Practical Strategies for Educators

by Shahid Abrar-ul-Hassan, Dan Douglas |

In language education programs, learners are assessed multiple times to determine how well they have learned. Their learning is typically expressed in a score or letter grade. These assessment results may become part of the permanent record of learners’ educational journeys and can influence progression, placement, graduation, and future opportunities.

Targeting “No Harm” in Assessment

Nearly two decades ago, Taylor and Nolen highlighted a fundamental principle for educators as ‘‘the ethical responsibility of educators is ‘first, do no harm’’’ in assessment.

Such harm can take many forms. For instance, a high-stakes test may create an excessive level of anxiety (commonly known as test anxiety). An assessment may lead to negative washback effect (i.e., an adverse impact of assessment on teaching and learning) by encouraging students to memorize information rather than develop meaningful language abilities. Moreover, a low level of score accuracy could lead to biases and misinterpretation of scores.

A single assessment format may consistently advantage some learners while limiting others’ opportunities to demonstrate what they know and can do. Moreover, unclear criteria may leave students uncertain about what competencies are involved. Even an otherwise well-designed assessment may create exclusivity based on learners’ access to the resources, technologies, time, or conditions needed to complete it successfully.

Educators need to make every reasonable effort to understand and minimize potential harm that may arise during the assessment process.

Recognizing Fairness and Inclusion

Before we begin, let’s clarify what fairness and inclusion do not mean. Fair assessment does not entail giving every learner a high score, lowering standards, or an easy and challenge-free assessment. Nor does inclusion mean that every learner must necessarily be given a completely different assessment.

Fairness means creating reasonable opportunities for learners to showcase the knowledge, skills, and abilities that the assessment is intended to measure (i.e., test validity).

Inclusion means recognizing that learners may have different educational, cultural, psycho-physical, and technological circumstances, which can impact their responses to assessment.

For instance, in a speaking assessment designed to measure a learner’s ability to communicate ideas orally, if the assessment requires students to interpret culturally unfamiliar humour, their performance will likely be influenced by something other than the language ability educators intend to assess. Similarly, if a writing task contains complicated or incomplete instructions, which might be unintended, performance may partly reflect students’ ability to decipher the task rather than their ability to demonstrate the targeted writing skills. These examples highlight fairness being impacted by low validity (i.e., measuring what it is intended to assess or something irrelevant getting in the way).

Tackling Subjectivity

One of the most important challenges in assessment is subjectivity. Language educators commonly assess performances for which there is no single objectively correct answer: essays, paragraphs, presentations, interviews, portfolios, multistep projects, discussions/debates, and other complex performances.

Nevertheless, the resulting grades in these assessments may appear to be remarkably precise when they are, in fact, to some degree unreliable. For example, an 8/10 score is interpreted as better performance than a 6/10, and a B+ is distinguished from a regular B. Once scores are entered into an institutional gradebook, these judgments, though subjective, are generally regarded as accurate representations of student achievement.

The challenge is not that educators rely on their professional judgments, which is an essential part of assessment practices. The challenge is making that judgment as evidence-informed, transparent, and consistent as possible. We offer the following four-step baseline fairness check in assessment:

1. Clarifying the criteria

Sharing with learners what evidence would distinguish different performance levels, such as excellent, competent, developing, and inadequate. If this level differentiation is not clear to educators, learners are unlikely to understand it either.

↓

2. Using a scoring guide/rubric

A scoring guide/rubric should not merely divide 10 marks, for example, among several categories but should refer to specific performance indicators that actually matter for the course learning outcomes.

↓

3. Calibrating judgments

After giving a test but before grading the entire class, grade a small sample of responses (together if a teaching team). Judge how well the test criteria match the responses and perhaps adjust the test rubric, establishing reasonable criteria that match the actual class level. Where possible, discuss any sample responses with a peer educator who teaches the same or a similar course.

↓

4. Rechecking the grading consistency

Reviewing or reflecting whether standards shifted after reading too many assignments or the same type of errors are graded differently. Revisit one or two earlier responses to check your consistency.

Diversifying the Evidence, Not Standards

Another practical approach toward fairness is to avoid making consequential judgments about learning from a single source of evidence. As we have discussed in previous posts on innovative alternatives in EAL classroom assessment, learners can demonstrate their academic achievement through multiple forms: tests, presentations, projects, portfolios, interviews, written work, demonstrations, and other meaningful tasks.

This diversification does not mean offering unlimited choices for every assessment. Rather, educators can diversify assessment while making them relevant to curricular objectives and institutional policies. For example, instead of basing 70% of a language course grade on a midterm and final examination, an educator might collect evidence through a reading-to-writing task, an oral presentation, a portfolio, a collaborative project, and selected quizzes or tests. A basic challenge associated with a single format is that one format is less likely to respond to diverse learning styles.

When learning outcomes allow, educators can give learners some choice in how they demonstrate their skills. For example, learners completing an oral communication task might choose from several comparable topics, while those completing a project might choose the topic or context they focus on. The key is to make sure all options address the same learning outcomes and are evaluated using equivalent criteria.

Making the Assessment Transparent

Fairness can also be strengthened even before learners begin an assessment. This can be done, for example, by

    • providing clear instructions,
    • explaining the purpose of the assessment and its relationship to the learning outcomes, and
    • sharing assessment criteria in advance.

It’s important to provide this information early on and to give students opportunities to ask questions.

One useful strategy is to spend 5 minutes in class asking learners to review the rubric and explain, in their own words, what successful performance would look like. Their responses may reveal expectations that aren’t as clear as the educator intended. This simple exercise helps make assessment criteria transparent—and transparency doesn’t compromise rigour.

Conducting Postassessment Reviews

Fairness and inclusion should also be considered after scores have been assigned and released. Take a few minutes to look at the results with these questions:

    1. Did an unexpectedly large proportion of the class perform poorly on one question or criterion?
    2. Did more than one student misunderstand the same instruction?
    3. Was one part substantially more difficult than anticipated?
    4. Did an assessment format appear to create barriers unrelated to the intended learning outcome?

Reflections related to these aspects can provide useful evidence for assessment revisions and refinement.

Learner feedback can help as well. Instead of simply asking, “Was the test fair?” ask more specific questions, either informally or through a brief survey:

    1. Were any instructions unclear?
    2. Which task best allowed you to show what you learned?
    3. Was there anything that prevented you from presenting what you knew?

The answers can inform the improvement of the next assessment cycle and might even require adjusting the results of the current test.

Initiating Educators’ Agentic Role

Institutional policies, standardized course requirements, grading systems, accessibility procedures, and program expectations all impact assessment practices. Educators can nevertheless play an agentic role for fairness and inclusivity. For instance, they can

    • advocate for assessment policies that support learning,
    • raise concerns about unnecessary barriers,
    • collaborate with accessibility services,
    • diversify evidence where course requirements permit, and
    • share effective practices with their colleagues as a resource.

Notably, fairness and inclusion are equal in significance to accuracy, consistency, or institutional standards. The capacity to make these distinctions is part of what we have emphasized throughout this series as language assessment literacy.

Assessment-literate educators do more than administer tests and assign grades. They understand what assessment results mean, recognize their limitations, make informed decisions about assessment design and scoring, consider consequences for learners, and continually reflect on their practices to develop appropriate opportunities for students to demonstrate learning.

Following Up With Fairness and Inclusivity

Drawing on research in culturally responsive pedagogy, Bennett identifies five design principles that can help educators make assessment practices fairer and more inclusive:

Principle 1

Present problem situations that connect to, and value, examinee experience, culture, and identity

Principle 2

Allow for multiple forms of representation and expression in problem stimuli and in responses

Principle 3

Promote instruction for deeper learning through assessment design

 

Principle 4

Adapt the assessment to student characteristics

 

Principle 5

Represent assessment results as an interaction among what the examinee brings to the assessment, the types of tasks engaged, and the conditions and context of that engagement

As you plan your next assessment, consider three questions:

    1. What exactly am I trying to assess?
    2. Does every learner have a reasonable opportunity to demonstrate it?
    3. Can I defend the consistency and accuracy of the judgment I make?

Fair and inclusive assessment doesn’t come from a single technique or checklist. It requires professional judgment, reflection, and a willingness to examine our own practices. As mentioned earlier, educators’ language assessment literacy is an important part of this process—and of our responsibility to the learners we assess.

Useful Resource: The International Language Testing Association (ILTA) Code of Ethics, which includes standards for fairness and inclusiveness in both professional and classroom-based language assessment, is a helpful reference for assessment practices.

 

About the author

Shahid Abrar-ul-Hassan

Shahid Abrar-ul-Hassan, PhD, has been working for over two decades across the globe as an English language educator (currently, associate professor), academic researcher, and faculty development professional. He is an alumnus of the Middlebury Institute of International Studies at Monterey (USA) and the University of British Columbia (Canada). His edited works include special issues on ESP assessment in the English for Specific Purposes Journal (Elsevier, 2025-) and language assessment literacy in System (Elsevier, 2023) as well as Volume 1 of the TESOL Encyclopedia of English Language Teaching (Wiley-Blackwell, 2018). His professional interests are EAP/ESP, (learner-oriented) language assessment, assessment literacy, learner motivation, and differentiated teacher development.

About the author

Dan Douglas

Dan Douglas, PhD, has worked as a university professor and a language testing professional. He received the Distinguished Achievement Award (2019) from Cambridge University Assessment and the International Language Testing Association, for distinguished service and scholarship in the field of language testing. His books include Assessing Languages for Specific Purposes (Cambridge, 2000), Assessing Language through Computer Technology (with C. Chapelle, Cambridge, 2006), Understanding Language Testing (Routledge, 2010), and Fundamental Considerations in Technology Mediated Language Assessment (with K. Sadeghi, Routledge, 2023). His current interests include assessing English for aviation and English for nursing.

comments powered by Disqus

This website uses cookies. A cookie is a small piece of code that gives your computer a unique identity, but it does not contain any information that allows us to identify you personally. For more information on how TESOL International Association uses cookies, please read our privacy policy. Most browsers automatically accept cookies, but if you prefer, you can opt out by changing your browser settings.