Date of Award

10-13-2025

Degree Type

Thesis

Degree Name

MSc in Nursing

First Advisor

Dr Laila Ladak

Second Advisor

Dr. Philip Moons

Third Advisor

Ms. Arjumand Rizvi

Department

School of Nursing and Midwifery, Pakistan

Abstract

Background: Pakistan continues to face a shortage of qualified faculty, which increases pressure on teaching and the processes involved in creating and maintaining high-quality assessments. Developing high-quality MCQs requires subject expertise, careful design of distractors, and a significant amount of time. However, when experienced faculty are limited and workloads are high, the assessments produced may not meet the desired standards. Artificial intelligence can be a valuable resource for creating assessments, especially when its use is guided by experts and supported by robust quality assurance procedures.
Purpose: The primary objective of this study was to compare the quality and cognitive-level classification of MCQs created by three sources: (1) AI-Objective-Driven prompts, (2) AIContent-Driven prompts, and (3) faculty-developed MCQs.
Method: This quantitative, quasi-experimental study used a three-arm comparative design to evaluate the quality and cognitive-level classification of 54 MCQs, with 18 in each arm. Each MCQ was systematically aligned with a Table of Specifications corresponding to Bloom’s taxonomy levels (C1–C3). Arm 1 consisted of AI-Objective-Driven MCQs created using the Prompt Canvas framework based on specified learning objectives. Arm 2 included AIContent-Driven MCQs, which were developed from the course textbook material using the Prompt Canvas framework. Arm 3 included faculty-developed MCQs selected from the institutional examination bank. Three pairs of reviewers, each consist of one education expert, and one subject matter expert evaluated the MCQs. The reviewers used a five-domain quality rubric that evaluated appropriateness, clarity, relevance, quality of alternatives, and suitability for assessment. Data were analysed using one-way ANOVA to compare mean quality scores across the three arms. Inter-rater reliability for quality ratings was assessed using the Intraclass Correlation Coefficient (ICC), while agreement on Bloom’s level classification was measured with Fleiss’ κ.
Findings: A one-way ANOVA showed that the difference in overall mean quality scores across the three arms was not statistically significant (p = .88). Inter-rater reliability for quality ratings was moderate (ICC ≈ 0.65), indicating acceptable agreement among reviewers. Agreement on Bloom’s level classification was poor to fair (Fleiss’ κ ≈ 0.20), reflecting variability in cognitive-level judgments.
Conclusion: The study demonstrates that when AI is guided by structured prompts and validated by experts, it can generate MCQs of quality similar to those created by experienced faculty. Because of the limited reliability in classifying higher cognitive levels, creating higher-order MCQs remains a challenge for AI, emphasizing the need for ongoing human involvement in the assessment process. A collaborative human-in-the-loop approach that combines AI efficiency with educators' contextual judgment offers a promising way to improve assessment quality in nursing education. Institutions should prioritize developing faculty skills in AI literacy and prompt engineering to ensure the responsible, effective, and ethical use of AI tools in academic assessment practices.

First Page

1

Last Page

157

Share

COinS