Authors - Nwagu Chima Ajanwachuku, Onyemaobi Bethram Chibuzo, Nwafor Franca Amaka, Divine Nnodim Oluchi Abstract - Tertiary institutions across the world are now adopting artificial intelligence-based academic detection systems to aid in detecting different forms of academic misconduct. For these detection systems, there is more focus on technical performance metrics such as detection accuracy, precision and recall and little focus on whether these systems validly measure the complex construct of academic misconduct. This study aims to examine how AI-based academic integrity detection systems operationalise, measure, and validate academic misconduct in higher education, focusing on construct operationalisation, measurement accuracy, and construct validity, through a systematic literature review. We conducted a systematic review and retrieved articles from ACM Digital Library, IEEE Xplore, Web of Science, and Google Scholar. Of 793 articles, 56 were selected using the PRISMA framework, and the findings were synthesised narratively. The 56 studies focused on plagiarism detection, AI-generated text detection, authorship verification, behavioural monitoring, biometric authentication, and multimodal detection systems. Across all the studies we considered, detection systems mainly measured observable digital signals. 55 of 56 studies showed evidence of construct misalignment between the measured signal and the claimed misconduct construct. Most studies we considered treated similarity as plagiarism, AI-generated probability as dishonesty, and behavioural anomalies as cheating, even though these signals could not capture intent, differentiate between acceptable collaboration and collusion, or account for disclosure practices or alignment with institutional policy. Also, we observed that most studies validated detection systems using technical metrics such as accuracy, precision, recall, and F1 score, and just a few directly addressed construct validity, bias, or robustness.