Supplementary materials for the 2026 SIGCSE-VIRTUAL paper "Poor Identifier Names in Introductory Programming: A Taxonomy and LLM Evaluation".
File dataset_ratings.csv contains the programs and the identifiers they contains, both human ratings and the LLM ratings from the Full prompt (the assigned score, reason and sub-reason from each), and the types of disagreement between the first and the second human rater and between the first rater and the LLM.
Folder prompts contains the texts of all the prompts described in the paper.
The submissions used in this study were anonymized prior to analysis and publication of the dataset to protect learners' privacy. The learners provided consent for their code to be used for research purposes, and all published materials do not present a risk of re-identification. The study involved analysis of anonymized educational data and formal ethics approval was deemed not required under the applicable guidelines.