--- title: "Construct-Rating Content Validation: Hinkin-Tracey to Colquitt" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Construct-Rating Content Validation: Hinkin-Tracey to Colquitt} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") set.seed(12) ``` ```{r setup} library(contentvalidR) ``` ## What the rating workflow asks The Hinkin and Tracey (1999) content-rating procedure asks judges to evaluate how well each item corresponds to **each** construct definition under consideration. The typical design is fully crossed within judges: the same judge rates an item against the intended definition and against one or more orbiting definitions. That design provides two complementary kinds of evidence: - **Definitional correspondence:** does the item strongly match its intended construct? - **Definitional distinctiveness:** does it match the intended construct more strongly than plausible orbiting constructs? `contentvalidR` keeps those questions separate rather than reducing the study to a single coefficient. ## Example data ```{r data} rating_dat <- expand.grid( item = c("A1", "A2", "A3", "B1"), rater = 1:24, construct = c("A", "B", "C") ) rating_dat$target_construct <- ifelse(rating_dat$item == "B1", "B", "A") rating_dat$rating <- ifelse( rating_dat$construct == rating_dat$target_construct, pmin(5, pmax(1, round(rnorm(nrow(rating_dat), 4.4, .6)))), pmin(5, pmax(1, round(rnorm(nrow(rating_dat), 2.2, .8)))) ) ``` Each item-judge combination appears once for every construct definition. A duplicated item-rater-construct row is treated as a data error. ## HTC: definitional correspondence Following Colquitt et al. (2019), the Hinkin-Tracey correspondence index is \[ HTC = \frac{\bar{x}_{target}}{a}, \] where \(a\) is the number of response anchors when ratings use a 1-to-\(a\) scale. `contentvalidR` can also accept an equally spaced integer scale such as 0-to-4; it shifts that scale internally to the equivalent 1-to-5 anchor metric before computing HTC. ```{r htc} htc(rating_dat, scale_min = 1, scale_max = 5) ``` Higher HTC means stronger correspondence with the intended definition. ## HTD: definitional distinctiveness HTD compares intended-definition ratings with orbiting-definition ratings: \[ HTD = \frac{\text{average}(x_{target} - x_{orbiting})}{a - 1}. \] It ranges from -1 to 1. Positive values favor the intended definition; negative values indicate that orbiting definitions are rated more highly on average. ```{r htd} htd(rating_dat, scale_min = 1, scale_max = 5) ``` The item-level table also identifies the strongest orbiting competitor. That is often more useful for revision than merely knowing that distinctiveness is weak. ## Repeated-measures item screening The same judges provide multiple construct ratings, so those observations are not independent. `anova_content()` uses a one-way repeated-measures ANOVA for the standard fully crossed design and follows it with planned paired comparisons of the intended definition against every orbiting definition. ```{r anova} aov_out <- anova_content(rating_dat, design = "within") aov_out attr(aov_out, "contrasts") ``` The omnibus F test asks whether the item's mean ratings differ somewhere across definitions. The planned contrasts ask the more direct content-validity question: is the target mean higher than each orbiting mean? With more than two construct definitions, the conventional repeated-measures F test assumes sphericity. `contentvalidR` reports that historical omnibus test but does not hide the assumption. The planned target-versus-orbiting comparisons are therefore important diagnostic evidence rather than decorative post-hoc tests. ## Recommended workflow ```{r workflow} fit <- rating_validity( rating_dat, scale_min = 1, scale_max = 5 ) fit summary(fit) ``` The item-level recommendation has deliberately limited meaning: - **Retain:** the item cleared the package's inferential screening rule in this pretest. - **Review:** the full screening rule was not met; inspect wording, construct overlap, and judge feedback. - **Insufficient data:** too few complete judge profiles are available for the repeated-measures comparison. `Review` is not an instruction to delete an item. Content coverage can be harmed by mechanical item deletion. ## Scale-level Colquitt norms Colquitt et al. (2019) created empirical norms from **scale-level averages** of HTC and HTD across 112 published scales. `rating_validity()` therefore averages item HTC/HTD within each target scale before assigning those descriptive normative labels. ```{r scale} fit$scale_summary colquitt_benchmarks("htc") colquitt_benchmarks("htd") ``` The overall bands are empirical percentile standing, not universal validity cutoffs. If the average correlation between a focal scale and its orbiting scales is known, correlation-conditional norms can be requested: ```{r conditional} rating_validity( rating_dat, orbiting_r = c(A = .42, B = .55) )$scale_summary ``` A given level of distinctiveness can be more impressive when the focal and orbiting constructs are known to correlate strongly. ## Naive versus expert judges Colquitt et al.'s normative distributions were developed using naive judges representative of substantive target populations. Their paper cautions against applying those norms to expert panels. The package therefore separates calculation from norm applicability: ```{r expert} rating_validity(rating_dat, judge_type = "expert")$scale_summary ``` HTC/HTD are still computed, but the Colquitt labels are suppressed. ## Missing ratings For HTD and repeated-measures inference, a judge must have a usable rating for every construct definition presented for that item. Incomplete profiles are excluded itemwise and counted explicitly in the output. This preserves the paired design rather than quietly treating incomplete repeated observations as independent data. ## Plotting The original one-index views remain available: ```{r plot, fig.width=7, fig.height=4} plot(fit, metric = "htc") plot(fit, metric = "htd") ``` A correspondence-distinctiveness evidence map displays HTC and HTD together: ```{r rating-map, fig.width=7, fig.height=5} plot(fit, type = "map") ``` Target-scale averages are shown as diamonds and items needing review are labeled by default. As with the item-sort map, Colquitt norm regions are not drawn across individual items because those benchmarks were constructed from scale averages. The **target-versus-competitor gap plot** makes the Hinkin-Tracey mean-rating logic more directly visible: ```{r rating-profile, fig.width=7, fig.height=5} plot(fit, type = "profile") ``` Filled points are intended-definition means, open points are the strongest orbiting-definition means, and the connecting segment is the observed content distinctiveness gap. A reversed segment immediately identifies an item whose strongest competitor outrates its intended definition. This is a graphical extension of the mean-rating tables used in the original procedure, not a new statistical cutoff. ## Reporting A useful report should identify: 1. the construct definitions and orbiting constructs shown to judges; 2. the judge population and recruitment method; 3. the response anchors and rating instructions; 4. item-level HTC, HTD, repeated-measures omnibus results, and planned contrasts; 5. the strongest orbiting competitor for items needing review; 6. target-scale mean HTC/HTD and the norm set used, if applicable; and 7. qualitative feedback and substantive decisions made after the pretest. The quantitative analysis is evidence about definitional correspondence and distinctiveness. It does not by itself demonstrate that the item pool comprehensively samples the full construct domain. ## References Hinkin, T. R., & Tracey, J. B. (1999). An analysis of variance approach to content validation. *Organizational Research Methods, 2*(2), 175-186. https://doi.org/10.1177/109442819922004 Colquitt, J. A., Sabey, T. B., Rodell, J. B., & Hill, E. T. (2019). Content validation guidelines: Evaluation criteria for definitional correspondence and definitional distinctiveness. *Journal of Applied Psychology, 104*(10), 1243-1265. https://doi.org/10.1037/apl0000406