--- title: "Design & Reporting Guide" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Design & Reporting Guide} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") library(contentvalidR) read_example <- function(name) { utils::read.csv( system.file("extdata", name, package = "contentvalidR"), stringsAsFactors = FALSE ) } ``` ## Choosing a method Use the design that matches the question posed to judges rather than selecting an index after data collection. - **Item sort:** judges assign each item to the construct definition that best represents it. Use `sort_validity()` for Psa/Csv, exact target-count screening, strongest-competitor diagnostics, and scale-level Colquitt norms. - **Construct rating:** the same judges rate each item against every focal and orbiting definition. Use `rating_validity()` for HTC/HTD, repeated-measures inference, planned target-versus-orbiting contrasts, and scale-level norms. - **Expert panel:** experts answer a relevance, essentiality, or congruence question. Use `expert_validity()` with the corresponding explicit `mode`. The three designs can complement one another during scale development, but their statistics are not interchangeable. ## Sample-size planning For item sorts, plan judge N in relation to the exact retention rule and a plausible true target-assignment probability. `sort_power()` provides exact planning probabilities; avoid a universal judge-count rule of thumb. For construct-rating studies, power depends on the number of judges, number of construct definitions, within-judge target-orbiting separation, and missing profiles. Report the effective complete-judge N itemwise. For expert panels, panel-size sensitivity is part of the statistic. CVR exact critical counts and common CVI review guidelines therefore need to be interpreted with the actual effective N, not a nominal panel size that ignores missingness. ## Recommended reporting sequence 1. **Design:** judge population, recruitment, construct definitions, item pool, instructions, and response format. 2. **A priori rules:** alpha, relevance threshold, target mapping, multiplicity adjustment if used, and any planned norms. 3. **Item-level evidence:** correspondence, distinctiveness, exact/paired inference, strongest competitor, and review status as appropriate. 4. **Scale-level evidence:** target-scale averages or S-CVI summaries where the method defines them. 5. **Substantive decisions:** revisions, removals, retained domain-coverage items, and how qualitative comments informed those decisions. 6. **Reproducibility:** package version, analysis settings, anonymized data when permitted, and the script used to reproduce tables/figures. The `reporting-examples` vignette expands this sequence into reusable methods and results scaffolds. Treat those examples as reporting patterns rather than fixed language that must be copied verbatim. ## Deterministic bundled examples ```{r example-data} sort_dat <- read_example("sort_example.csv") rating_dat <- read_example("rating_example.csv") expert_rel <- read_example("expert_relevance_example.csv") sort_fit <- sort_validity(sort_dat) rating_fit <- rating_validity(rating_dat, scale_min = 1, scale_max = 5) expert_fit <- expert_validity( as.matrix(expert_rel[setdiff(names(expert_rel), "expert")]), mode = "relevance", lo = 1, hi = 4 ) ``` These files are synthetic and generated by `data-raw/build-example-data.R` in the source repository. They deliberately include both supported and review-worthy items so documentation exercises realistic output paths without depending on random-number generation. ## Table templates ### Item sort ```{r sort-table} sort_fit$results[c( "item", "target", "n", "n_target", "competitor", "psa", "csv", "p_value", "status", "recommendation" )] ``` At the target-scale level, report mean Psa/Csv and the benchmark set actually used. Do not convert Colquitt's scale-level norms into individual-item cutoffs. ### Construct rating ```{r rating-table} rating_fit$results[c( "item", "target", "n_complete", "strongest_competitor", "htc", "htd", "p_value", "max_contrast_p", "status", "recommendation" )] rating_fit$scale_summary ``` Report the repeated-measures design and target-versus-orbiting contrasts. For a review item, naming the strongest competitor is often more informative than a standalone p value. ### Expert relevance ```{r expert-table} expert_fit$results[c( "item", "N", "V", "ci_low", "ci_high", "I_CVI", "kappa_mod", "status", "recommendation" )] expert_fit$scale_summary ``` Essentiality and congruence require different expert tasks. Do not place CVR, CVI, Aiken V, and IOC in one generic threshold table as if they answer the same question. ## Diagnostics and later validation `signal_detection()` and `reproducibility_phi()` remain auxiliary helpers. They can be useful when researchers later compare pretest decisions with CFA/IRT retention or an independent replication pretest, but they are not required parts of the flagship workflows. ```{r diagnostics} pretest_supported <- sort_fit$results$status == "Supported" later_retained <- c(TRUE, TRUE, TRUE, FALSE, TRUE, FALSE) signal_detection(pretest_supported, later_retained) replication_supported <- c(TRUE, TRUE, TRUE, FALSE, TRUE, TRUE) reproducibility_phi(pretest_supported, replication_supported) ``` ## Power quick check ```{r power} sort_power(N = c(20, 30, 40), true_p = c(.60, .70, .80)) ``` ## Good practices Pre-register quantitative screening criteria when feasible, preserve qualitative judge feedback, report effective N after missingness, and archive the construct definitions and item wording used in the pretest. A content-validation statistic is evidence from a designed judgment task; it is not a substitute for defining and sampling the construct domain.