--- title: "Discrete, mixed, and zero-inflated data" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Discrete, mixed, and zero-inflated data} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set(collapse = TRUE, comment = "#>") library(rvinecopulib) set.seed(401) ``` Copula likelihoods for variables with atoms require two probabilities at every observation: the CDF value `F(x)` and its left limit `F(x-)`. For an integer-valued variable, `F(x-) = F(x - 1)`. Continuous variables satisfy `F(x-) = F(x)`. The high-level `vine()` interface computes these values from fitted margins. The copula-only `bicop()` and `vinecop()` interfaces accept them explicitly. ## Discrete variables on the original scale All ordinary discrete variables are assumed to have integer support. Declare a numeric or integer column through `var_types = "d"`. ```{r integer-data} n <- 100 latent <- rnorm(n) x <- data.frame( value = latent + rnorm(n), count = rpois(n, exp(0.5 + 0.3 * latent)) ) fit <- vine( x, var_types = c("c", "d"), copula_controls = list(family_set = "onepar") ) summary(fit)$margins dvine(x[1:5, ], fit) ``` Explicitly discrete columns are checked for finite integer-valued observations. This catches accidental declarations of rounded-looking continuous data. ## Ordered factors Ordered factors are discrete variables whose levels are mapped internally to the integer support `0, 1, ...`. Their class and levels are restored after simulation. ```{r ordered-data} survey <- data.frame( rating = ordered( sample(c("low", "middle", "high"), n, replace = TRUE), levels = c("low", "middle", "high") ), score = rnorm(n) ) ordered_fit <- vine( survey, copula_controls = list(family_set = "indep") ) ordered_draws <- rvine(5, ordered_fit) str(ordered_draws) ``` No `var_types` entry is needed because `ordered()` carries the declaration in the column class. ## Zero-inflated variables `zero_inflated()` marks a continuous variable with an atom at zero. It behaves like a numeric vector in a data frame and survives ordinary subsetting. ```{r zero-inflated} amount <- zero_inflated(c(rep(0, 25), rexp(75))) zi_data <- data.frame(amount = amount, score = rnorm(100)) inherits(zi_data$amount, "zero_inflated") zi_fit <- vine( zi_data, copula_controls = list(family_set = "indep") ) summary(zi_fit)$margins[, c("name", "type")] ``` The atom is always located at zero. A custom zero-inflated margin must return the atom probability from `dmargin(0, margin)`. rvinecopulib then computes the left limit as ``` pmargin(0, margin) - dmargin(0, margin) ``` and uses the ordinary CDF away from zero. Alternatively, declare the type with `var_types = "zi"`. ## Copula-scale data layouts Let `d` be the dimension and `k` the number of discrete variables. Copula methods accept two layouts: - **expanded:** `2d` columns, with all `F(x)` columns followed by all `F(x-)` columns; - **compact:** `d + k` columns, with the `F(x)` block followed only by left-limit columns for discrete variables, in variable order. Consider a three-dimensional model with discrete variables 1 and 3. ```{r layouts} n <- 80 raw <- data.frame( first = rpois(n, 2), middle = rnorm(n), last = rpois(n, 4) ) values <- cbind( ppois(raw$first, 2), pnorm(raw$middle), ppois(raw$last, 4) ) left_discrete <- cbind( ppois(raw$first - 1, 2), ppois(raw$last - 1, 4) ) compact <- cbind(values, left_discrete) left_all <- cbind( ppois(raw$first - 1, 2), pnorm(raw$middle), ppois(raw$last - 1, 4) ) expanded <- cbind(values, left_all) copula_fit <- vinecop( compact, var_types = c("d", "c", "d"), family_set = "indep" ) stopifnot(isTRUE(all.equal( dvinecop(compact, copula_fit), dvinecop(expanded, copula_fit) ))) ``` For a mixed bivariate model, this means three columns: the two CDF values and the left limit of the discrete variable. Fully discrete bivariate models need four columns. ## What the density means `dbicop()` and `dvinecop()` return the copula density contribution defined as the joint density or mass divided by the marginal densities or masses. This definition is consistent across continuous, discrete, and mixed models. For a bivariate model, write \[ u_j = F_j(x_j), \qquad u_j^- = F_j(x_j^-), \qquad \Delta_j = u_j - u_j^-. \] If both variables are discrete, the copula contribution is the rectangle probability divided by the two marginal masses: \[ c^*(u_1,u_2) = \frac{C(u_1,u_2)-C(u_1^-,u_2)-C(u_1,u_2^-)+C(u_1^-,u_2^-)} {\Delta_1\Delta_2}. \] If only the first variable is discrete, the corresponding mixed contribution is \[ c^*(u_1,u_2) = \frac{\partial_2 C(u_1,u_2)-\partial_2 C(u_1^-,u_2)}{\Delta_1}. \] The continuous density `c(u1, u2)` is recovered when both jump widths vanish. Vine likelihoods apply these continuous, mixed, or rectangle contributions edge by edge to the appropriate conditional CDF values. `dvine()` multiplies that contribution by the fitted marginal densities or masses and therefore returns the full joint density or probability mass on the original scale. In general, \[ f_X(x) = c^*(u,u^-)\prod_{j=1}^d f_j(x_j), \] where each `f_j` denotes a density or a probability mass as appropriate. ## Simulation `rvinecop()` always returns one uniform-scale value per variable. It does not invent a marginal distribution for a discrete copula variable. Use `rvine()` when simulated values should be returned on the original scale with their integer, ordered, or zero-inflated margins. ```{r simulation} copula_draws <- rvinecop(4, copula_fit) dim(copula_draws) original_draws <- rvine(4, ordered_fit) is.ordered(original_draws$rating) ``` ## Rosenblatt transforms with atoms For an atom, the Rosenblatt transform is an interval rather than a single value. By default, `rosenblatt()` randomizes uniformly within that interval so the transformed variable is uniform under a correct model. Set `randomize_discrete = FALSE` to return the deterministic upper endpoint. More precisely, if `G_j(ยท | x_