--- title: "Getting Started with ecoGLMM" author: "André Felipe Carneiro dos Santos, Bruna Martins Bezerra, and Bárbara Lins Caldas de Moraes" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Getting Started with ecoGLMM} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r setup, include=FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>", eval = TRUE ) library(ecoGLMM) ``` ## Overview `ecoGLMM` provides a reproducible workflow for fitting and comparing ecological generalized linear mixed models. For each response and environmental predictor, the package can fit an additive model and a model containing an interaction with a temporal factor. Candidate models are compared using AICc, and the selected models can be summarized, diagnosed, plotted, and exported. This vignette presents a complete analysis with simulated data. The lightweight model example is evaluated when the vignette is built, whereas computationally intensive diagnostics and file exports are displayed without being run. Users can copy and run every example in an interactive R session. ## Install and load the package ```{r installation, eval=FALSE} # Install the CRAN version when available: install.packages("ecoGLMM") # Install the development version from GitHub: # install.packages("remotes") remotes::install_github("andre-fcsantos/ecoGLMM", upgrade = "never") library(ecoGLMM) ``` ## Required data structure The input must be a data frame with: - one or more numeric response columns; - numeric environmental predictors; - a temporal factor, such as sampling period; - a grouping variable for the random intercept, such as site. The final analysis data must contain all these columns in the same data frame. The generic template can import either one analysis-ready table or separate acoustic and environmental tables. In separate mode, it validates the join keys and combines the tables without changing the number of acoustic observations. Each row of the resulting data set represents one observation. The example below creates 150 observations from 10 sites and five sampling periods. ```{r simulate-data} set.seed(123) sites <- paste0("Site_", seq_len(10)) period_levels <- c("T-0", "T-1", "T-2", "T-3", "T-4") example_data <- expand.grid( site = sites, period = period_levels, replicate = seq_len(3), stringsAsFactors = FALSE ) n <- nrow(example_data) example_data$forest <- runif(n, 10, 95) example_data$urban <- runif(n, 0, 60) site_effect <- rnorm(length(sites), 0, 0.25) names(site_effect) <- sites linear_predictor <- -0.5 + 0.015 * example_data$forest - 0.012 * example_data$urban + site_effect[example_data$site] example_data$acoustic_index <- plogis( linear_predictor + rnorm(n, 0, 0.35) ) example_data$acoustic_index <- adjust_beta(example_data$acoustic_index) head(example_data) ``` ## Configure the responses The configuration object requires the columns `response` and `family`. An optional `label` column provides a readable name for tables and figures. Supported family names are `"beta"`, `"gamma"`, `"gaussian"`, `"poisson"`, and `"nbinom2"`. ```{r configuration} response_config <- data.frame( response = "acoustic_index", family = "beta", label = "Simulated acoustic index" ) ``` For ecoacoustic analyses, `prepare_acoustic_indices()` can transform NDSI from the interval [-1, 1] to (0, 1) and move boundary values away from zero and one before beta regression. ```{r acoustic-preparation, eval=FALSE} # analysis_data <- prepare_acoustic_indices( # data = raw_data, # ndsi = "NDSI", # beta_responses = c("AEI", "ACT"), # ndsi_output = "NDSI_beta" # ) ``` ## Run the candidate-model analysis Here, additive and period-interaction models are fitted separately for `forest` and `urban`. Predictors are standardized by default. All candidate models use the same complete-case data set, allowing valid AICc comparisons. ```{r run-analysis} fit <- run_ecoglmm( data = example_data, config = response_config, predictors = c("forest", "urban"), period = "period", group = "site", period_levels = period_levels, standardize = TRUE, complete_cases = TRUE, include_additive = TRUE, include_interactions = TRUE, competitive_delta = 2 ) fit summary(fit) ``` ## Inspect selection and coefficients ```{r inspect-results} fit$selection$acoustic_index fit$interaction_tests$acoustic_index fit$coefficients fit$preparation ``` A model with the smallest AICc is ranked first. Models with `Delta_AICc <= 2` are marked as competitive. Akaike weights express relative support within the candidate set. The likelihood-ratio tests compare the nested additive and interaction models for each predictor. The columns `R2_marginal` and `R2_conditional` are calculated with `MuMIn::r.squaredGLMM()`. Alternative estimates from `performance::r2_nakagawa()` are provided in columns ending in `_performance`. ## Run residual diagnostics ```{r diagnostics, eval=FALSE} diagnostics <- diagnose_models( fit, nsim = 1000, seed = 123 ) diagnostics$summary # Inspect the DHARMa residual plot: plot(diagnostics$residuals$acoustic_index) ``` The summary reports uniformity, dispersion, and outlier tests. A `"Passed"` status means that all three p-values exceed 0.05; it does not replace graphical inspection or ecological judgment. ## Create figures ```{r figures} plot_effect(fit, "acoustic_index") plot_selection(fit, "acoustic_index", metric = "delta") plot_selection(fit, "acoustic_index", metric = "weight") plot_coefficients(fit) ``` ## Export the analysis ```{r export, eval=FALSE} exported_files <- export_ecoglmm( object = fit, path = file.path(tempdir(), "ecoGLMM_results"), diagnostics = diagnostics, figures = TRUE ) exported_files ``` The output directory contains an Excel workbook with the overview, coefficients, model-selection results, interaction tests, and diagnostics. When `figures = TRUE`, publication-quality PNG files are also created. ## Use the complete generic analysis template The installed package includes a complete generic template that accepts either one analysis-ready table or separate acoustic and environmental tables in CSV or Excel format. When separate files are used, the template joins them by a user-defined site key and checks for duplicated or unmatched values. The example uses fictional column names and does not contain data or settings from any unpublished study. Locate, copy, and open an editable version of the template with: ```{r locate-template, eval=FALSE} template_path <- system.file( "examples", "ecoGLMM_template.R", package = "ecoGLMM" ) file.copy( from = template_path, to = file.path(tempdir(), "ecoGLMM_template.R"), overwrite = FALSE ) file.edit(file.path(tempdir(), "ecoGLMM_template.R")) ``` For a persistent copy, choose your own writable destination instead of `tempdir()`. Edit only Section 01 of the copied file before running the complete script. The full template is reproduced below directly from the file distributed with the package, ensuring that the vignette and script remain synchronized. ```{r display-generic-template, echo=FALSE, results="asis"} template_path <- system.file( "examples", "ecoGLMM_template.R", package = "ecoGLMM" ) if (!nzchar(template_path)) { stop("The generic analysis template was not found in the installed package.") } template_code <- readLines(template_path, warn = FALSE) cat( "```r\n", paste(template_code, collapse = "\n"), "\n```\n", sep = "" ) ``` Users should adapt file paths, column mappings, responses, distributions, predictors, temporal levels, output folders, and analysis options only in their own copy. After changing any setting, run the script again from Section 01. For two input tables, set `INPUT_MODE <- "separate"` and provide the exact join-column name used by each table. The template verifies uniqueness in the environmental table, reports unmatched keys, preserves the original acoustic-row order, removes configured non-analytical columns, and confirms the final column names. For ecoacoustic data requiring NDSI and beta-response preparation, set `PREPARE_ACOUSTIC_INDICES <- TRUE` in Section 01. The corresponding column names and boundary adjustment can be configured without editing the processing code. If the option is left disabled while the configured NDSI output is requested, the template returns a targeted instruction instead of a generic missing-column error. ## Reproducibility recommendations - Preserve the original data and analyse a separate working copy. - Record the response distribution selected for each ecological index. - Keep `complete_cases = TRUE` when comparing models by AICc. - Inspect convergence messages and DHARMa plots, not only p-values. - Report the candidate set, AICc, delta AICc, Akaike weights, coefficients, confidence intervals, and marginal and conditional R-squared. - Save `sessionInfo()` with the exported results.