--- title: "Getting Started with rumenGP" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Getting Started with rumenGP} %\VignetteEngine{knitr::rmarkdown} \usepackage[utf8]{inputenc} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>" ) ``` # Introduction **rumenGP** provides a complete workflow for analyzing *in vitro* rumen gas production experiments. The package supports: - ANKOM RF datasets - Manual gas-volume datasets - Pressure-based datasets - Fifteen built-in kinetic models - User-defined kinetic models - Model comparison and ranking - Treatment-level model evaluation - Diagnostic and visualization tools This vignette demonstrates a complete workflow using the packaged ANKOM example dataset. ```{r} library(rumenGP) ``` # Load Example Data The package includes a small example dataset. ```{r} files <- example_data() files ``` # Import ANKOM Data Import the ANKOM RF output file and metadata table. ```{r} raw_data <- read_ankom( files$ankom ) metadata <- read_metadata( files$metadata ) ``` # Validate Metadata Before processing data, validate the metadata table. ```{r} metadata <- validate_metadata( metadata ) ``` # Process ANKOM Data Convert pressure measurements into cumulative gas production. ```{r} gp <- process_ankom( raw_data, metadata, headspace_ml = 210, temperature_c = 39, zero_negative_pressure = TRUE ) ``` # Validate Processed Data The resulting dataset is a standardized `rumen_gp` object. ```{r} gp <- validate_ankom( gp ) class(gp) ``` Inspect the data: ```{r} head(gp) ``` # Visualize Raw Gas Production Individual bottle profiles can be visualized. ```{r, eval = FALSE} plot_gp( gp, head = "1" ) ``` # Fit Kinetic Models Several built-in models are available. ```{r} groot_fit <- fit_groot(gp) mm_fit <- fit_mm(gp) gompertz_fit <- fit_gompertz(gp) brody_fit <- fit_brody(gp) burr_fit <- fit_burr_xii(gp) inverse_paralogistic_fit <- fit_inverse_paralogistic(gp) ``` # Summarize Model Fits Each model provides parameter estimates and diagnostic statistics. ```{r} summary(groot_fit) ``` # Identify Potentially Problematic Bottles ```{r} flags <- flag_model( groot_fit ) head(flags) ``` # Plot Model Fits Observed and predicted values can be visualized. ```{r, eval = FALSE} plot_fit( groot_fit, head = "1" ) ``` # Plot Residuals Residual plots help identify systematic deviations. ```{r, eval = FALSE} plot_residuals( groot_fit, head = "1" ) ``` # Compare Models Compare model performance using multiple metrics. ```{r} comparison <- compare_models( Groot = groot_fit, MichaelisMenten = mm_fit, BurrXII = burr_fit, InverseParalogistic = inverse_paralogistic_fit, Gompertz = gompertz_fit, Brody = brody_fit ) comparison ``` The comparison table includes: - Mean R-squared - Mean RMSE - Mean RSS - Mean AIC - Mean BIC - Number of successful fits # Rank Models ```{r} rank_models( comparison ) ``` # Compare Models by Treatment Treatment-level comparisons are also available. ```{r} treatment_comparison <- compare_models_by_treatment( Groot = groot_fit, MichaelisMenten = mm_fit, BurrXII = burr_fit, InverseParalogistic = inverse_paralogistic_fit, Gompertz = gompertz_fit, Brody = brody_fit ) treatment_comparison ``` # Rank Models by Treatment ```{r} ranked_treatments <- rank_models_by_treatment( treatment_comparison ) ranked_treatments ``` # Determine the Best Model per Treatment ```{r} best_models <- best_model_by_treatment( ranked_treatments ) best_models ``` # Model Win Frequency ```{r} model_win_frequency( best_models ) ``` # Quality Control Workflow A typical workflow is: ```text Import data ↓ Validate metadata ↓ Process ANKOM data ↓ Validate processed data ↓ Fit multiple candidate models ↓ Flag problematic bottles ↓ Inspect residuals ↓ Exclude problematic bottles ↓ Refit models ↓ Compare RMSE ↓ Compare AIC and BIC ↓ Select final model ``` Example bottle exclusion: ```{r, eval = FALSE} gp_clean <- exclude_heads( gp, heads = c("10"), reason = "Sensor malfunction" ) ``` # Available Models Current built-in models: - Brody - Dual Logistic - EXP0 - EXPL - Gompertz - Groot - LE0 - LEL - Logistic - Mitscherlich - Generalized Michaelis-Menten - Ørskov and McDonald - Burr XII - Inverse Paralogistic # Model Equivalence ## Groot and Generalized Michaelis-Menten The Groot and generalized Michaelis-Menten models are mathematically equivalent. Parameter correspondence: - VF = A - b = K - k = c Both formulations produce identical fitted values, residuals, diagnostics, AIC, BIC, RMSE, and R-squared when convergence is achieved. Researchers may choose either formulation depending on the terminology commonly used in their field. ## Groot, Generalized Michaelis-Menten, and Log-logistic The Log-logistic formulation \[ V(t) = VF \frac{(rt)^a} { 1 + (rt)^a } \] can be rewritten as \[ V(t) = VF \frac{t^a} { t^a + (1/r)^a } \] which is mathematically identical to both the Groot and generalized Michaelis-Menten models. Parameter correspondence: | Groot | Generalized Michaelis-Menten | Log-logistic | |---------|---------|---------| | VF | A | VF | | b | K | 1/r | | k | c | a | Therefore: ```text Groot = Generalized Michaelis-Menten = Log-logistic ``` These formulations describe the same underlying curve and differ only in parameterization. For this reason, rumenGP does not currently implement a separate Log-logistic fitting routine. The Log-logistic curve family is already represented through the existing Groot and generalized Michaelis-Menten implementations. # Recommended Model Selection Workflow No single gas-production model should be considered universally superior. A recommended workflow is: 1. Fit several biologically plausible models. 2. Verify convergence. 3. Inspect fitted curves. 4. Examine residuals. 5. Compare RMSE. 6. Compare AIC and BIC. 7. Evaluate biological plausibility of parameter estimates. 8. Select the model most appropriate for the scientific objective. Recent comparative work has identified Burr XII, Inverse Paralogistic, and Log-logistic formulations among the strong-performing models across diverse feed datasets. Because the Log-logistic formulation is mathematically equivalent to the existing Groot and generalized Michaelis-Menten models, rumenGP already provides this curve family through those parameterizations. # New Additions ## Burr XII Equation: \[ V(t) = VF \left[ 1 - \left( 1+(rt)^a \right)^{-p} \right] \] Parameters: - VF: asymptotic gas production - r: rate parameter - a: shape parameter - p: shape parameter Potential advantages: - Highly flexible curve shape - Accommodates diverse fermentation profiles - Often produces excellent goodness-of-fit - Useful for comparative model evaluation Potential limitations: - Four-parameter model - Greater risk of overfitting than simpler models - Parameter interpretation may be less intuitive ## Inverse Paralogistic Equation: \[ V(t) = VF \left[ 1 + (rt)^{-a} \right]^{-a} \] Parameters: - VF: asymptotic gas production - r: rate parameter - a: shape parameter Potential advantages: - Flexible sigmoidal behavior - Relatively simple parameterization - Performs well across diverse kinetic profiles Potential limitations: - Less common in rumen literature - Shape parameter can be difficult to interpret - Requires positive incubation times Neither Burr XII nor Inverse Paralogistic should be considered universally superior. Model performance depends on feed type, experimental design, data quality, and model-selection criteria. # Next Steps Additional package capabilities include: ## Importing Manual Datasets ```r gp <- as_rumen_gp( data = my_data, head_col = "Bottle", time_col = "Time", gas_col = "Gas" ) ``` ## Importing Pressure Data ```r gp <- as_rumen_gp( data = my_data, head_col = "Bottle", time_col = "Time", pressure_col = "PSI", pressure_unit = "psi", headspace_volume = 60 ) ``` ## User-Defined Models ```r custom_fit <- fit_custom( data = gp, formula = Gas_mL ~ A * ( Time_h / ( Time_h + K ) ), start = list( A = 150, K = 10 ), lower = c( A = 0, K = 0 ), model_name = "Hyperbolic" ) ``` See: ```r ?as_rumen_gp ?fit_custom ``` for additional details. # Summary rumenGP provides a complete workflow for importing, processing, visualizing, fitting, comparing, and interpreting in vitro rumen gas-production data. Current capabilities include: - ANKOM RF workflows - Manual gas-volume workflows - Pressure-based workflows - Fifteen built-in kinetic models - User-defined models - Model comparison and ranking - Treatment-level model evaluation - Diagnostic tools - Visualization tools Researchers are encouraged to compare multiple biologically plausible models and to consider both statistical performance and biological interpretation before selecting a final model.