--- title: "Streaming, cache and file generations" output: rmarkdown::html_vignette vignette: > %\VignetteIndexEntry{Streaming, cache and file generations} %\VignetteEngine{knitr::rmarkdown} %\VignetteEncoding{UTF-8} --- ```{r, include = FALSE} knitr::opts_chunk$set( collapse = TRUE, comment = "#>" ) ``` ```{r} library(raisr) ``` ## How an archive is read Each `.7z` archive holds a single text file: Latin-1, one header line, one record per line. Extracted, a regional file is several gigabytes of text (the establishments file of 2023 alone is 1.3 GB), and a naive `read.csv()` on it needs many times that in memory. `rais_read()` never extracts the archive. It opens a read connection on the compressed entry through `archive::archive_read()`, so `libarchive` decompresses on the fly, and hands that connection to `readr::read_delim_chunked()`. Every chunk (500,000 lines by default) goes through a callback that: 1. keeps only the records whose municipality code starts with one of the requested states (`uf`); 2. keeps only the requested `columns`. Only what survives the callback is accumulated. Everything is read as character first, and typed once at the end (wages as doubles, integer codes as integers, the Ministry's ignored marker as `NA`), because the two file generations write numbers differently. The chunk size is a memory/speed trade-off: 500,000 lines of 60 short columns is a few hundred megabytes at peak. Lower it on a small machine; the result is the same: ```{r} f <- system.file("extdata", "2022", "RAIS_VINC_PUB_NORDESTE_sample.7z", package = "raisr") a <- rais_read(f, chunk_size = 5L, verbose = FALSE) b <- rais_read(f, chunk_size = 100000L, verbose = FALSE) identical(a, b) ``` ## Two header generations Up to the RAIS 2022 (and in the partial and legacy editions) the files are `;`-separated with decimal comma, padded fields and headers such as `Vínculo Ativo 31/12`; the ignored marker is `{ñ class}`. From the RAIS 2023 onwards the files are `,`-separated with quoted strings, decimal point and headers such as `Ind Vínculo Ativo 31/12 - Código`, and two columns were added (`ind_vinculo_abandonado`, `categoria_trabalhador`). `rais_read()` detects the generation from the header line and maps both to the same normalized names, so different years stack directly: ```{r} new <- rais_read(system.file("extdata", "2024", "RAIS_VINC_PUB_NORDESTE_sample.7z", package = "raisr"), columns = c("municipio", "vinculo_ativo_31_12", "vl_remun_media_nom"), verbose = FALSE) old <- rais_read(f, columns = c("municipio", "vinculo_ativo_31_12", "vl_remun_media_nom"), verbose = FALSE) rais_stock(rbind(new, old)) ``` `rais_layout()` lists the two spellings of every column: ```{r} subset(rais_layout(), original != original_2023)[, c("column", "original", "original_2023")] ``` Before 2018 the files are per state and shorter: the monthly wage columns exist from 2015 (named `vl_rem__cc` up to 2017, `vl_rem__sc` after), and the oldest years use other classifications (`cbo_ocupacao`, `grau_instrucao_2005_1985`). Columns absent from a year come back as `NA` when years are stacked with `rais_fetch()`. ## The cache `rais_download()` stores every archive under `rais_cache_dir()`, mirroring the server: one sub-folder per year and edition, the server's file name inside. ``` /2024/RAIS_VINC_PUB_NORDESTE.7z /2024/RAIS_ESTAB_PUB.7z /2023-legado/RAIS_VINC_PUB_NORDESTE.7z /2017/PE2017.7z ``` A file already in the cache is not downloaded again (`status = "cached"`) unless `force = TRUE`. This is what lets you call `rais_fetch()` for the same years repeatedly, with different states or columns, and pay for the download only once. `rais_read()` takes the reference year from that folder name (or from the file name of pre-2018 files); pass `year` explicitly for an archive kept elsewhere. The cache location is resolved from the `cache_dir` argument, the `RAISR_CACHE_DIR` environment variable, the `raisr.cache_dir` option, or a folder under `tempdir()` (the CRAN-compliant default, wiped with the R session). A persistent cache is strongly recommended: a regional file is hundreds of megabytes and the Ministry's server is slow. ```{r, eval = FALSE} Sys.setenv(RAISR_CACHE_DIR = "~/dados/rais") rais_cache_list() rais_cache_clear(2019) # drop one year ``` ## Editions: final, partial and legacy Some years are published first as a preliminary edition, in a folder ` Parcial`, and when a year is re-published the previous files are kept in a `Legado` sub-folder. `rais_available()` reports them with `edition = "parcial"` and `"legado"`, and `rais_download()`, `rais_fetch()` and `rais_files()` accept `edition` to fetch them; they are cached under `-parcial` and `-legado`. The default, `"final"`, is the current official file of the year. ```{r} rais_files(2023, uf = "PE", edition = "legado")$url ``` ## Network The server is a plain FTP server. Outbound access on port 21 **and** on the high ports used by passive mode is required; corporate firewalls often block one or both. When the server cannot be reached `rais_available()` returns an empty tibble with a warning, and `rais_download()` reports `status = "error"` after three attempts, without raising, so that a long loop over years can continue and be inspected afterwards through the `"download"` attribute of `rais_fetch()`.