The output is gzipped by default. The published table is one row per grid
coordinate with every dimension's label repeated on each of them, which is
about as compressible as tabular data gets – an order of magnitude off the
plain text, for a format every reader (including readr, arrow, pandas
and duckdb) opens directly. Pass compress = FALSE for plain .csv.
Arguments
- x
An
ivtobject fromread_ivt(), or the path to an.ivtfile – then the table is decoded and written a chunk at a time and the completed grid is never held whole, so a table far larger than memory converts on a small machine. The written file is the same one the in-memory path produces. Seechunk_cells.- path
Output
.csv.gzpath. Defaults to<product_id>.csv.gzin the data cache (ivt_cache_dir("data")). As withivt_write_parquet(), passing the path of an.ivtfile asxconverts it a chunk at a time rather than materialising the whole table.- labels
Passed to
ivt_tidy(): write labelled columns (TRUE, default) or the compact integer-id table (FALSE).- missing
Also write the cell-status table (
ivt_tidy_missing()) to a<name>_missing.csv.gznext topath(TRUE, default), whenxcarries one (i.e. was read withread_ivt(missing = TRUE)). Silently skipped otherwise.- dim_names
How to name the data-dimension columns (passed to
ivt_tidy()and the member sidecar):"slug"(default, the terse structural slug) or"label"(the full dimension name). Slug columns can be labelled on read withlabel_ivt_columns().- language
Output language for labels and label-derived column names (passed to
ivt_tidy()):"en"(default) or"fr". The member sidecar carries both languages regardless.- compress
Gzip the output (
TRUE, default). The extension always tells the truth about the file:.gzis appended to apaththat lacks it, and apaththat already ends in.gzis compressed whatevercompresssays. The written path – which may not be the one passed – is what is returned.- chunk_cells
Only when
xis a file path: how many grid rows to decode, write and drop at a time (defaultgetOption("canivt.chunk_cells", 5e6)). Chunks are cut along the outermost paged dimension – usually geography – which is the axis the file itself pages on, so a chunk is a contiguous run of output rows and no page is read twice. A table whose layout pages on nothing but the straddle window cannot be sliced and is held whole, subject to the usualcanivt.max_cellsguard.- ...
Passed to the CSV writer (
readr::write_csv()if available, elseutils::write.table()).
Examples
path <- system.file("extdata", "98100044.ivt", package = "canivt")
ivt <- read_ivt(path)
out <- ivt_write_csv(ivt, file.path(tempdir(), "98100044.csv"))
basename(out) # .gz appended
#> [1] "98100044.csv"
head(readLines(out), 3) # ... and read back transparently
#> [1] "geo_label,geo_name,geo_uid,geo_level,type,collective,value,symbol,status"
#> [2] "Canada,Canada,2021A000011124,Country,Total - Type of collective dwelling,Collective dwellings occupied by usual residents,24140,NA,NA"
#> [3] "Canada,Canada,2021A000011124,Country,Total - Type of collective dwelling,Population in collective dwellings,657920,NA,NA"
# the same file, decoded and written a chunk at a time, uncompressed
ivt_write_csv(path, file.path(tempdir(), "streamed.csv"), compress = FALSE)
#> Error in readr::write_csv(df, p, ...): unused argument (compress = FALSE)
