
Determine regions to join to correct for likely geocoding anomalies
Source:R/tongfen_anomalies.R
tongfen_anomaly_joins.RdLooks for regions with surprising drops in the timeline of a count variable that are complemented by a neighbouring region, as explained in `tongfen_detect_anomalies`, and joins them. This gets repeated on the joined regions until there are no more regions left that qualify to get joined. In each round a region only gets joined with one other region, the most surprising regions go first.
Joining regions trades geographic detail for timelines that are consistent over time. The parameters control how aggressively regions get joined and are best calibrated on the data at hand, erring on the side of joining too few regions risks keeping geocoding problems, erring on the other side risks removing real changes and needlessly coarsens the geography.
The result can be used to join the regions via `tongfen_join_regions`, or to update a correspondence via `tongfen_join_correspondence`.
Usage
tongfen_anomaly_joins(
data,
variables,
id = "TongfenID",
neighbours = NULL,
rel_scale = 0.25,
abs_scale = 200,
p = 4,
surprise_cutoff = 0.15,
total_surprise_cutoff = 0.75,
cutoff_fact = 0.6,
surprise_reduction_const = 0.15,
sum_fact = 0.7
)Arguments
- data
data on a common geography, with one row per region, for example as returned by `tongfen_aggregate` or `get_tongfen_ca_census`. Needs to be of class sf unless `neighbours` is specified.
- variables
names of the columns holding the timeline of a count variable like population or dwellings, in temporal order. Changes from or to a missing value are not surprising and don't make up for surprising changes in neighbouring regions, and joined regions are missing a value if one of the regions they are made up of is. Replace missing values by zero beforehand if they stand for regions where nothing got counted
- id
name of the column that uniquely identifies the regions, default is "TongfenID"
- neighbours
optional, neighbouring regions as a table with the identifiers of pairs of neighbouring regions in the first two columns, or as a neighbours list like the ones returned by `spdep::poly2nb`. By default all regions with intersecting geometries are neighbours, which can miss neighbours if the geometries have been simplified and don't share their boundaries any more.
- rel_scale
relative decrease that is half way to full surprise, default is `0.25` for a 25% drop
- abs_scale
absolute decrease that is half way to full surprise, default is `200`
- p
exponent of the norm used to combine the surprises across the timeline into the total surprise. Large values focus on the most surprising change, 1 adds up the surprises of all changes, default is `4`
- surprise_cutoff
changes with larger surprise count as surprising, default is `0.15`. Only regions with at least one surprising change are candidates
- total_surprise_cutoff
only regions with larger total surprise are candidates, default is `0.75`
- cutoff_fact
join regions if the total surprise after joining is lower than this share of the total surprise of the candidate region, default is `0.6`
- surprise_reduction_const
join regions if joining lowers the total surprise by more than this, default is `0.15`
- sum_fact
join regions if the total surprise after joining is lower than `cutoff_fact * sum_fact` times the sum of the total surprises of both regions, default is `0.7`
Value
A tibble with one row for each region that gets joined with other regions, with the identifier of the region, the identifier of the joined region it becomes part of in the column named like the identifier with suffix `_joined`, by default `TongfenID_joined`, and the `round` in which the region first got joined to another region. The identifier of a joined region is the smallest identifier of the regions it is made up of.
Examples
# Correct 2001 through 2021 dissemination area level population timelines in the
# City of Vancouver for likely geocoding problems
if (FALSE) { # \dontrun{
datasets <- c("CA01","CA06","CA11","CA16","CA21")
meta <- meta_for_additive_variables(datasets,"Population")
data <- get_tongfen_ca_census(regions=list(CSD="5915022"),meta=meta,level="DA",base_geo="CA21")
joins <- tongfen_anomaly_joins(data,paste0("Population_",datasets))
corrected_data <- tongfen_join_regions(data,joins,meta)
} # }