18  APIs and Public Data Sources

Not every analysis starts with data your agency collects. The denominators for your rates come from the Census Bureau, the national counts you compare against come from CDC, and both live on someone else’s server, published for exactly this kind of use.

The usual way to get that data is to browse to the portal, click through some filters, hit Export, and download a CSV, which can leave the download settings unrecorded. Six months later, nobody remembers which filters were applied or whether the provider has revised the numbers since, and the download is the one step of the analysis that cannot be re-run.

Scripted downloads record the request and make it repeatable. In R, this means using a dedicated package when one exists, building requests with httr2 when one does not, handling API keys, and structuring the pull so the rest of the analysis stays reproducible.

Note

This chapter is about public data sources, where the data is intended for download and the main concerns are reproducibility and politeness. Getting access to your agency’s internal databases is covered in Chapter 16, and the governance process around restricted data in Chapter 25.

18.1 What an API Request Looks Like

A web API (application programming interface) lets code request data or operations from a service. Its endpoints are URLs; the example here returns data. CDC publishes NNDSS weekly notifiable disease counts on its open data portal, and this URL asks for the 2024 pertussis records:

https://data.cdc.gov/resource/x9gk-5huc.json?label=Pertussis&year=2024

The pieces: data.cdc.gov is the portal, x9gk-5huc identifies the dataset, .json is the response format, and everything after the ? is a set of query parameters that filter the result. Paste it into a browser to see JSON, a plain-text format for structured data. The response contains records and named fields.

JSON values can be text even when they represent numbers. Inspect the provider’s field definitions before converting types (Section 18.7).

18.2 Where Public Health Data Lives

A few sources cover most of what public health teams pull from outside their agency:

Source What’s there How to access it
data.cdc.gov CDC’s open data portal: NNDSS, provisional deaths, vaccination coverage, BRFSS, and hundreds more Socrata API; RSocrata or httr2
US Census Bureau ACS, decennial census, population estimates (your rate denominators) tidycensus (free API key required)
CDC WONDER Mortality, natality, and other query systems Web query tool; an API that accepts XML-formatted requests
HealthData.gov Cross-agency HHS datasets Varies by dataset
State and local open data portals Many jurisdictions run their own Socrata or CKAN portals Same Socrata patterns as data.cdc.gov

18.3 Use a Package When One Exists

Check for a maintained package before implementing requests yourself. tidycensus handles Census queries and returns estimates with margins of error. RSocrata handles pagination and date conversion for Socrata portals. Confirm the package supports the dataset and fields you need; a wrapper does not remove the need to validate the result.

For example, this tidycensus call requests county population estimates for Virginia from the 2023 five-year ACS. Request a free key at https://api.census.gov/data/key_signup.html and store it with tidycensus::census_api_key("YOUR_KEY", install = TRUE). This writes the key to .Renviron for future sessions (Section 18.5). Restart R, or run readRenviron("~/.Renviron"), before making a request in the current session. Then run:

va_pop <- tidycensus::get_acs(
  geography = "county",
  state = "VA",
  variables = c(population = "B01003_001"),
  year = 2023,
  survey = "acs5"
)

For a Socrata dataset, RSocrata can handle the page requests:

pertussis <- RSocrata::read.socrata(
  "https://data.cdc.gov/resource/x9gk-5huc.json?label=Pertussis&year=2024"
)

Kyle Walker’s Analyzing US Census Data (Walker 2023) covers Census data and tidycensus in detail. The httr2 example below gives you control over the request and archiving steps when a package does not provide the options you need.

18.4 Building Requests with httr2

When no suitable wrapper exists, httr2 lets you construct and inspect a request before sending it. Record the endpoint, filters, selected fields, and any version parameter. Keep requests in a separate acquisition script so changing a report does not repeatedly contact the service.

The complete example below uses httr2, jsonlite, and dplyr. Install those packages before running it. The function definitions can be run without contacting the API; the download happens only when you call fetch_nndss() in Section 18.8.

First, define the request for the 2024 pertussis records:

nndss_request <- function() {
  req <- httr2::request("https://data.cdc.gov/resource/x9gk-5huc.json") |>
    httr2::req_url_query(
      `$where` = "label='Pertussis' AND year='2024'",
      `$limit` = 5000,
      `$order` = ":id"
    ) |>
    httr2::req_retry(max_tries = 3) |>
    httr2::req_throttle(capacity = 30, fill_time_s = 60)
  token <- Sys.getenv("SOCRATA_APP_TOKEN")
  if (nzchar(token)) {
    req <- httr2::req_headers(req, `X-App-Token` = token)
  }
  req
}

req_url_query() encodes the filter and other query parameters. The page size and ordering will stay fixed while the download function changes the offset. The optional token comes from your environment (Section 18.5); it is not part of the URL or the saved metadata.

18.4.1 Pagination

Check the API’s row limit and completion rule. A successful HTTP response can contain only the first page. For offset-based pagination, use a stable ordering and continue until the documented completion condition is met. Socrata recommends an explicit order such as :id. Changes to the dataset during pagination can still affect results; use a provider snapshot when available.

Set a maximum number of requests to catch accidental loops, and treat reaching it as an incomplete pull. Compare the final row count and coverage with expectations. A round count is a reason to investigate, but neither a round nor an irregular count proves completeness.

18.4.2 Retries and rate limits

Use bounded retries for transient errors and obey the provider’s request limits. An authentication failure or invalid query requires correction, not repeated requests. The request above allows at most three attempts per page. By default, req_retry() retries HTTP 429 and 503 responses and respects a Retry-After header when the server supplies one. req_throttle() limits the request rate to 30 requests per minute here; choose a rate appropriate to the provider. Section 18.8 shows how to archive each page and record whether the pull completed.

18.5 API Keys and Tokens

Many APIs require a key, and others work better with one. The Census API requires a free key. Socrata portals work anonymously but throttle anonymous traffic aggressively; a free app token moves you to a much more generous rate limit. Keys are credentials, and the discipline is the same as for database passwords in Section 16.5: keep them in .Renviron, read them with Sys.getenv(), and never commit them.

In a scheduled pipeline, supply the credential through the job’s approved secret store (Section 11.3). Avoid writing credentials into request logs, filenames, or archived metadata.

18.6 Separate the Pull from the Analysis

Do not put an API call at the top of an analysis script. Put it in its own script, write the raw response to data/raw/ under a dated filename or directory (Section 2.3), and have the analysis read the file.

Saving the response preserves the data used by the analysis. An API reflects the data as of right now, and many public health sources revise: CDC’s provisional death counts, for example, are updated for recent weeks as certificates arrive, so the same query run a month apart returns different numbers. The archived pages record what you retrieved in that run. Read and combine those files when reproducing the analysis. Separating the pull also means the analysis renders offline and keeps working when the API is down or the dataset moves, and iterating on a report no longer re-downloads the same data on every render.

In an automated pipeline, the pull becomes its own step, scheduled with the tools in Chapter 11 or declared as a targets node (Section 11.4) so downstream steps rerun only when a new file arrives.

18.7 Validate What Comes Back

Check the response schema, types, and values before using the data. Required columns may be renamed or removed; a field may disappear from every record when all its values are missing. Distinguish an allowed all-missing field from a missing required identifier.

Define how to interpret missing values and suppression flags from the provider’s documentation. Do not turn a suppressed count into zero. Reject unexpected text during numeric conversion; a warning that becomes NA can otherwise disappear into later calculations.

Also check the requested year, geography, disease, row count, and reporting-period coverage. These checks catch a syntactically valid response to the wrong query. Chapter 3 covers implementing checks and stopping the pipeline when required checks fail.

For this example, require the disease label, year, and reporting week. Convert the count field m1 only after checking that its nonmissing values contain integer text. The validator allows an entirely missing count field but rejects missing identifiers:

validate_nndss <- function(x) {
  required <- c("label", "year", "week")
  if (!is.data.frame(x) || !all(required %in% names(x)) || nrow(x) == 0L) {
    stop("Expected a nonempty table with label, year, and week")
  }
  if (anyNA(x[required]) || any(x$label != "Pertussis")) {
    stop("Unexpected or missing label, year, or week")
  }
  # Socrata can omit a field from every row when all values are missing.
  if (!"m1" %in% names(x)) {
    x$m1 <- NA_character_
  }
  for (field in c("year", "week", "m1")) {
    raw <- as.character(x[[field]])
    present <- !is.na(raw)
    if (any(!grepl("^[0-9]+$", raw[present]))) {
      stop("Unexpected non-integer text in ", field)
    }
    value <- suppressWarnings(as.integer(raw))
    if (any(present & is.na(value))) {
      stop("Integer conversion failed in ", field)
    }
    x[[field]] <- value
  }
  if (any(x$year != 2024L) || any(x$week < 1L | x$week > 53L)) {
    stop("Unexpected year or reporting week")
  }
  # Missing m1 values remain missing; retain provider flag fields for interpretation.
  x
}

The original response files will retain the provider’s fields and values. This validator leaves missing counts missing and keeps any flag fields in the table. Check those flags and the dataset’s field definitions before interpreting a missing value. The checks here establish the expected schema and basic values; they do not establish geographic coverage or completeness of reporting.

You can check the missing-count behavior without making a request:

example_records <- data.frame(
  label = c("Pertussis", "Pertussis"),
  year = c("2024", "2024"),
  week = c("1", "2"),
  m1 = c("3", NA_character_)
)

validate_nndss(example_records)
      label year week m1
1 Pertussis 2024    1  3
2 Pertussis 2024    2 NA

The count becomes numeric, and the missing count stays missing. Replacing "3" with an unexpected value such as "suppressed" would stop validation. Decide how to interpret such a flag from the provider’s documentation before adding a conversion rule.

18.8 Download and Reuse

Use one reader for both a new download and later reuse. It combines the saved JSON pages and applies the validator:

read_nndss_pages <- function(paths) {
  pages <- lapply(paths, function(path) {
    jsonlite::fromJSON(path, simplifyVector = TRUE)
  })
  validate_nndss(dplyr::bind_rows(pages))
}

The download function creates a new archive directory and saves each response before moving to the next page. It continues until the API returns an empty page. Reaching max_pages stops the pull with an error, so a loop limit cannot be mistaken for successful completion:

fetch_nndss <- function(out_dir, max_pages = 100L) {
  stopifnot(
    length(max_pages) == 1L,
    is.finite(max_pages),
    max_pages >= 1,
    max_pages == floor(max_pages)
  )
  if (!dir.create(out_dir, recursive = TRUE, showWarnings = FALSE)) {
    stop("Choose a new output directory")
  }
  req <- nndss_request()
  paths <- character()
  complete <- FALSE
  for (i in seq_len(max_pages)) {
    resp <- httr2::req_perform(httr2::req_url_query(
      req,
      `$offset` = (i - 1L) * 5000L
    ))
    path <- file.path(out_dir, sprintf("page-%04d.json", i))
    writeBin(httr2::resp_body_raw(resp), path)
    paths <- c(paths, path)
    if (length(httr2::resp_body_json(resp)) == 0L) {
      complete <- TRUE
      break
    }
  }
  if (!complete) {
    stop("Page limit reached; incomplete archive, do not analyze")
  }
  result <- read_nndss_pages(paths)
  metadata <- list(
    retrieved_utc = format(Sys.time(), tz = "UTC", usetz = TRUE),
    endpoint = "https://data.cdc.gov/resource/x9gk-5huc.json",
    filter = "label='Pertussis' AND year='2024'",
    order = ":id",
    page_size = 5000L,
    rows = nrow(result),
    complete = TRUE
  )
  # No credential is written to the archive.
  jsonlite::write_json(
    metadata,
    file.path(out_dir, "manifest.json"),
    auto_unbox = TRUE,
    pretty = TRUE
  )
  result
}

The completion manifest records the endpoint, filter, ordering, page size, retrieval time, and row count. It is written only after pagination and validation succeed. If a request fails partway through, the pages already downloaded remain available for inspection, but there is no completion manifest. Start a fresh archive when retrying so files from different pulls are not mixed.

To download, run the four function definitions above, then call fetch_nndss() from an acquisition script or the console. This example puts the archive in a directory named with its UTC retrieval time:

archive_dir <- file.path(
  "data",
  "raw",
  format(Sys.time(), "%Y%m%dT%H%M%S", tz = "UTC")
)
pertussis <- fetch_nndss(archive_dir)

Record the chosen archive path with the analysis. For a later offline run, set archive_dir to that existing directory and use the saved pages:

metadata <- jsonlite::read_json(file.path(archive_dir, "manifest.json"))
stopifnot(isTRUE(metadata$complete))
page_files <- sort(list.files(
  archive_dir,
  pattern = "^page-[0-9]+[.]json$",
  full.names = TRUE
))
pertussis <- read_nndss_pages(page_files)
stopifnot(nrow(pertussis) == metadata$rows)

A missing manifest stops this reuse step. The row-count check detects some accidental changes to an archive; it does not prove the source data was complete or unchanged during pagination. Keep the saved responses with the analysis and apply the coverage checks described in Section 18.7. Report rendering should read this archive without making a new API call.