21 Documentation
Code that runs correctly but cannot be understood is a liability. Six months after writing an analysis, even the original author may not remember what a variable represents, why a particular filter was applied, or where a dataset came from. For a new team member inheriting existing work, undocumented code is essentially opaque.
Documentation is the investment that makes analytical work reusable, reviewable, and transferable. Leave enough information that the next person (or your future self) can pick up the work without starting over.
21.1 READMEs
Every project repository should have a README at the root. This is the first thing anyone sees when they open a project, and it should answer the questions a new reader has immediately:
- What does this project do?
- What data does it use, and where does that data come from?
- How do you run the analysis?
- What are the outputs and where do they go?
A README does not need to be long. A one-page README that answers these questions is more valuable than a ten-page document that describes the history of the project and the background literature.
21.1.1 A Minimal README Template
# Project Title
One or two sentences describing what this analysis does and why it was done.
## Data
What data is used. Where it comes from (source, date accessed, any access
requirements). If data files are not in the repository (e.g., because
they contain PII), explain where they should be placed.
## Running the Analysis
How to reproduce the results. What to run first, what runs next.
Any packages or environment requirements (see environments chapter).
## Outputs
What the analysis produces. Where outputs are written.Commit the README with the project. Update it when something material changes. A stale README is worse than a brief one, because it creates false confidence that the instructions are current.
21.2 Inline Comments
Use comments to explain the reasons for decisions. The code itself shows what is happening; a comment adds value by explaining the reason behind a decision that might otherwise look arbitrary.
A comment that adds no information:
# filter to active cases
cases <- cases |> filter(status == "active")The code already says this. The comment is noise.
A comment that adds information:
# exclude cases with status "pending" and "transferred": these are not yet
# reportable and would inflate counts. see data dictionary for full status codes.
cases <- cases |> filter(status == "active")This explains why only active cases are kept, information that is not in the code itself and that a reviewer or future maintainer needs.
Comment when:
- A non-obvious decision was made and the reason matters (a filter that excludes something, a constant that came from an external source)
- A known data quality issue is being worked around
- An approach was chosen over an obvious alternative for a specific reason
Do not comment:
- Code that reads naturally and does what it says
- Every line as a matter of habit
- To explain what R functions do (that is what documentation is for)
21.2.1 Sectioning Longer Scripts
For analysis scripts that run longer than a screen, section headers help readers navigate. In R, headers recognized by RStudio and Positron use # ---- or # ==== suffixes:
# Load data ---------------------------------------------------------------
# Clean and filter --------------------------------------------------------
# Calculate rates ---------------------------------------------------------
# Produce outputs ---------------------------------------------------------These create an outline in the document navigator (the panel at the bottom of the script editor in Positron) that lets readers jump directly to a section without scrolling.
21.3 Documenting Functions with roxygen2
When a function will be reused across scripts, projects, or team members, it should be documented at the function definition, not in a separate file. The standard tool for this in R is roxygen2.
roxygen2 documentation lives in comments immediately above the function definition, starting with #'. The most important tags:
@param: documents each argument: its name, expected type, and what it does@return: describes what the function returns@examples: shows how to call the function, ideally with a runnable example
A minimal documented function:
#' Calculate age-adjusted rate per 100,000
#'
#' Applies direct standardization to compute an age-adjusted rate using
#' the 2000 U.S. standard population weights.
#'
#' @param cases A data frame with columns `age_group`, `events`, and `population`.
#' @param std_pop A data frame with columns `age_group` and `std_weight`. Defaults
#' to the 2000 U.S. standard population.
#'
#' @return A single numeric value: the age-adjusted rate per 100,000 population.
#'
#' @examples
#' calculate_age_adjusted_rate(flu_cases_2023)
calculate_age_adjusted_rate <- function(cases, std_pop = us_std_pop_2000) {
# ... function body
}Even a brief description, one @param per argument, and one @return is substantially better than no documentation at all. A future reader knows what the function expects without having to read and mentally execute the implementation.
21.3.1 When Functions Are More Than Utilities
Keep project-specific helpers with the analysis. For functions shared across projects, Chapter 14 explains packaging and distribution; Section 13.9 covers how to divide the shared code into maintainable packages.
21.4 Project-Level Documentation
Analysis projects involve more than code: there are data sources, methodological decisions, known data quality issues, and interpretive context that do not belong in code comments or README files but that someone needs to know.
21.4.1 Data Dictionaries
If a project uses a dataset with non-obvious columns (abbreviated names, coded values, or context-dependent meanings), maintain a data dictionary alongside the analysis. This can be a simple table in a Markdown or Quarto document:
| Column | Type | Description |
|---|---|---|
case_id |
character | Unique case identifier; not persistent across system updates |
epi_class |
character | Epidemiologic classification: confirmed, probable, suspect |
rpt_cnty_cd |
character | FIPS code of the county of report, stored as text to preserve leading zeros (not necessarily the county of residence) |
onset_dt |
date | Symptom onset date as reported; missing if not collected for this condition |
A data dictionary does not need to document every column. Document the columns that are confusing, the ones that carry important caveats, and the ones where the name does not tell the full story.
A machine-readable dictionary can also supply labels and definitions to the pipeline. Keeping metric definitions in a machine-readable file that both the R code and the dashboard consume is covered in Section 10.7.
21.4.2 Decision Logs
Analyses involve choices that are not visible in the code: why a particular date range was chosen, why a jurisdiction was excluded, why one method was preferred over another when both were defensible. These decisions are worth recording, because they will be questioned by a reviewer, by a stakeholder, or by a future analyst who is maintaining the work.
A simple decision log entry might look like:
2024-11-14: Excluded 2020 data from trend analysis. COVID-related disruptions to case reporting caused substantial undercounting across multiple conditions; the 2020 data points would distort trend estimates in ways that are not informative about the underlying epidemiology. Analysis covers 2018–2019 and 2021–2023.
This does not need to be elaborate. A short paragraph with a date and a rationale is enough to reconstruct the reasoning later.
21.4.3 Methods Documentation
For analyses with non-trivial statistical methods, a methods document that describes the approach in plain language (separate from the code) is worth maintaining. This is the document a reviewer (see Chapter 27) or stakeholder can read to understand what was done without reading R code. It should describe the data sources, the analytical approach, known limitations, and any validation steps taken.
This document can be a section of the final report, a standalone Quarto document in the repository, or a README section, wherever it is most likely to be found and updated.
21.5 Documentation as a Team Practice
Update project documentation in the same change as the code or data definition it describes. A reviewer should be able to find the inputs, understand the decisions, and reproduce the expected outputs (Section 27.2).
Record the team’s minimum documentation requirements in its handbook (Section 23.4), and link to each project’s README and data dictionary. Keep setup and team procedures in the handbook; keep project-specific explanations with the project.