11 Automation and Scheduling
Every Monday morning the respiratory surveillance report goes out. Someone opens the project, pulls the latest line list, renders the Quarto document, uploads the HTML, and emails the link to the program team. It takes twenty minutes if nothing goes wrong, longer if the data moved or a package updated over the weekend. It happens fifty-two times a year, and it depends entirely on a person remembering to do it.
The reporting chapter (Chapter 6) showed how to build a report whose numbers update when the data changes, and the dashboards chapter (Chapter 9) showed how to publish one to the web. This chapter is about the next step: making those things happen on a schedule, or in response to an event, without a person in the loop. The recurring weekly report becomes a job that runs at 6am whether or not anyone remembers, and the team finds out only if something breaks.
11.1 What to Automate (and What Not To)
Automation pays off when the same work happens the same way, repeatedly. A weekly surveillance render, a nightly data refresh from a public API (Chapter 18), a validation check that should run every time a file lands: all of these are worth the up-front effort to set up, because the effort is amortized over dozens or hundreds of runs.
The same logic applies to publishing steps that sit downstream of the analysis. A dashboard extract that gets refreshed by hand every Monday is a scheduled job waiting to be written, and both Tableau and Power BI expose APIs for exactly this (Section 10.3).
It pays off poorly, or not at all, for one-off analyses, for work that requires a human decision partway through, and for exploratory work where the whole point is that you do not yet know what you are going to do. A useful test: would you run this code, in this same way, more than a handful of times? If yes, automate it. If no, the time spent automating is time not spent on the analysis.
An automated pipeline can repeat errors across many outputs. Automation presumes the underlying work is already validated (Chapter 3) and reproducible (Chapter 15). Test the pipeline before scheduling it and keep validation in each run.
11.2 Scheduling on Your Own Machine
The simplest form of automation is a saved script that the operating system runs on a schedule. The first requirement is that the script run without anyone clicking anything. Instead of opening the project and rendering interactively, put the render in an R script that can be run headlessly:
# render_report.R
quarto::quarto_render(
input = "respiratory_report.qmd",
execute_params = list(week = Sys.Date())
)This is the same parameterized render from Section 6.6, just invoked from a script instead of by hand. From a terminal, Rscript render_report.R runs it start to finish with no interaction.
Once the work is a single command, the operating system can run it on a schedule. On macOS and Linux that scheduler is cron. A crontab entry is a schedule followed by a command; this one runs the render every Monday at 6am:
0 6 * * 1 cd /Users/you/projects/surveillance && /usr/local/bin/Rscript render_report.R
The five fields before the command are minute, hour, day-of-month, month, and day-of-week. The cd sets the working directory before R starts, so the relative report path and project startup files resolve correctly. Replace both paths for your machine; command -v Rscript in a terminal shows the R executable you normally use. Quote paths containing spaces. The cronR package can create and manage these entries from R if you would rather not edit a crontab by hand. On Windows the equivalent is Task Scheduler, which the taskscheduleR package wraps in the same way.
Machine scheduling is easy to set up and good enough for plenty of internal work, but it has real limits. The machine has to be powered on and awake at 6am Monday, which rules out a laptop that goes home in a bag on Friday. Paths must be absolute, because a scheduled job does not start in your project directory or with your usual environment. If nobody receives an alert when the job fails, the report simply does not appear, and you find out when the program team asks where it is. The next two sections address each of these in turn.
11.3 Continuous Automation with GitHub Actions
For work whose code already lives on GitHub (Chapter 1), GitHub Actions runs jobs on GitHub’s machines instead of yours. A workflow is a YAML file in the .github/workflows/ directory of the repository. It describes when to run (a trigger) and what to do (a sequence of steps), and GitHub spins up a fresh virtual machine each time to carry it out. Because the machine is GitHub’s, it does not matter whether your laptop is open.
A workflow can be triggered by a push, by a manual button, or by a schedule. This one renders a Quarto report every Monday at 6am UTC, can also be run on demand, and publishes the result to GitHub Pages (Section 9.4):
# .github/workflows/render.yml
name: Render surveillance report
on:
schedule:
- cron: "0 6 * * 1" # every Monday at 06:00 UTC
workflow_dispatch: # also allow manual trigger
jobs:
render:
runs-on: ubuntu-latest
permissions:
contents: write
steps:
- uses: actions/checkout@v4
- uses: quarto-dev/quarto-actions/setup@v2
- uses: r-lib/actions/setup-r@v2
- uses: r-lib/actions/setup-renv@v2
- uses: quarto-dev/quarto-actions/publish@v2
with:
target: gh-pages
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}The steps read top to bottom: check out the repository, install Quarto and R, restore the exact package versions recorded by renv (Section 15.1), then render and publish. The maintained building blocks here are r-lib/actions for R setup and quarto-actions for rendering and publishing; the Quarto documentation on GitHub Actions, linked in Appendix A, is the reference for the publishing options.
The cron schedule in a GitHub Actions workflow runs on UTC, not your local time zone, and does not adjust for daylight saving. A “6am Monday” job will land at a different local hour depending on the season. For a surveillance report this rarely matters; when it does, schedule against UTC deliberately.
When a job needs a credential (an API token as in Section 18.5, a database password), that secret never goes in the workflow file or anywhere else in the repository. GitHub stores it under the repository’s Settings > Secrets and variables, and the workflow reads it as an environment variable at run time (the secrets.GITHUB_TOKEN reference above is the built-in example). This is the same discipline as the .Renviron pattern in Chapter 16: credentials live outside version control, and code reads them from the environment.
11.4 Multi-Step Pipelines with targets
A single render is one step. Many real projects are a chain: read the raw extract, clean it, validate it, summarize, then render a report from the summary. When any one input changes, you want to rerun the steps that depend on it and skip the ones that do not. Rerunning everything every time wastes compute on a large dataset; rerunning the wrong subset can leave results based on stale intermediate files.
The targets package tracks dependencies between pipeline steps. Files must be declared as file targets, helper functions must be loaded, and report dependencies must be visible to the pipeline.
This small project reads invented case counts, checks and cleans them, and produces a weekly report. Create an empty project directory with the following files and two subdirectories, R/ and data/. Each code block below gives the complete contents of one file:
surveillance/
_quarto.yml
_targets.R
R/
functions.R
data/
line_list.csv
report.qmd
Install the Quarto command-line application and the R packages targets, tarchetypes, quarto, knitr, and rmarkdown. This example runs locally and does not publish anything.
Save the input as data/line_list.csv:
data/line_list.csv
week,county,cases
1,Fairfax,23
1,Arlington,41
2,Fairfax,30
2,Arlington,38
In R/functions.R, define the helpers. These checks require the expected columns, at least one row, no missing values in those columns, and nonnegative numeric counts. A production pipeline should add checks for its reporting periods, allowed county names, and other input requirements (Chapter 3).
R/functions.R
clean_line_list <- function(x) {
needed <- c("week", "county", "cases")
stopifnot(all(needed %in% names(x)), nrow(x) > 0L,
!anyNA(x[needed]), is.numeric(x$cases), all(x$cases >= 0))
x$county <- trimws(x$county)
x
}
summarize_by_week <- function(x) {
stats::aggregate(cases ~ week, data = x, FUN = sum)
}The pipeline definition goes in _targets.R. Declaring the CSV with format = "file" lets targets detect changes to its contents. Loading the helper functions makes their code available for dependency tracking:
_targets.R
library(targets)
library(tarchetypes)
source("R/functions.R")
list(
tar_target(raw_file, "data/line_list.csv", format = "file"),
tar_target(raw, read.csv(raw_file)),
tar_target(clean, clean_line_list(raw)),
tar_target(summary, summarize_by_week(clean)),
tar_quarto(report, path = "report.qmd", quiet = FALSE)
)The report reads the summary by name in an active code cell. Save this as report.qmd:
report.qmd
---
title: "Simulated weekly case summary"
format: html
---
This report uses invented counts for a pipeline demonstration.
```{r}
#| echo: false
summary <- targets::tar_read(summary)
knitr::kable(summary)
```tarchetypes::tar_quarto() detects that target read and tracks the report’s input and output files. Keep reads explicit in the report: dynamically constructed target names or reads hidden in helper functions may not be detected. Declare additional file dependencies when automatic detection does not cover them.
Save the following configuration as _quarto.yml. Keep Quarto freezing and caching disabled so they do not prevent execution when the pipeline requests a render:
_quarto.yml
project:
type: default
format: html
execute:
freeze: false
cache: falseFrom the project root, run:
targets::tar_make()Open report.html. It should show 64 cases for week 1 and 68 for week 2. Running the pipeline again without changing anything should skip all targets.
Before scheduling the pipeline, check which steps rerun after each kind of change. Changing a count in the CSV should rebuild the data preparation and report. Changing summarize_by_week() should rebuild the summary, and the report if the summary’s value changed. Editing only the report text should rerender the report without repeating data preparation. targets::tar_outdated() lists targets that need updating, and targets::tar_visnetwork() displays the dependency graph. A dependency the pipeline does not know about cannot trigger a rebuild.
11.5 Running in a Reproducible Environment
A scheduled job that relies on whatever packages happen to be installed on the machine can fail after an unrecorded package update. Eventually a package updates, a function’s default changes, and the Monday report renders differently or fails outright. And because nobody changed the code, the cause is hard to find.
The fix is to pin the environment so that the scheduled run matches the one you tested. Recording dependencies with renv (Section 15.1) lets the job restore the exact package versions it was built against; the GitHub Actions example above did this with the setup-renv step. A Docker image can also record the system libraries and tools the job needs (Section 15.2). Section 15.4 covers when each approach is worth the overhead. Record the environment used by the scheduled job.
11.6 Knowing When It Breaks
An automated job can fail without anyone noticing. A dashboard may continue displaying old results for weeks after a refresh fails. Alert the person responsible so they can investigate.
Build in failure signals from the start. Configure GitHub Actions notifications for the people responsible for the workflow and verify that they receive failure alerts. For higher-stakes pipelines, add a step that posts to a Slack or Teams channel on failure so the alert lands where the team actually looks.
Validation belongs inside the automated job in addition to interactive work. A pointblank stop-on-fail check (Section 3.2) placed before publishing can halt the run and trigger an alert when the data fails a required check.
11.7 Sensitive Data and Automation
Most of this chapter’s cloud examples assume the data is safe to send to GitHub’s machines. Much public health data is not. Protected health information, restricted-use files, and data governed by a use agreement generally must not run on GitHub-hosted runners or any other shared cloud infrastructure, because doing so moves the data outside the boundary where its handling is authorized.
This does not rule out automation; it changes where the automation runs. The options are to use self-hosted runners that sit inside the agency’s secure network (while preventing protected data from appearing in uploaded logs, caches, or artifacts), to schedule the job on an approved on-premises server with cron or Task Scheduler, or to use whatever managed analytics platform the agency already operates. The mechanics from earlier in this chapter still apply; only the location changes.
The decision of where a scheduled job may run is a governance decision before it is a technical one. The terms of the relevant data use agreement, the HIPAA considerations, and the secure-environment requirements in Chapter 25 determine what is permissible, and provisioning a scheduled job on a shared server is an IT request that should start early (Section 24.4). Work that out before building the pipeline, not after, because a pipeline built for the wrong environment cannot simply be moved into the right one.
11.8 Further Reading
- The
targetsuser manual (Landau 2021) is the definitive guide to building reproducible pipelines in R, covering branching, cloud storage, and high-performance computing well beyond the minimal example here. - The
r-lib/actionsandquarto-actionsrepositories document the GitHub Actions building blocks for R and Quarto, including ready-made example workflows. - Building Reproducible Analytical Pipelines with R, linked in Appendix A, connects scheduling, pipelines, and reproducible environments into a single framework and goes substantially deeper on each.