db <- available.packages()
deps <- tools::package_dependencies(
"dplyr",
db = db,
recursive = TRUE,
which = c("Depends", "Imports", "LinkingTo")
)
base_recommended <- rownames(installed.packages(
priority = c("base", "recommended")
))
length(setdiff(unique(unlist(deps)), base_recommended))13 Managing Package Dependencies
Chapter 12 covers what happens when you call library(): the search path, masking, and how to keep a script’s package list accurate. Chapter 15 covers pinning versions with renv and freezing an entire environment into a container. Before loading or pinning a package, decide whether the project needs it.
Each dependency adds code maintained outside your project. Most of those commitments are worth making; R packages provide tested implementations of common tasks such as reading CSV files. The ones that go badly tend to fail at inconvenient times, in an automated job, on a Linux server, months after the person who introduced the dependency has moved on.
13.1 The Cost of a Dependency
Consider installation, system requirements, compatibility, image size, and maintenance.
Install time. Packages install from binaries in seconds on macOS and Windows. On Linux servers and inside Docker builds they usually compile from source. A dependency tree that installs in ten seconds on a laptop can take several minutes in a container, and uncached rebuilds repeat that work.
System libraries. Some R packages wrap C or C++ libraries that must be present on the machine. Those libraries must also be installed, which may require help from IT. Missing system libraries can prevent installation on another machine.
Compatibility. Changes in indirect dependencies can affect your results, even if the package you chose has not changed.
Container size. Images grow with the dependency tree. A 400 MB image pulls faster and costs less to store than a 4 GB image, and in a scheduled pipeline that pull recurs (Section 11.5).
Maintenance attention. Each dependency is something to update, something whose release notes may matter, and something to explain to the next analyst.
13.2 Measuring Dependency Weight
The recursive tree can be counted before committing to anything. tools::package_dependencies() ships with base R:
pak::pkg_deps_tree() presents the same information as a tree, which identifies which package contributes the weight:
pak::pkg_deps_tree("httr2")Record the package version and repository snapshot or retrieval date with a dependency count. The code above explicitly excludes installed base and recommended packages. Counts alone do not establish whether a dependency is appropriate: consider what it provides and what maintaining an alternative would require.
Package counts omit system requirements. Installing sf also requires geospatial libraries such as GDAL, GEOS, and PROJ, whose versions and compatibility affect a container build. Measure installation time in the intended environment.
Check the SystemRequirements field on a package’s CRAN page before adding anything geospatial, anything that renders images, anything that connects to a database, or anything built on Rcpp. It lists dependencies that an R package manager may not install for you.
Dependency weight matters most where builds are automated and least where a person installs once. Convenience may justify more dependencies during exploration. For a containerized job, check installation time and system requirements before adding a package.
13.3 Meta-Packages
tidyverse is a meta-package: installing it brings in a broad dependency tree, and library(tidyverse) attaches its core packages. Section 12.2 presents the search-path argument for naming packages individually; the dependency argument points the same direction.
Meta-packages are appropriate for interactive work, teaching, exploratory analysis, and any script that genuinely uses most of what they load. The convenience is real and the cost is not noticeable in those settings.
Naming individual packages is appropriate for production pipelines, containers, packages under development, and anything another team must install. A header reading library(dplyr), library(readr), library(ggplot2) states its requirements precisely. library(tidyverse) does not show which attached packages the script uses. The same reasoning applies to tidymodels: if a modeling script uses parsnip and recipes, name them.
13.4 Vetting a Package
Before adopting a package, review its maintenance and compatibility.
For a deep dive on this, see my recent paper, “Ten quick tips to SNIFF out sustainable and secure scientific software” (Nagraj et al. 2026).
Date of last release. A long interval since the last release can reflect stability or inactivity. Read recent issues and maintainer responses for context.
Issue response. Sort the issue tracker by most recent. Maintainer replies on the last several issues indicate an active package. Issues open for eighteen months without comment indicate otherwise.
Reverse dependencies. A package that a hundred other CRAN packages depend on will not disappear without many people noticing. CRAN lists these on each package page.
CRAN check status. The package page reports check results across platforms. Persistent failures on a platform you use are a warning.
Dependency version constraints. Check for exact-version requirements or upper bounds that conflict with other packages you need. A minimum version requirement alone does not pin a package to that version.
Maintainer. Many R packages have a single maintainer. Consider who could maintain the code if that person becomes unavailable. A single-maintainer utility package is unremarkable; a single-maintainer package at the center of a surveillance pipeline warrants a contingency plan.
Independent upstream changes. A package wrapping a vendor’s REST API depends on both the maintainer’s attention and the vendor’s release schedule, and the vendor has no obligation to the package.
The last two criteria compound. A wrapper package that is on CRAN, currently functional, singly maintained, not recently updated, and pinned to an older dependency has no present defect. Its failure mode is that the vendor ships an API change and no one updates the wrapper, at which point a pipeline that ran reliably for a year stops. Using httr2 directly (Section 10.3) gives your team responsibility for maintaining the API calls.
None of these findings is automatically disqualifying. Use the review to identify maintenance needs before adopting the package.
13.5 Auditing an Existing Project
You may inherit a project with an existing dependency list. Before deciding what to add, establish what is already present.
renv::dependencies() scans a project directory for every library(), require(), and :: call and reports which packages the project uses and which file each reference came from:
renv::dependencies()Look for these cases. Packages appearing exactly once are candidates for :: in place of library(), or for removal. Investigate unfamiliar packages before removing them; they may still be needed. Packages appearing in a script but not in the environment documentation are how a pipeline that runs today fails on a new machine.
A first pass over an inherited pipeline commonly removes a substantial fraction of the library() calls without changing any result. Doing this before containerizing is worthwhile, since removing unused dependencies can shorten uncached builds.
13.6 Lifecycle Stages
Tidyverse packages label functions with lifecycle stages: stable, experimental, superseded, and deprecated. The badges appear in the documentation and carry specific meanings.
Stable functions are safe to build on. Superseded functions continue to work indefinitely and receive no new features, so you can plan migration around the project’s needs. Experimental functions may still change their interface; they are reasonable in analysis code you control and worth avoiding in a package other people install. Deprecated functions will eventually error, so plan migration before the function is removed.
tidyr::separate() is superseded, not deprecated. It remains available and receives critical bug fixes. If you adopt separate_wider_delim() or a related replacement, ensure everyone has a compatible version of tidyr.
packageVersion("dplyr")[1] '1.2.1'
Printing the version in a setup chunk helps diagnose differences between installations. It does not enforce a version requirement; use an explicit check or restore the project lockfile for that.
13.7 Version Consistency Across a Team
After choosing dependencies, record and test updates in the project environment. Section 15.1 is the main reference for lockfiles, restoration, and coordinated updates. Function-name conflicts are a separate issue covered in Section 12.4.
13.8 Writing Your Own Instead
Occasionally the appropriate choice is no dependency.
A twenty-line function you control, understand, and test (Chapter 4), living in your own repository, is preferable to a package with forty dependencies that provides the same capability plus ninety others. This applies most clearly to small utilities: a parser for an agency’s file naming convention, a function implementing local suppression rules, a wrapper around a single API endpoint.
Use established implementations for tasks with complex correctness requirements. Date and time zone arithmetic, statistical methods, character encoding, and HTTP all reward using an established implementation, because writing your own can introduce errors that are difficult to detect.
A rough test: code you could write correctly in an afternoon and explain to a colleague in five minutes is often better written than depended upon. Code requiring you to consult a specification first is better taken as a dependency.
Vendoring is the intermediate option, meaning copying a specific function into your own codebase with attribution and a comment recording its origin and date. This suits the case where a package offers one needed function and forty unneeded dependencies, and where the license permits it. Check the license, retain the attribution, and recognize that you have assumed maintenance of the copy.
13.9 Organizing Your Own Packages
Once a team accumulates shared helper functions, the appropriate container is a package (Chapter 14), which provides namespacing, documentation, tests, and a single versioned copy of the shared functions.
The question that follows is how many packages, and the organizing principle matters more than the count. Consider organizing shared utilities by function, such as plotting or standardization. Separate packages for plotting conventions and for standardization routines will each be used across many projects and maintained coherently. Packages split by originating data source will each accumulate a little plotting, a little cleaning, and a little reporting, duplicated across all of them. The tidyverse is organized on this principle: dplyr transforms, ggplot2 plots, readr reads.
Separating independent functions lets developers maintain plotting conventions in one place and gives each package a clear scope. Getting it wrong produces a set of packages that depend on one another and must be updated together, which is a single package with additional overhead.
A new data source should mean a new function in an existing package, or at most a thin ingest package depending on the shared ones, without duplicating the shared utilities.
13.10 Further Reading
- The dependencies chapter of R Packages discusses how to evaluate the costs and benefits of a dependency.
- The
lifecycledocumentation explains each stage badge and what it commits the maintainer to. -
R Packages (Wickham and Bryan 2023) covers the
DESCRIPTIONfile, the distinction betweenImportsandSuggests, and keeping a package’s own dependency footprint small. - Teams migrating from SAS need to learn how R manages dependencies, because the two systems distribute extensions differently. The habit that transfers poorly is loading everything by default; the habit to build is naming what you use and qualifying with
pkg::fnwhere conflicts exist (Section 12.3, Section 19.8).