23  Onboarding and the Team Handbook

Staff leave with two weeks’ notice. Fellows rotate on an annual cycle. Someone goes on parental leave in the middle of a surveillance season, and the person covering has never opened the script. In each case a recurring product has to change hands in a window too short to write documentation from scratch.

Section 22.4 describes the risks of relying on one person. Maintain a written handbook a new analyst can follow on the first day, which continues to function after its author departs.

23.1 Why Keep a Handbook

Most teams know they should document their work. What they produce is scattered: a README in one repository, a data dictionary in a shared drive folder, setup instructions in an old email thread, and procedures that only one person knows.

A handbook differs in having a single entry point. A new staff member is directed to one document, and everything required is either in it or linked from it. This gives a new analyst instructions they can follow even when the person who wrote them is unavailable.

A handbook reduces time spent repeating setup instructions, leaving colleagues more time to discuss the work itself.

Note

A handbook is distinct from project documentation. Chapter 21 covers READMEs, inline comments, and data dictionaries, which document specific code and specific datasets. The handbook documents how the team operates: where things are, who approves what, and what to do when an extract fails.

23.2 Contents of the Handbook

A workable table of contents for a small public health analytic team:

  1. Orientation: what the team does, which products it delivers and on what cadence, and who the stakeholders are.
  2. Environment setup: installing and configuring R, Positron, Git, and any database drivers on an agency-managed machine. Chapter 15 covers the reproducible-environment side of this.
  3. Access requests: every system the team touches, who owns it, who approves access, which form to file, and typical approval time. This section saves the most calendar time, because a new analyst who files all access requests on the first day becomes productive weeks earlier than one who discovers each requirement upon encountering it.
  4. Repository layout and naming conventions: where code lives, where data lives, and how projects are named (Chapter 2, Chapter 1).
  5. Update protocols, one document per recurring product, covered below.
  6. Data source inventory: every source the team uses, its origin, its update frequency, its upstream owner, and its known defects.
  7. Glossary and metric definitions: what the team means by a case, a rate, a denominator, and a reporting week. Disagreements here produce figures that do not reconcile across products.
  8. Contacts: names and roles for upstream data owners, IT contacts, the disclosure reviewer, and whoever understands why the county denominators are adjusted.

Keep the handbook in version control alongside the code, written in markdown and rendered if desired. Whatever format you choose, record changes and assign responsibility for keeping it current.

23.3 Testing Setup Documentation

Test the setup instructions with a colleague.

Write the setup instructions, then hand them to someone who has never performed the setup, sit beside them, and observe. Provide no assistance. Record every point at which they stop, guess, or ask a question. Those points are the gaps in the documentation, including assumptions the author may have overlooked.

An experienced colleague unfamiliar with the setup can be a useful first reader. Make clear that you are testing the instructions, so readers feel comfortable reporting unclear steps.

Common gaps include assumed software or permissions, screenshots from another version, Unix commands that do not work on a managed Windows machine, and Git instructions that assume prior experience.

Tip

Date screenshots, or record the software version they depict. Screenshots can clarify instructions, but interfaces change. A line reading “screenshots current as of GitLab 17.2, October 2025” tells the reader whether to trust what they are looking at.

23.4 Conventions

Record the choices specific to your team: branch naming, commit size, file naming, documentation requirements, and where restricted data belongs. Link to Section 1.2, Chapter 2, and Chapter 21 for the general practices. Use examples from your own repositories where a rule is easy to misinterpret.

23.5 Update Protocols

For each recurring product the team delivers, write one document covering the steps required to produce it, the inputs and their sources, the expected outputs and their appearance, the known failure modes and the response to each, and whom to contact for anything not listed.

These documents serve two purposes. They are the coverage plan when the owner is unavailable, and they are the training material when someone new assumes the product.

Writing the protocols can reveal knowledge gaps between roles. Where a team has specialized, with one person in biosurveillance, one writing R, and one maintaining the dashboard, the incomplete section of the protocol document may concern a handoff between those roles. Gaps in protocol documentation indicate where knowledge is not shared.

Warning

Include instructions for known failures. A line reading “if the query returns zero rows, check whether the facility list changed; this occurs roughly twice a year” gives the next analyst a specific check to try. Add an entry each time something breaks.

23.6 Practice Repositories and Training Data

New analysts require somewhere to make mistakes. A practice repository seeded with a small realistic project allows someone to learn the team’s Git workflow by branching, breaking something, and opening a merge request without touching a production pipeline. This gives new analysts practice with the workflow.

The data equivalent is a set of standardized training datasets mirroring the structure of real data without the protected information: the same column names, the same types, the same awkward date formats, the same known quirks, with fabricated values. Their utility extends well beyond onboarding. They are what you develop against before entering a secure environment (Section 25.5), what you use to demonstrate a problem to a colleague, and what you provide to a support request or an AI assistant in place of real records (Chapter 26).

Building them is not free, and the task never becomes urgent on its own. The argument that funds it is that the same datasets serve onboarding, development, and debugging. A team establishing shared enterprise datasets (Section 17.9) is already performing most of the specification work, so the training versions should be built in the same effort.

23.7 Cross-Training

Specialization occurs without anyone deciding on it. One person becomes the dashboard person, another the R person, another the one who understands the syndromic surveillance feed. The arrangement is efficient until one of them is unavailable.

Rotating ownership of recurring products is the practical countermeasure, so reserve time for it. Select one product per quarter. The current owner does not run it; they sit with a colleague while that colleague runs it from the protocol document, and every inaccuracy in the document is corrected as it is found. The result is a tested protocol and a second person who has performed the work, which differs from a second person who has read about it.

Complete a rotation before a known absence. Teams that handle parental leave, overlapping PTO, and the question of who covers a script have generally answered that question in advance.

Note

Pair rotation with code review (Chapter 27). Running a product yourself gives you useful context for reviewing changes to it.

23.8 Handoff and Offboarding

Two weeks is the customary notice period and is insufficient for writing documentation from scratch. It is sufficient for executing a checklist that already exists.

Inventory what the departing person owns: every recurring product, repository, scheduled job, credential, and relationship with an upstream data owner. Compile this before it is needed and keep it in the handbook.

Name the person taking over each item.

Run each recurring product once, with the departing person observing and the new owner operating from the protocol document, correcting the document wherever it fails.

Transfer or reissue credentials. Scheduled jobs running under a departing employee’s account are a common cause of a pipeline dying several weeks after their last day, and the failure is particularly damaging because nothing errors visibly until an expected report does not arrive (Section 11.6).

Record the undocumented reasoning. Half an hour spent asking why a county is excluded and why the analysis uses the March file recovers knowledge that is otherwise lost.

Update the handbook, including the contacts section, before the last day.

23.9 Adapting an Existing Handbook

Look for existing handbooks you can adapt. Other agencies have written most of what you need, and many publish it.

Adapting a borrowed handbook is faster than drafting one, and the translation is where much of the value lies. Converting instructions written for a hosted Git service into instructions for a self-hosted instance requires stating how your instance differs. Converting instructions written for a team with a managed analytics server requires stating where your team actually executes code. Each substitution helps identify assumptions your handbook should explain.

Ask through whatever professional network or community platform your team participates in. Agencies are generally willing to share, and they are solving the same problems in parallel without any means of discovering one another.

23.10 Maintaining the Handbook

Without regular maintenance, a handbook can give new staff outdated instructions.

Assign an owner by name. Review on a schedule, quarterly being reasonable for a small team, and attach the review to something that already occurs so that it requires no independent motivation. The natural triggers are onboarding a new person, which provides a chance to test the instructions, and any occasion when a protocol document proves wrong during an incident.

Keep edits inexpensive. A handbook requiring a merge request, three reviewers, and a rendering step will not be updated during the ten minutes when someone notices a broken instruction. Markdown files in the team repository with a low review threshold strike an appropriate balance.

Tip

The best occasion to write a handbook section is the first time you explain something to a colleague. The explanation is already being composed; recording it is cheaper than delivering it twice.

23.11 The IT Liaison

Where possible, involve a named IT representative early enough to understand the team’s systems and priorities. They can identify approval lead times and suggest supported alternatives before infrastructure constraints delay a project. Record the contact’s responsibilities in the handbook, along with the procedures for requesting access, installing approved software, and escalating incidents. Section 24.4 covers planning these requests.

23.12 Further Reading

  • Chapter 21 covers code and project documentation; link to those documents from the handbook.
  • Chapter 22 covers roles, hiring, and retention, of which onboarding is one component.
  • Appendix B covers funding the training that a new analyst’s onboarding plan identifies as a gap.
  • GitLab publishes its complete company handbook at handbook.gitlab.com. It is far larger than a small analytic team requires and shows how an organization can document its operations in a shared handbook. The sections describing how the handbook itself is maintained are the relevant ones.