22 Building and Sustaining a Data Science Team
A health department may depend on one analyst to maintain a surveillance pipeline and explain its assumptions. If that person leaves, colleagues may struggle to update it. A department building a new team faces related questions: which skills to hire for, how to classify the positions, and where the team should report.
The aim is building a data science capability that does not depend on any single person and that fits the realities of a public health agency. The project management chapter (Chapter 24) covers running a project well. This one covers building and keeping the team that runs the projects.
22.1 Roles on a Public Health Data Science Team
The titles overlap and the boundaries are fuzzy. An epidemiologist, a data analyst, a data scientist, a data engineer, and an informatician can have job descriptions that read almost identically, and in a small agency the same person may answer to all five. It helps to think less about titles and more about the functions that have to be covered:
- Subject-matter framing: turning a program question into an answerable analytic question. This requires epidemiology and domain expertise.
- Data engineering: getting data out of source systems, moving it, and keeping pipelines running. Often the least visible function and the one whose absence hurts most.
- Analysis and modeling: the statistical and computational work of producing the result.
- Communication and delivery: reports, dashboards, and the translation of findings for people who did not run the analysis (Chapter 28).
Analysts often specialize in one of these areas while maintaining working knowledge of the others. The practices in this book (version control, reproducible reporting, validation) are the working competence that lets one person cover more than their specialty while maintaining quality.
On a small team, each person may cover several functions. Check that the team can handle all four. A lightweight skills matrix (functions down one axis, people across the other, proficiency in the cells) is a quick way to see where training or additional staff would help.
22.2 Hiring and Position Descriptions
Hiring a data scientist into a government agency runs straight into a classification problem. Many civil-service systems have no “data scientist” job classification at all, so the role gets filed under whatever existing title is closest: epidemiologist, IT specialist, statistician, research analyst. Each comes with a salary band and a set of required credentials that may not match the work. A position classified as “epidemiologist III” may require a degree the best candidate lacks, or cap salary below what the labor market commands for the skills you actually need.
You cannot always fix the classification, but you can write the position description around the work. State what the person will do: build reproducible analytic pipelines, work in version control, query institutional databases, produce reports and dashboards, communicate findings to program staff. These descriptions help applicants assess whether they are qualified and give interviewers specific skills to evaluate.
When evaluating candidates, a portfolio tells you more than a transcript. A GitHub profile, a published dashboard, or a short take-home exercise1 that mirrors the actual work reveals whether someone can do the job in a way that a list of degrees cannot. A junior hire should show they can write clean, working code and learn quickly; a senior hire should show judgment about when an analysis is good enough, how to structure a project so others can pick it up, and how to mentor.
Choose between a full-time hire and a contractor based on the duration and type of work. A contractor can deliver specialized skills quickly and is the right call for a bounded, well-specified build. A full-time employee accumulates the institutional knowledge (which data sources have known limitations, which stakeholders need what) that makes the next project faster. For anything ongoing, especially recurring surveillance work, continuity usually wins.
22.3 Growing Skills on the Team You Have
Hiring is not the only way, and often not the best way, to build capacity. Experienced analysts may know Excel, SAS, or SPSS well and need time to learn R, Git, and Quarto. (Migrating the workflows themselves, as opposed to the people who run them, is the subject of Chapter 19.) Upskilling the people you already have is frequently more practical than competing for scarce outside hires, and it builds on institutional knowledge that a new hire would have to acquire from scratch.
Free resources are available (Appendix A), but staff also need protected learning time. Procurement rules may affect access to paid training (Appendix B). Learning to work reproducibly while also delivering the weekly reports is nearly impossible if the weekly reports consume every hour. Capacity-building that is not given real, defended time on the calendar does not happen, no matter how good the intentions.
Teams can combine several approaches:
- Structured curricula for foundational skills, working through a book or course as a group (Appendix A).
- Programs built for this purpose, like the DSTT program this book grew out of, which pair training with coaching on the team’s actual projects.
- Communities of practice inside or across agencies, where people doing similar work share patterns and help each other solve problems.
- Pairing and code review (Chapter 27) for learning as well as quality control: reviewing a colleague’s pull request is one of the most effective ways for both people to learn.
It also helps to be explicit that the practices in this book are learned incrementally. Nobody adopts version control, reproducible reporting, validation, and reproducible environments all at once. A team that picks up one practice per quarter is moving at a healthy pace.
22.4 Onboarding and Knowledge Continuity
One measure of this risk is the bus factor: how many people could leave before the work stops? A bus factor of one, the lone analyst who is the only one who understands the pipeline, leaves the team vulnerable to an absence or departure. The goal of onboarding and continuity practices is to get that number above one and keep it there.
Assign backup coverage for each recurring product and give staff time to learn work outside their specialty. Track which products still depend on one person. Chapter 23 covers the handbook, practice runs, and handoff procedures that make this coverage usable.
22.5 Where the Team Sits
Where a data science function reports within an agency shapes what it can do. There is no single right answer, but the common options trade off in predictable ways:
- Inside an epidemiology or surveillance program keeps the team close to the subject-matter questions and the people who will use the findings, at the cost of distance from the data infrastructure and the IT relationships that provision it.
- Inside informatics or IT puts the team next to the databases, servers, and access controls, at the cost of distance from the program questions and sometimes from the analytic culture.
- A standalone analytics unit serving the whole agency can set its own standards and serve many programs, at the cost of having to build relationships with each program.
A related choice is the service model. A centralized team is a shared resource that programs request work from; it builds deep technical capability and consistent standards but can become a bottleneck and can drift away from any one program’s needs. Embedded analysts sit inside programs, close to the work, but can become isolated and reinvent each other’s solutions. Many agencies land on a hybrid: a small central team that sets standards and handles cross-cutting infrastructure, with analysts embedded in the largest programs.
Work with IT and data governance staff regardless of where the team reports. The team’s ability to access data, run pipelines, and publish results depends on it, which is why Section 24.4 and Chapter 25 treat those relationships as part of project planning.
22.6 Retention and Sustainability
Public sector salaries may be lower than industry salaries. Agencies should also consider the working conditions they can offer. People stay in public health data science for the mission, for autonomy and ownership over meaningful work, for the chance to keep growing, and for the conditions that let them do the work well. Each of those is something an agency can actually offer.
Tooling is one of those conditions, and an underrated one. Skilled people who are forced to work in outdated, click-heavy, irreproducible environments (exporting from one system, pasting into another, hand-checking numbers) will leave for somewhere that lets them work the way they know how. The modern practices in this book are not only about output quality; giving capable people good tools is itself a retention strategy.
Succession planning, documentation, and cross-training help the team continue its work when someone leaves. Shared knowledge also lets staff take leave or change responsibilities without leaving a critical system unsupported.
22.7 Further Reading
- Executive Data Science (Peng et al. 2015) is a short, practical book on structuring and managing data science teams and working with executive stakeholders. Useful for team leads and for analysts who want to understand the decisions being made above them.
- Build a Career in Data Science (Robinson and Nolis 2020) covers hiring, job descriptions, interviewing, and career growth from both sides of the table, and is a useful reference for managers writing position descriptions and evaluating candidates.
- CSTE and the Public Health Informatics Institute publish workforce and competency resources specific to public health data and informatics roles, which are helpful when mapping these general practices onto agency classifications.
Take-home exercises should be followed up with an oral interview to walk through the exercise, since candidates may use AI tools to complete them.↩︎