Overview of data collection
The objective of the U.S. COVID-19 County Policy (UCCP) Database is to systematically gather, characterize, and assess variation in U.S. county-level COVID-19-related policies. This study summarizes the first phase of data collection, in which the research team gathered policy data for 171 counties in California, Louisiana, Mississippi, New Jersey, New York, Texas, and Utah (Supplemental Table 1). The total population of these counties is 90.4 million. While these counties are not nationally representative, they include over a quarter of the U.S. population and are diverse with respect to geography, race/ethnicity, and politics [19]. The reasons for the selection of the given counties is described in the Supplemental Methods.
For these counties, we gathered data in January-March 2021 on COVID-19-related policies that were in effect at that time, capturing a cross-sectional picture of county policies during that period. Data collection proceeded for approximately eight weeks. Data were also collected for the corresponding states in which these counties are nested, with two waves of state policy data collection conducted in January 2021 just before county data collection, and again in February–March 2021 just after county data collection was completed.
Policy coding
We gathered data on 22 policies within three overarching domains: containment and closure, economic support, and public health measures (Table 1). These were in part modeled on national and state policy data currently collected through the Oxford COVID-19 Government Response Tracker, excluding those not applicable to counties (e.g., border control), and including additional policies that are primarily relevant at the county level that may affect health (e.g., housing support). For each policy, the study team assessed a sample of current policies across rural and urban counties in several states, and developed scoring criteria to assess comprehensiveness of the policy. These were aligned with the way in which policies and associated restrictions were framed at the county level (Supplemental Table 2). For example, one policy indicator captured public events with the following categories: minimal (≥ 50% capacity) limitations, major (< 50% capacity) limitations, recommended cancellation, and required cancellation. Each indicator also included a category to capture the scenario where there was explicitly no relevant restriction or program in place, and a category for missing where there was no information available to determine the policy/program in place. For school closures in this first phase of data collection, in counties with more than one school district or university, data collectors coded the district or university with the most stringent/comprehensive policy.
Data collection
Data collectors used the Research Electronic Data Capture (REDCap) data entry and management platform [20]. For each policy, data collectors abstracted scoring related to comprehensiveness, the effective date of the policy in place at the time of data collection, and source documentation. Data collectors gathered data from a variety of sources, including government websites, policy and government response summaries and databases, press releases, news articles, and social media posts by government organizations. Details on data collection are provided in the Supplemental Methods.
Imputation of missing data
Despite a thorough search of multiple sources, in some cases there was no information on the county policies of interest. This ranged from 4.7% for school closures to 50.3% for utility support (Supplemental Table 3). This was especially the case in rural counties (39.9% missingness across all policies), which are less likely to have robust public health departments than urban counties (23.8% missing) (Supplemental Table 4). For these, data on state policies were used to impute county policies. Since state data were gathered in two waves—just before and after county data collection—we used the closest state survey date to fill in the corresponding policy for missing county data. If there was no information on the policy at the state level either, then the policy remained coded as missing.
Additionally, we combined policies that were missing and those with clear documentation of no relevant restrictions, in essence assuming that there was no policy in place for those counties with no documented policy. While this assumption may not always be accurate, if policies were difficult for our trained staff to find after a thorough search, they were also likely to be difficult for residents to locate, resulting in no substantive restrictions in place. For example, counties with no clear documentation of a face covering policy and those with clear documentation that no face coverings were required were combined.
Primary analysis
First, we calculated univariate distributions of each policy within each of the three overarching policy domains, documenting the range of comprehensiveness of each policy.
We then calculated bivariate Spearman’s correlations between each pair of policies. This type of non-parametric analysis examines the extent to which two ordinal ranked variables are associated with one another. In this case, it assessed the extent to which policies co-occurred in a given county, reflecting the fact that governments often implement bundles of policies on related issues [21, 22].
Next, we examined geographic distribution of the policies. First, for each state, we tabulated the mean number of policies implemented by counties in that state in each of the three policy domains. We then produced heatmaps for each state, documenting the number of policies present in each county (range 0–22). For these plots, policies were coded as binary (i.e., no policy versus any policy).
Secondary analyses
We next conducted several secondary analyses to account for possible bias introduced by the imputation process. To do so, for each analysis above, we calculated the results using data obtained directly from counties only, without imputation using state data.
Second, to allow for greater variation and nuance in policy landscapes, we calculated an overall index of comprehensiveness, rather than simply the number of policies implemented. For this analysis, if there was no policy, this was coded as 0, the most comprehensive were coded as 1, and intermediate categories were fractions thereof. For example, for public events, no restriction was coded as 0, minimal (≥ 50% capacity) limitations was 0.25, major (< 50% capacity) limitations was 0.50, recommended cancellation was 0.75, and required cancellation was 1. When summing the policies to achieve a total policy score for each county, the range was again 0–22, with non-integer values possible.
Finally, we used principal component analysis (PCA) as an alternative technique to create a composite index of policy comprehensiveness (see Supplemental Methods for details). We also examined the contributing policies to each of the principal components created by this technique to assess the relationships between the different policies in a different manner than the pairwise correlations described above.
