This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| The COVID Tracking Project | |
|---|---|
| Name | The COVID Tracking Project |
| Formation | March 2020 |
| Dissolution | March 2021 |
| Purpose | Collecting and publishing COVID-19 data for the United States |
| Headquarters | Brooklyn, New York |
| Founders | - The Atlantic - volunteers |
| Website | (archived) |
The COVID Tracking Project The COVID Tracking Project was a volunteer-run initiative that compiled, standardized, and published data on the COVID-19 pandemic in the United States during 2020–2021. Drawing on state and territorial reports, independent journalism, and contributions from civic technologists, public health researchers, and data scientists, it became a widely cited source for policymakers, journalists, and researchers during the early phase of the COVID-19 pandemic in the United States. The project coordinated with academic institutions, non-profit organizations, and media outlets to make datasets and visualizations broadly available.
The project began in March 2020 in response to inconsistent reporting by many state governments and local agencies; it sought to provide an authoritative national compilation similar to efforts by Centers for Disease Control and Prevention, Johns Hopkins University, and international trackers like World Health Organization dashboards. Staff and volunteers included contributors from organizations such as The Atlantic, New York Times, Washington Post, ProPublica, FiveThirtyEight, Harvard University, Columbia University, Yale University, and civic-tech groups associated with OpenStreetMap, GitHub, and DataKind. The team published daily updates, methodological notes, and an open-source codebase that intersected with projects from Kaiser Family Foundation, Robert Wood Johnson Foundation, and academic consortia.
Data collection relied on manual and automated scraping of official sources including state health department websites, territorial dashboards, and municipal reports such as those from New York City Department of Health and Mental Hygiene, Los Angeles County Department of Public Health, and Chicago Department of Public Health. Methodological decisions were informed by epidemiologists from Johns Hopkins Bloomberg School of Public Health, Harvard T.H. Chan School of Public Health, and analytic teams at CDC Foundation and National Institutes of Health. The project categorized metrics including cases, hospitalizations, tests, test positivity, and deaths, reconciling divergent definitions from jurisdictions like Florida Department of Health, Texas Department of State Health Services, and California Department of Public Health. Version control and transparency used tools and platforms such as GitHub, Python (programming language), R (programming language), and continuous integration from contributors affiliated with MIT, Stanford University, Princeton University, and UC Berkeley.
Published outputs included daily time-series datasets, state-level summaries, archival snapshots, and interactive visualizations that informed reporting by outlets including CNN, NBC News, BBC News, Reuters, and Associated Press. Technical artifacts comprised CSV files, JSON feeds, API endpoints, and code repositories used by developers at Tableau Software, Microsoft, Google, and research teams at Carnegie Mellon University, University of Washington, Imperial College London, and University of Oxford. Tools for analysis exploited statistical packages from CRAN, machine-learning frameworks like TensorFlow and scikit-learn, and mapping via Leaflet (JavaScript library) and D3.js for interactive charts.
The datasets underpinned modeling and forecasting by groups including IHME, CDC COVID-19 Response Team, The COVID Analysis and Modeling Center, and independent modelers at Los Alamos National Laboratory and Harvard Global Health Institute. Journalists used the data in reporting across The New Yorker, Bloomberg, Politico, Vox, and The Guardian. Policymakers and public officials in jurisdictions such as New York (state), California, Illinois, Washington (state), and Massachusetts referenced the project’s compilations in briefings and press conferences, and academic studies in journals like The Lancet, JAMA, and Nature cited its datasets in analyses of testing, hospitalization, and racial disparities linked to COVID-19 outbreaks.
Though volunteer-driven, the project coordinated with organizational partners including The Atlantic, Knight Foundation, Mozilla Foundation, and academic partners at Columbia University Mailman School of Public Health. Financial and logistical support came via grants, in-kind contributions, and platform services from technology partners including GitHub, Google Cloud Platform, and Amazon Web Services. Collaboration extended to non-profits and data organizations such as Civic Hall, Sunlight Foundation, Open Knowledge Foundation, Data for Democracy, and research consortia at NYU Langone Health and Mount Sinai Health System.
Critics noted limitations common to second-party compilations: dependency on inconsistent reporting by agencies such as Ohio Department of Health and Georgia Department of Public Health, delays in death certification from county coroners, and divergent definitions of tests and hospitalizations across jurisdictions like Puerto Rico Department of Health and Alaska Department of Health and Social Services. Methodological debates involved academic groups at Columbia University, Yale School of Public Health, and think tanks like Brookings Institution about the handling of probable cases, data smoothing, and backlog reconciliation. The project acknowledged gaps in demographic completeness, echoing concerns raised by civil-rights advocates and public-health researchers at NAACP, KFF, and Urban Institute.
After winding down active collection in March 2021, archives of datasets, code, and documentation were preserved in public repositories on GitHub and institutional archives used by Library of Congress and university libraries including collections at Harvard Library and Columbia University Libraries. The project’s artifacts continue to inform retrospective analyses by historians, epidemiologists at Johns Hopkins University, data scientists at Argonne National Laboratory, and policy researchers at RAND Corporation. Its approach influenced later data standardization efforts by federal entities like Centers for Disease Control and Prevention and inspired civic-data initiatives in other countries modeled on the project's open-data practices.
Category:COVID-19 pandemic in the United States Category:Open data Category:Citizen science