This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| ProPublica Data Store | |
|---|---|
| Name | ProPublica Data Store |
| Type | Online data repository |
| Owner | ProPublica |
| Launched | 2013 |
| Country | United States |
ProPublica Data Store
The ProPublica Data Store is an online repository and distribution platform for datasets, investigations, and supporting materials produced by the nonprofit investigative newsroom ProPublica. It aggregates data underlying investigations, supplements reporting with structured files, and distributes datasets to journalists, researchers, and the public through downloadable formats and APIs. The Data Store supports collaboration between journalists, academic institutions, advocacy organizations, and civic technologists tied to major investigations and transparency projects.
The Data Store functions as an archival and distribution node for ProPublica's investigative work and related datasets, connecting output from newsroom projects such as the Nobel Prize-covered investigations, collaborations with the New York Times, the Washington Post, and partnerships echoing methods used by the Center for Investigative Reporting, International Consortium of Investigative Journalists, and the Sunlight Foundation. It catalogs datasets that intersect with entities like the Internal Revenue Service, Centers for Disease Control and Prevention, Federal Bureau of Investigation, Department of Justice (United States), and major institutions such as Harvard University, Stanford University, and Massachusetts Institute of Technology. The platform's scope includes data on public policy topics involving the Affordable Care Act, the Dodd–Frank Wall Street Reform and Consumer Protection Act, and the Patriot Act as they relate to specific investigations and datasets.
Launched in the early 2010s amid a broader shift toward data-driven reporting championed by outlets such as the Guardian (news organization), the Data Store was developed to institutionalize ProPublica’s approach to open data pioneered alongside initiatives by the BBC, the Associated Press, and the Los Angeles Times. Early datasets reflected collaborations with academic partners at Columbia University, University of California, Berkeley, and Princeton University and fed into investigations related to entities like the Social Security Administration, Environmental Protection Agency, and Food and Drug Administration. The platform evolved during a period marked by major events such as the 2010s United States debt-ceiling crisis and the 2016 United States presidential election, adapting to new standards from organizations like the Open Knowledge Foundation and the Data Journalism Handbook community.
Collections hosted in the Data Store have included datasets accompanying investigations into topics involving the Internal Revenue Service filings, Medicare billing patterns, Federal Aviation Administration enforcement actions, and corporate disclosures filed with the Securities and Exchange Commission. Notable projects paralleled reporting on the Katrina aftermath, analyses of Hurricane Katrina-related spending, examinations tied to the Enron scandal period regulatory reforms, and datasets used in cross-border investigations with the International Consortium of Investigative Journalists like the Panama Papers-era analyses. The Data Store has released material on campaign finance connected to the Federal Election Commission, lobbying records intersecting with the Lobbying Disclosure Act of 1995, and public health datasets referencing outbreaks cataloged by the Centers for Disease Control and Prevention and research from institutions such as the Johns Hopkins University and Yale University.
Datasets in the Data Store are typically distributed under open licenses aligned with norms used by the Open Knowledge Foundation, the Creative Commons framework, and similar open-data advocates such as Data.gov. Access modalities evolved with practices from projects by the Knight Foundation and standards promoted by the World Wide Web Consortium. Users drawn from institutions including Library of Congress, Brookings Institution, Pew Research Center, and university libraries can download files in formats compatible with tools used at MIT Media Lab and archives maintained by the National Archives and Records Administration.
The Data Store has amplified reporting by ProPublica and partners, influencing coverage in outlets such as the New York Times, Washington Post, Los Angeles Times, Chicago Tribune, and international partners like the BBC and The Guardian (UK). Researchers at Harvard Kennedy School, Columbia Journalism School, and think tanks such as the Annenberg Public Policy Center and Urban Institute have used datasets for peer-reviewed studies, policy briefs, and graduate theses. The datasets have informed congressional hearings involving committees such as the United States House Committee on Oversight and Reform and regulatory reviews by agencies including the Securities and Exchange Commission and Department of Health and Human Services.
Technical implementation drew on open-source components used by projects at GitHub, Apache Software Foundation-hosted projects, and mapping libraries from organizations like Mapbox and the OpenStreetMap community. The Data Store supports formats and tools familiar to practitioners at Google Research, Microsoft Research, and civic-technology groups such as Code for America; it integrates with analytics workflows relying on software from the Python Software Foundation, the R Project for Statistical Computing, and platforms like Jupyter Notebook and Tableau. Backend hosting and content delivery practices reflect patterns used by major repositories operated by Internet Archive and institutional data services at universities such as University of Michigan and University of California, Los Angeles.
Governance of the Data Store aligns with ProPublica’s editorial and nonprofit structures and interacts with funders and partners including philanthropic organizations like the John D. and Catherine T. MacArthur Foundation, the Ford Foundation, the Knight Foundation, and collaborators such as the Investigative Reporters and Editors (IRE), Open Society Foundations, and media partners including the Associated Press and NPR. Academic collaborations have included researchers from Yale Law School, Georgetown University, and University of Pennsylvania, while technology partnerships track with organizations such as Mozilla Foundation and the Linux Foundation.
Category:Data repositories Category:Investigative journalism