This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| The Guardian Datablog | |
|---|---|
| Name | The Guardian Datablog |
| Type | Data journalism blog |
| Owner | The Guardian |
| Launch | 2009 |
| Language | English |
| Country | United Kingdom |
The Guardian Datablog was a data journalism initiative of The Guardian that combined structured datasets, statistical analysis and visualisation to accompany reporting on politics, health, crime and public policy. It blended newsroom journalism with open data practices, publishing datasets, code and interactive charts to support investigations into subjects ranging from elections and budgets to climate change and public spending. The project intersected with developments in data journalism, open-source software communities and transparency campaigns led by organisations and figures across media and civic technology.
The Datablog emerged amid a broader surge in digital news innovation that included projects by ProPublica, FiveThirtyEight, The New York Times, The Washington Post and Al Jazeera; it launched as part of Guardian News & Media's digital expansion under editors influenced by figures such as Alan Rusbridger and managers within Guardian Media Group. Early work drew on collaborations with advocates from Open Knowledge Foundation, MySociety, Mozilla Foundation, Sunlight Foundation, and researchers at universities like University of Oxford, London School of Economics, University College London and University of Cambridge. The blog's timeline intersected with major events including the 2009 United Kingdom parliamentary expenses scandal, the 2010 United Kingdom general election, the 2014 Scottish independence referendum, and the 2016 United Kingdom European Union membership referendum, shaping its editorial priorities and dataset releases.
Editors and reporters on the Datablog emphasised reproducibility, licensing, and sourcing, adopting practices from communities around Creative Commons, OpenStreetMap, Wikidata, and standards developed at institutions such as International Organization for Standardization and research labs at Massachusetts Institute of Technology and Stanford University. Journalists worked with data scientists versed in tools from R (programming language), Python (programming language), D3.js, and libraries associated with projects at Google's research groups and the European Organization for Nuclear Research. The editorial framework referenced methods used in investigative pieces by outlets like BBC News, Channel 4, Der Spiegel, Le Monde, and El País, and aligned with open records litigation strategies used by litigants in cases before courts such as the High Court of Justice (England and Wales) and the European Court of Human Rights.
The Datablog published datasets and investigations on topics that attracted attention across political and civic spheres, often linking to primary sources such as releases from UK Treasury, Office for National Statistics, NHS England, Met Office, HM Revenue and Customs, and international agencies like World Bank, International Monetary Fund, World Health Organization, and United Nations. High-profile datasets covered election results with ties to Electoral Commission (United Kingdom), public spending and the Parliamentary expenses scandal, regional housing datasets linked to reports by Department for Communities and Local Government, crime statistics from Home Office (United Kingdom), and environmental data referencing Intergovernmental Panel on Climate Change and programmes by United Nations Environment Programme. The blog collaborated on investigative projects resonant with work by Richard Stallman-adjacent open-data advocates, journalists such as Burt Herman-style innovators, and research partnerships with groups like Centre for Data Ethics and Innovation.
The Datablog influenced how newsrooms incorporated data-driven reporting, contributing to the professional practices in organisations including The Times, Financial Times, Bloomberg, Reuters, Politico, and nonprofit outlets such as The Bureau of Investigative Journalism and OpenCorporates. Academics at Harvard University, Columbia University, University of Pennsylvania, and University of California, Berkeley cited Datablog pieces in studies of media transparency and computational journalism. Its work fed public debates involving politicians like David Cameron, Theresa May, Boris Johnson, Jeremy Corbyn, and institutions such as House of Commons. Critics and commentators in forums including Nieman Lab, Tow Center for Digital Journalism, and at conferences like International Journalism Festival assessed the blog's role in shaping expectations for data availability and civic accountability.
The Datablog used and promoted tools from ecosystems including GitHub, GitLab, Google Public Data Explorer, and mapping platforms like Carto and Mapbox. Analysts employed statistical environments and libraries linked to projects at RStudio, scikit-learn, TensorFlow experiments from Google Brain, and visualisation frameworks inspired by work at Observablehq and academic labs at MIT Media Lab. Data release practices referenced metadata standards developed by W3C and dataset cataloguing systems similar to those at data.gov.uk and regional portals such as data.gov (United States). The blog contributed to open-source codebases and reusable workflows shared with organisations like DataKind, DrivenData, and training programmes at Knight Foundation-supported initiatives.
Reporting and tools produced by the Datablog received recognition alongside awards won by The Guardian in competitions such as the British Journalism Awards, European Press Prize, Society for News Design, and honors acknowledged by organisations including Association for Computing Machinery conferences, Online News Association, and the Data Journalism Awards. Individual contributors were shortlisted for prizes administered by institutions like Royal Statistical Society, British Academy, and media fellowships at Reuters Institute for the Study of Journalism and Nieman Foundation.
Category:Data journalism Category:The Guardian