This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| GDELT Project | |
|---|---|
| Name | GDELT Project |
| Type | Open data initiative |
| Founder | Kalev Leetaru |
| Established | 2013 |
| Headquarters | United States |
| Focus | Global event data, media analysis |
GDELT Project The GDELT Project is an open-data initiative that monitors worldwide broadcast, print, and web news in multiple languages to produce structured event, sentiment, and network data. It aggregates sources across continents and updates at high frequency to support research in political science, journalism, and international relations. The project is used by academic institutions, technology companies, and media organizations for large-scale empirical analysis of international events.
GDELT ingests mass media content from sources such as The New York Times, BBC News, Al Jazeera, Reuters, The Guardian and international wire services to convert unstructured text into coded records. It applies automated coding akin to projects like Integrated Crisis Early Warning System, ICEWS and datasets such as Correlates of War and ICEWS to map actors and actions across space and time. Researchers from universities including Harvard University, Stanford University, Massachusetts Institute of Technology, Oxford University and Columbia University have used its outputs alongside tools from Google, Microsoft, Amazon Web Services, IBM and Hewlett Packard Enterprise for computational social science and digital humanities.
The pipeline collects feeds from aggregators and publishers including Associated Press, Agence France-Presse, The Washington Post, Le Monde, Der Spiegel, and regional outlets covering areas such as Middle East, Sub-Saharan Africa, East Asia, South Asia and Latin America. It leverages natural language processing methods related to work by teams at Google Research, Facebook AI Research, OpenAI, Stanford NLP Group and Carnegie Mellon University to perform named-entity recognition, geocoding and event extraction. Processing stages echo methodologies from computational projects like Event Registry, Media Cloud and ICEWS and use gazetteers referencing GeoNames, Wikidata and United Nations place lists to assign coordinates and country codes. The system outputs coded event records, tone metrics, and network ties that mirror frameworks from CAMEO coding schemes and diplomatic datasets maintained by organizations such as World Bank, International Monetary Fund and United Nations Educational, Scientific and Cultural Organization.
GDELT produces multiple public files including event datasets, mentions files, and graph files in tabular and compressed formats designed for large-scale analysis. Formats are compatible with tools used by researchers at National Institutes of Health, European Commission, RAND Corporation and Brookings Institution, and interoperable with statistical environments like R (programming language), Python (programming language), MATLAB, SAS and Stata. The datasets contain fields for actors linked to identifiers similar to entries in Wikidata, attributes paralleling CAMEO event codes, geolocation data referencing GeoNames, and temporal stamps used in time-series research by centers such as Pew Research Center and Annenberg Public Policy Center.
Users analyze GDELT outputs with cloud services from Google Cloud Platform, Amazon Web Services, Microsoft Azure and with big-data frameworks like Apache Spark, Hadoop, Dask and Apache Flink. Visualization and network analysis often employ software from Tableau (software), Gephi, Pajek, Cytoscape and libraries such as NetworkX, D3.js and matplotlib. Academic collaborations integrate GDELT with platforms developed at Stanford University, MIT Media Lab, University of California, Berkeley and Princeton University for projects in computational geopolitics, conflict forecasting, and media studies.
GDELT datasets have been applied in studies of political instability, election monitoring, humanitarian response, and media framing by scholars at Yale University, University of Chicago, London School of Economics, University of Toronto and National University of Singapore. Public health researchers at Johns Hopkins University and Centers for Disease Control and Prevention have coupled signals with epidemiological data from World Health Organization and Pan American Health Organization for outbreak situational awareness. Nonprofits such as International Committee of the Red Cross and Human Rights Watch have used media-derived event data for crisis mapping, while financial analysts at firms like Goldman Sachs and Citigroup have experimented with sentiment indicators for market intelligence alongside datasets from Bloomberg and Thomson Reuters.
Critiques have focused on coverage bias, source selection transparency, event-coding accuracy, and challenges noted by researchers at MIT, University of Oxford, London School of Economics and Princeton University. Comparisons with hand-coded resources like Correlates of War and adjudicated event collections from Uppsala Conflict Data Program highlight issues in false positives, geolocation ambiguity and media reporting skew in regions such as Central Africa and Middle East. Methodological debates reference standards used by American Political Science Association members and point to reproducibility concerns raised in venues like NeurIPS and ACL (conference).
The project was developed beginning in the early 2010s by teams led by Kalev Leetaru and collaborators who engaged with communities at venues including AAAI, ICWSM, International Communication Association, Association for Computational Linguistics and International Studies Association. It evolved through stages influenced by advances from groups at Google, Microsoft Research, Stanford University and Carnegie Mellon University and has been cited in work alongside datasets from ICEWS, Correlates of War, Uppsala Conflict Data Program and initiatives at Harvard Kennedy School.
Category:Data projects Category:Computational social science