This article was accepted into the corpus but its outbound wikilinks were never NER-processed — typical at the deepest BFS hop or when the run's entity cap was reached. No expansion funnel to show.
| KONECT | |
|---|---|
| Name | KONECT |
| Type | Dataset Repository |
| Country | Netherlands |
| Established | 2010s |
| Domain | Network Science |
KONECT
KONECT is a collection of network datasets used for research in graph theory, network science, social network analysis, machine learning, and data mining. It aggregates diverse networks from domains such as social network, citation network, transportation network, communication network, and biological network, enabling comparative studies across datasets used by researchers affiliated with institutions like University of Szeged, Max Planck Society, MIT, Stanford University, and University of Oxford. The collection supports analyses that intersect work from labs such as the Santa Fe Institute, Lawrence Berkeley National Laboratory, Microsoft Research, and Google Research.
KONECT provides curated network datasets drawn from sources including online platforms like Twitter, Facebook, YouTube, Wikipedia, and Stack Overflow; bibliographic venues like arXiv, DBLP, and Web of Science; infrastructure projects like OpenStreetMap and European Space Agency maps; and biological repositories such as GenBank and Protein Data Bank. Researchers from Imperial College London, ETH Zurich, Princeton University, Columbia University, and Harvard University use KONECT data to evaluate algorithms from fields represented by groups at Bell Labs, Carnegie Mellon University, Georgia Institute of Technology, and University of California, Berkeley.
Datasets in KONECT originate from collaborations with organizations and from public exports of services operated by companies such as Amazon.com, eBay, Netflix, LinkedIn, Flickr, and Reddit. The collection includes temporal networks, static snapshots, directed networks, undirected networks, weighted graphs, bipartite graphs, and multiplex structures similar to datasets used in studies at Los Alamos National Laboratory, National Institutes of Health, European Organization for Nuclear Research, and NASA. Metadata captures provenance comparable to standards endorsed by Digital Object Identifier, Creative Commons, and domain repositories like Zenodo and Figshare.
KONECT distributes files in formats interoperable with tools developed at institutions such as The Apache Software Foundation, Python Software Foundation, R Consortium, and MATLAB. Available formats include edge lists, adjacency matrices, and sparse representations compatible with libraries like NetworkX, igraph, Graph-tool, TensorFlow, and PyTorch. Schema conventions align with serializations used by JSON, CSV, and HDF5 ecosystems and follow citation norms similar to ACM SIGMOD, IEEE, and SIAM publications.
Researchers apply algorithms for centrality, community detection, link prediction, and temporal analysis using implementations from projects like Gephi, Cytoscape, Neo4j, OrientDB, and frameworks developed at Facebook AI Research and OpenAI. Analytical methods benchmarked on KONECT datasets include spectral clustering related to work by von Luxburg, stochastic block models from research by Holland, Laskey, and Leinhardt, PageRank originating with Google founders Sergey Brin and Larry Page, and modularity measures influenced by studies at University of California, Santa Barbara. Visualization approaches draw on techniques used in D3.js demos and tools emerging from Visualization Research groups at Microsoft Research Redmond and Tableau Software.
KONECT datasets support applications in recommender systems like those studied in Netflix Prize, fraud detection methods used by Visa and Mastercard, epidemiological modeling informed by work at World Health Organization and Centers for Disease Control and Prevention, urban planning projects undertaken by UN-Habitat and City of New York, and biological network analysis in research by National Institutes of Health and European Bioinformatics Institute. The repository aids reproducibility campaigns advocated by organizations such as Open Science Framework and methods compared in competitions like Kaggle and NeurIPS challenges.
Access policies for datasets mirror licensing practices from sources such as Creative Commons Attribution, Creative Commons Zero, proprietary terms used by Twitter, Inc. and Facebook, Inc., and data-sharing agreements similar to those negotiated with Elsevier and Springer Nature. Distribution mechanisms leverage mirrors and archival strategies employed by Internet Archive and LOCKSS while promoting citation practices aligned with Digital Object Identifier assignments and archival norms in repositories like Zenodo.
KONECT emerged in the 2010s amid increased interest in empirical studies by groups at University of Pisa, Technical University of Munich, University of Cambridge, and Tel Aviv University. Its growth parallels milestones such as the release of large-scale datasets by Twitter, publication of influential works like "Networks" by Mark Newman, and methodological advances from conferences including KDD, ICML, NeurIPS, WWW Conference, and SIGMOD. Development has been influenced by community efforts exemplified by Open Data initiatives, reproducibility discussions at AAAS meetings, and dataset curations led by repositories like UCI Machine Learning Repository.
Category:Network datasets