Wikidata and Wikibase as a hub for identifiers

From DHWiki

Wikidata and Wikibase as a Hub for Identifiers

The world's cultural and academic knowledge is stored in a multitude of separate systems that have shown limited interconnection for years, until the Semantic Web started to slowly shape form.

In such complex ecosystem, Wikidata (with its underlying software, Wikibase) has become an indispensable component of the global Semantic Web infrastructure, thanks to the amount of entities it describes (as of October 2025, more than 119 M[1]) and the number of external resources to which it is connected (as of October 2025, its properties with datatype external-id are more than 9 thousands[2]).

This chapter examines the role of Wikidata and Wikibase instances in the landscape of Semantic Web: a model that bridges existing resources while leveraging a new generation of dynamic, collaborative databases of authority files, controlled vocabularies, and GLAM data repositories.

Authority files

Wikidata and authority files

GLAM institutions, most notably libraries, maintain catalogues of their collections in order to make them accessible by the public. Each resource can be found through many keywords, including the title, the author(s), etc. However, the names of these entities are often ambiguous: an author can have multiple names (e.g., due to different transliterations) and/or homonym authors can exist. In order to ensure an effective indexing of the catalogued resources, the keywords used most frequently by the users to find them are treated as controlled access points. The process of creating controlled access points, known as authority control, allows users to find all and only the resources they are looking for. Each controlled access point (e.g., each author) has an entry, known as authority record, containing basic data and linked to all and only the pertinent resources. Authority records play a crucial role in disambiguating homonym entities and in collecting the multiple names of each entity, in order to ensure consistency and accuracy in indexing resources. The entire collection of the authority records curated by an institution, or group of institutions, is known as authority file.

Since 2013, Wikidata items link, among various types of databases, to authority files; the second and the third external-id properties created in Wikidata were the ones for ISNI (P213) and VIAF (P214). As of October 2025, Wikidata has 225 external-id properties for library authority files,[3] used as value in 23.2 M statements[4] contained in 5.8 M unique items,[5] the majority of them (4.2 M) being personal items (i.e. having P31 = Q5, "instance of = human").[6] Among these 225 properties, the 5 most used by number of unique items are VIAF (3.8 M items), GND (2.4 M items), ISNI (2.2 M items), LC (1.7 M items) and IDREF (1.0 M items).[7]

Wikidata persistent identifiers (PIDs) have become de facto a central pillar in the landscape of authority files (Van Veen 2019).

The workflows regarding the use of Wikidata in authority control, in the framework of a cooperation between cataloguers and Wikidata editors, have been illustrated in many presentations (most recently, Pellizzari di San Girolamo 2025). These workflows can be divided into two parts:

  1. reconciliation;
  2. cooperation.

Reconciliation

Authority records can always link to Wikidata, in the 017 field of UNIMARC and 024 field of MARC 21.

Conversely, Wikidata can link to authority records only if they have a URI. If an authority file does not provide URIs for its authority records, it cannot be linked from Wikidata, but it can link to Wikidata.

The reconciliation between authority records and Wikidata items can be performed in three ways: manually; importing a previous list of matches (e.g., into Wikidata through QuickStatements[8] or QuickStatements 3.0[9]); with the help of a reconciliation tool. The two main reconciliation tools used in Wikidata are Mix'n'match[10] and OpenRefine.[11] There are some relevant differences between them: Mix'n'match is a web tool, so it can be used by multiple users at the same time, and it allows to create new items putting together entries from different catalogues; conversely, OpenRefine is a software running on an individual computer, so each reconciliation project is developed by one individual user, and contains functions allowing to clean messy data. Mix'n'match, in general, is more suitable for projects which are meant to be completed in a long span of time and require collaboration, and when most of the entries in the dataset to be reconciled with Wikidata are already present in Wikidata; conversely, OpenRefine is more suitable for projects managed by one person, and when most of the entries need to be imported as new items into Wikidata. Only datasets licensed under CC0 can be massively imported into Wikidata.

Cooperation

The main areas of cooperation include:

  1. editing individual authority records and Wikidata items;
  2. comparing and editing both authority records and Wikidata items;
  3. editing massively authority records and Wikidata items.

Regarding point 1: everyone can edit Wikidata items, but only cataloguers can usually edit authority records, so if a Wikidata editor finds in Wikidata a mistake imported from an authority record, there should be some mistake-reporting procedure, so that, once the mistakes are corrected in Wikidata, the same also happens in the authority record - the so-called "data round-tripping" (cf. also CHAPTER).[12] The mistake-reporting email or webform can be indicated directly in the Wikidata property of the authority file through P10923. On the other side, cataloguers can (as Wikidata editors) improve Wikidata items linking to their authority records; this ensures that the authors they care about are well described in Wikidata, a source that can be used by other librarians to find data for their authority records, and by the library itself to display more data for the catalogue's users (e.g., links to other resources). Finding Wikidata items with missing and/or unreferenced data can be easily done through queries.

Regarding point 2: the correspondence between Wikidata items and authority records is usually biunivocal (1:1), although exceptions are possible, particularly in the case of pseudonyms, whose treatment differs between Wikidata and the various authority files (cf. Chen, Yuxuan 2024, Chou 2025). Consequently, most of the cases of non-biunivocal correspondence, i.e. unique-value constraint violations (2 or more Wikidata items to 1 authority record) and single-value constraint violations (1 Wikidata items to 2 or more authority records) are clues of mistakes. These mistakes can be duplications or conflations, in Wikidata (cf. Pellizzari di San Girolamo 2024a) or in the authority file. In order to empty the lists of constraint violations, fix the issues of one side is insufficient, because the other issues remain cluttering the list; only a conjunct effort can produce an effective workflow, allowing to improve both Wikidata and the authority file. If the authority file has a SPARQL endpoint, making federated queries with Wikidata allows to easily spot inconsistencies (e.g., in gender, dates) and work on them (Kerboul 2025).[13]

Regarding point 3: Wikidata can be used as a source to massively enrich authority records; this can be done in two ways, retrieving the data on-the-fly through the API and showing it in an infobox, or extracting the data periodically and copying it into the authority records (especially the empty or nearly-empty ones). Both these ways have pros and cons: the second gives more control on the data showed, whilst the first has the advantage of being always up-to-date. Examples of the first include the AuthorityBox in the ILS Koha (cf. Bargioni 2020, Q809) and the use of Wikidata inside the ILS Primo,[14] whilst the main example of the second is the enrichment of SBN authority records with data from Wikidata (cf. Ravelli 2024). On the other side, authority records can be used to improve Wikidata through massive imports (if licensed under CC0), adding new high quality items and raising the quality of the existing ones; the main project of data import from an authority file, currently continuing on a regular basis through a bot, is the one managed by the National Library of the Czech Republic (cf. Dostál 2021, Jansová, Maixnerová, Šťastná 2024, Dostál 2025).

Stable cooperation between the Wikidata community and authority files involve, besides the National Library of the Czech Republic, the Italian National Library Service (SBN)[15] and the Union of Roman Ecclesiastical Libraries (URBE).[16]

Wikibase and authority files

Wikibase is currently used to maintain two authority files, the Semantic Name Authority Repository Cymru (SNARC) by the National Library of Wales,[17] started in 2022, and the National Library of Nigeria Semantic Name Authority Repository (NLN SNAR),[18] started in 2025 (cf. Osuigwe, Boakye-Achampong, Appiah 2025). Wikibase is also being experimented as a software for the authority file of the National Library of Greece; a proof of concept was deemed successful in 2024 (cf. Zapounidou et al. 2024) and, as of 2025, the project is progressing (cf. Gerolimos et al. 2025).

Controlled vocabularies

Wikidata and controlled vocabularies

Wikidata can be used as a hub to interlink controlled vocabularies, both general-purpose ones and thematic ones. The main ongoing project regards the Nuovo Soggettario thesaurus, developed by the National Central Library of Florence, Italy:[19] the two focuses of the project have been fixing the previously accumulated non-biunivocal links, also through federated SPARQL queries and improving the interconnection on a specific topic area, photography, which proved successful under many aspects (cf. Cencetti, Pellizzari di San Girolamo, Viti 2025, Viti, Pellizzari di San Girolamo 2025).

Wikibase and controlled vocabularies

In 2025 the National Library of the Czech Republic exported one of its controlled vocabularies, TDKIV (the Czech Explanatory Terminology Database of Library and Information Science), into a Wikibase instance,[20] with the aim of discarding the previous platform and continuing the development of the thesaurus only on Wikibase.

GLAM data

GLAM institutions have been publishing digital collections during the last decades including a wide diversity of content such as images, paintings, postcards, metadata, text and videos. Cultural Heritage institutions have been exploring new ways to make the content available in the form of reusable data (Tasovac et al, 2020). Recent trends such as Collections as data,[21] or networks such as Research Data Alliance[22] and the International GLAM Labs Community[23] promote the publication of digital collections able to support computational use. Relevant examples of institutions making their digital collections available in the form of reusable data are the National Library of Scotland[24] and the National Library of Luxembourg[25]. In a recent collaborative effort, best practices for digital collections supporting computational use and including collaborative edition platforms such as Wikidata and Wikibase have been made available (Candela et al, 2023, Candela et al, 2023b). Collaborative approaches in the field of Digital Humanities and beyond include the Social Sciences and Humanities Open Marketplace,[26] a platform for publishing workflows and tool descriptions. Other initiatives are focused on the adoption of new techniques based on Artificial Intelligence and the provision of training materials for undertaking digital scholarship and data science in GLAM.[27][28]

Integration of initiatives - dhwiki

Wikidata and GLAM data

GLAM (Galleries, Libraries, Archives and Museums) have been exploring new ways to publish and enrich their data. In this context, Wikidata has played a leading role as a collaborative edition approach to connect and enrich their resources (Candela, 2024). In Digital Humanities, Wikidata is used in data quality assessment, enrichment, and visualisation (Zhao, 2023; Candela, 2023; Candela, 2025).

Wikidata enables the creation of dedicated external id properties to link to identifiers of artists, cultural heritage objects, etc. In general, Wikidata has been used to connect people, works, locations, and subjects. The following table shows a list of properties to link resources based on relevant institutions.

Examples of properties in Wikidata to link resources in GLAM institutions
Property Description
P5361 British National Bibliography person ID
P268 Bibliothèque nationale de France ID
P950 Biblioteca Nacional de España ID
P8565 British Museum object ID
P227 German National Library
P244 Library of Congress authority ID
P10373 Identificador de Mnemosine: Biblioteca Digital de la otra Edad de Plata
P5321 Museo del Prado Artist ID

Some of the benefits of using Wikidata are: i) enabling the community to have a relevant role to curate digital collections; ii) Wikidata functions as an identifier hub, avoiding the creation the data silos and promoting the connection of repositories (Dişli, 2025); iii) Wikidata content is being used in applications involving AI (see, for example, Kunpeng et al, (2023) or the Wikidata Embedding Project[29]; and iv) Wikidata enables the creation of visualisations such as maps and timelines.

Despite all these efforts, there is still room for improvement in terms of coverage. For instance, and considering the completeness dimension, not all records in each collection are linked to Wikidata. In addition, most of the properties are dedicated to Global North institutions, reducing the visibility of relevant institutions in different parts of the world. Additional work is required in this sense to cover underrepresented languages and cultures (Carroll, 2020).

#title: External-id properties related to Museums (from Wikidata)
# https://qlever.dev/wikidata/IT9Tyc gives usage counts for each property
#
PREFIX wikibase: <http://wikiba.se/ontology#>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
PREFIX bd: <http://www.bigdata.com/rdf#>

SELECT ?museum ?museumLabel ?countryLabel ?property ?propertyLabel WHERE {
  SERVICE <https://query.wikidata.org/sparql> {
    SELECT DISTINCT ?museum ?museumLabel ?countryLabel ?property ?propertyLabel WHERE {
      ?museum (wdt:P31/(wdt:P279*)) wd:Q33506;
        wdt:P17 ?country;
        wdt:P1687 ?property.
      ?property wikibase:propertyType wikibase:ExternalId.
      SERVICE wikibase:label { bd:serviceParam wikibase:language "[AUTO_LANGUAGE],mul,en". }
    }
  }
}

Try it!

Other CH and GLAM experiments are based on the extraction of data from Wikidata according to particular topics such as the Spanish Civil War and Exile (Candela et al., 2025b) and the identification of objects that were lost or stolen during the Second World War using provenance documentation from sale catalogues and archival inventories such as the Getty Provenance Index.[30]

These examples provide a wide diversity of experimental approaches, all of them using Wikidata to some extent for different purposes. As described, Wikidata is a powerful and valuable resource for GLAM institutions.

Wikibase and GLAM data

Recently, Wikibase has been adopted in a range of use cases that involve GLAM datasets. The following incomplete list points to some examples:

Resources

Authority files

Wikidata pages
Articles
Presentations

Controlled vocabularies

Wikidata pages
Articles
Presentations

Publications related to GLAM data

Use the following query to retrieve bibliographical items indexed with tag "GLAM".

#title: All bibliographical items for indexed domain label "GLAM"
PREFIX dhwb: <https://dhwiki.wikibase.cloud/entity/>
PREFIX dhdp: <https://dhwiki.wikibase.cloud/prop/direct/>
PREFIX dhp: <https://dhwiki.wikibase.cloud/prop/>
PREFIX dhpq: <https://dhwiki.wikibase.cloud/prop/qualifier/>

select ?item ?typeLabel ?date ?itemLabel ?langLabel (group_concat(distinct ?authorname;SEPARATOR=", ") as ?authors) (iri(concat("https://www.zotero.org/groups/5639268/dariah_wg_dhwiki/items/",?zot,"/item-details")) as ?zotero) (iri(concat(str(wd:),?wd)) as ?wikidata)
       
where { 
  ?item dhdp:P5 dhwb:Q2; dhdp:P8 ?type; dhdp:P41 dhwb:Q850; # Q850 GLAM
        dhdp:P26 ?date; dhdp:P11 ?lang; dhdp:P6 ?zot; dhp:P29 [dhpq:P10 ?authorname].
  optional {?item dhdp:P1 ?wd.}
  
  SERVICE wikibase:label { bd:serviceParam wikibase:language "[AUTO_LANGUAGE],en". }
} group by ?item ?typeLabel ?date ?itemLabel ?langLabel ?authors ?zot ?wd
order by desc(year(?date))

Try it!


References

  1. See https://www.wikidata.org/wiki/Special:Statistics.
  2. See https://www.wikidata.org/wiki/Help:Data_type.
  3. See https://qlever.dev/wikidata/LI6Z15.
  4. See https://qlever.dev/wikidata/0hgYoM.
  5. See https://qlever.dev/wikidata/PuO0It.
  6. See https://qlever.dev/wikidata/8OgPd2.
  7. See https://qlever.dev/wikidata/XxlFHJ.
  8. See https://quickstatements.toolforge.org/#/batch.
  9. See https://qs-dev.toolforge.org/.
  10. See https://mix-n-match.toolforge.org/.
  11. See https://openrefine.org/.
  12. Cf. https://www.wikidata.org/wiki/Wikidata:Data_round-tripping.
  13. Cf. e.g. https://www.wikidata.org/wiki/Property_talk:P269/Gender_mismatches.
  14. See https://knowledge.exlibrisgroup.com/Primo/Product_Documentation/020Primo_VE/Primo_VE_(English)/150End_User_Help/Searching_Linked_Open_Data_-_Person_Entity.
  15. See https://www.wikidata.org/wiki/Wikidata:Gruppo_Wikidata_per_Musei,_Archivi_e_Biblioteche/SBN.
  16. See https://www.wikidata.org/wiki/Wikidata:Gruppo_Wikidata_per_Musei,_Archivi_e_Biblioteche/Parsifal.
  17. See https://snarc-llgc.wikibase.cloud/.
  18. See https://nlnsnar.wikibase.cloud/.
  19. Cf. https://www.wikidata.org/wiki/Wikidata:Gruppo_Wikidata_per_Musei,_Archivi_e_Biblioteche/Nuovo_soggettario.
  20. See https://tdkiv.wikibase.cloud/.
  21. See https://collectionsasdata.github.io/.
  22. See https://www.rd-alliance.org/groups/collections-as-data-ig.
  23. See https://www.glamlabs.io/.
  24. See https://data.nls.uk.
  25. See https://data.bnl.lu.
  26. See https://marketplace.sshopencloud.eu/.
  27. See https://libereurope.github.io/ds-topic-guides/.
  28. See https://sites.google.com/view/ai4lam.
  29. See https://www.wikidata.org/wiki/Wikidata:Embedding_Project
  30. See https://www.getty.edu/research/provenance/.
  31. See https://cidoc-crm.org/.