The University of Southampton
University of Southampton Institutional Repository

Provenance information in a collaborative knowledge graph: an evaluation of Wikidata external references

Piscopo, Alessandro, Kaffee, Lucie-Aimee, Frimelle, Phethean, Christopher and Simperl, Elena (2017) Provenance information in a collaborative knowledge graph: an evaluation of Wikidata external references In The Semantic Web – ISWC 2017: 16th International Semantic Web Conference, Vienna, Austria, October 23–25, 2017.

Record type: Conference or Workshop Item (Paper)


Wikidata is a collaboratively-edited knowledge graph; it expresses knowledge in the form of subject-property-value triples, which can be enhanced with references to add provenance information. Understanding the quality of Wikidata is key to its widespread adoption as a knowledge resource. We analyse one aspect of Wikidata quality, provenance, in terms of relevance and authoritativeness of its external references. We follow a two-staged approach. First, we perform a crowdsourced evaluation of references. Second, we use the judgements collected in the first stage to train a machine learning model to predict reference quality on a large-scale. The features chosen for the models were related to reference editing and the semantics of the triples they referred to. 61% of the references evaluated were relevant and authoritative. Bad references were often links that changed and either stopped working or pointed to other pages. The machine learning models outperformed the baseline and were able to accurately predict non-relevant and non-authoritative references. Further work should focus on implementing our approach in Wikidata to help editors find bad references.

PDF WD_sources_iswc(7) - Version of Record
Restricted to Repository staff only until 31 October 2017.
Download (387kB)

More information

Accepted/In Press date: 14 July 2017
Venue - Dates: The 16th International Semantic Web Conference, Vienna, Austria, 2017-10-23 - 2017-10-25


Local EPrints ID: 412923
PURE UUID: 9bb52375-ee81-4c07-8b03-579381bcd2e2
ORCID for Alessandro Piscopo: ORCID iD
ORCID for Lucie-Aimee, Frimelle Kaffee: ORCID iD
ORCID for Christopher Phethean: ORCID iD
ORCID for Elena Simperl: ORCID iD

Catalogue record

Date deposited: 08 Aug 2017 16:31
Last modified: 08 Aug 2017 16:31

Export record


Author: Alessandro Piscopo ORCID iD
Author: Lucie-Aimee, Frimelle Kaffee ORCID iD
Author: Elena Simperl ORCID iD

University divisions

Download statistics

Downloads from ePrints over the past year. Other digital versions may also be available to download e.g. from the publisher's website.

View more statistics

Atom RSS 1.0 RSS 2.0

Contact ePrints Soton:

ePrints Soton supports OAI 2.0 with a base URL of

This repository has been built using EPrints software, developed at the University of Southampton, but available to everyone to use.

We use cookies to ensure that we give you the best experience on our website. If you continue without changing your settings, we will assume that you are happy to receive cookies on the University of Southampton website.