You might find this meets many needs: https://query.wikidata.org/querybuilder/ e...

LeonardoTolstoy · on March 25, 2023

I use wikidata a lot for movie stuff. Ideally I imagine the wiki foundation itself will be looking into using LLMs to help parse their own data and convert it into wikidata content (or confirm it, or keep it up to date, etc.)

Wikidata is incredibly useful for things that I would considered valuable (e.g. the tMDb link for a movie) but due to the curation imposed upon Wikipedia itself isn't typically available for very many pages. An LLM won't help with that but another bit of information like where films are set would be a perfect candidate for an LLM to try and determine and fill in automatically with a flag for manual confirmation.

worldsayshi · on March 25, 2023

To be fair, I was quite confused by wikidata query notation when I tried it as well.

rjh29 · on March 26, 2023

I used that when building a database of Japanese names, but found that even wikidata is inconsistent in the format/structure of its data, as it's contributed by a variety of automated and human sources!

riku_iki · on March 25, 2023

its wikidata, not wikipedia, they are two disjoint datasets.

ZeroGravitas · on March 25, 2023

Basically every wikipedia page (across languages) is linked to wikidata, and some infoboxes are generated directly from wikidata, so they're seperate, but overlapping and increasingly so.

https://en.wikipedia.org/wiki/Category:Articles_with_infobox...

edit: slightly wider scope category pointing to pages using wikidata in different ways:

https://en.wikipedia.org/wiki/Category:Wikipedia_categories_...

riku_iki · on March 25, 2023

I agree there is strong overlap between entities, and also infobox values, but both wikidata and wikipedia has many more disjoint datapoints: many tables, factual statements in wikipedia which are not in wikidata, and many statements in wikidata which are not in wikipedia.