Skip to content
RAG Repo

WIT (Wikipedia-based Image Text)

WIT draws on Wikipedia to connect images with the text around them. For each image it captures the caption, the reference description, and passages from the article the image sits in, so a model gets far more context than a one-line caption provides.

Its defining feature is language coverage. WIT spans well over 100 languages, which makes it one of the few large image-text datasets suitable for training and evaluating multilingual multimodal systems rather than English-only ones.

The dataset is released by Google Research under CC BY-SA 4.0, a share-alike licence, meaning any derivative database you distribute has to carry the same terms. The underlying images come from Wikimedia Commons under their own varied licences, so the picture licences can differ from the dataset licence.

image-textmultilingualwikipediamultimodalshare-alike

Related sources