Skip to content
RAG Repo

OBELICS is a curated collection of 141M English web documents where text and images appear together in the order they had on the original page. Most image-text datasets reduce each image to a single caption, but OBELICS keeps the full document, so a model sees images in the context of the paragraphs around them.

In total it holds 353M images and 115 billion tokens of text. This interleaved structure is what made it suitable for training the IDEFICS family of multimodal models, which read documents rather than isolated image-caption pairs.

The dataset is openly accessible on the Hugging Face Hub under CC BY 4.0.

image-textinterleavedweb-scalemultimodalenglish

Related sources