Where this comes from

Sources

Open Archives is built entirely from public archives that allow free reuse. Nothing here is scraped from behind a paywall, and every image we keep carries a public-domain or Creative Commons licence recorded next to it. Below is every collection we draw on, what it gives us, and how it is licensed.

In use today

These archives supply the cards, images and metadata you see across the site right now.

Wikipedia

In use

6.9M+ English articles

Card titles, summaries, dates and the links that connect cards to one another.

Text under CC BY-SA 4.0, credited back to Wikipedia on every card.

Visit Wikipedia

Wikidata

In use

115M+ structured items

Classification — is a card a person, place, thing or event — plus countries, roles, genres and tags.

CC0. Free for any use.

Visit Wikidata

Wikimedia Commons

In use

100M+ media files

Nearly every card image on the site. We only accept freely licensed files and drop fair-use ones.

Public domain or Creative Commons, with artist and license recorded per image.

Visit Wikimedia Commons

Smithsonian Open Access

In use

5M+ open items

Museum objects, specimens and photographs with dates, places and institutional credit — searchable under Databases.

CC0. We import only records the Smithsonian marks CC0.

Visit Smithsonian Open Access

Notion research bases

In use

Curated lists

Hand-built subject lists (for example the insect species master list) used to seed themed pages.

Used as a checklist only; the content itself still comes from the open archives above.

Visit Notion research bases

Google News RSS

In use

Continuous feed

The latest news headlines shown alongside cards, collections and web pages.

Headlines and links only, always pointing back to the publisher.

Visit Google News RSS

Next up

Large public archives with real APIs and clear licensing that we are building importers for.

Library of Congress Digital Collections

Planned

Millions of items

Historic photographs, maps, prints and manuscripts with rich dates and places.

Mostly public domain; rights are stated per item.

Visit Library of Congress Digital Collections

Chronicling America

Planned

All 50 states, through 1963

Digitised American newspapers, full-text searchable.

Public domain.

Visit Chronicling America

Internet Archive

Planned

1 trillion+ web captures, millions of texts

Books, film, audio, software and the Wayback Machine.

Varies by item; we would take only clearly free material.

Visit Internet Archive

Europeana

Planned

60M+ items

Aggregated European galleries, libraries, archives and museums — 34M images, 24.5M texts.

Varies by contributing institution; filtered to open licences.

Visit Europeana

DPLA

Planned

53.5M+ items

The US equivalent of Europeana, pulling from state and regional hubs.

Metadata CC0; the object itself lives on the partner site.

Visit DPLA

Gallica (BnF)

Planned

Millions of items

The French national library's digitised books, press, maps and images.

Largely public domain with a documented API.

Visit Gallica (BnF)

Trove

Planned

National scale

Australia's national newspaper and heritage archive.

Free with an API key.

Visit Trove

National Archives Catalog (NARA)

Planned

Hundreds of millions of pages

US federal records, photographs and film.

Overwhelmingly public domain as US government work.

Visit National Archives Catalog (NARA)

HathiTrust

Planned

18M+ volumes

Digitised research library books; about a third readable in full.

Public domain volumes only.

Visit HathiTrust

Music discovery directory

A directory of 59 places to find music — databases, radio stations, record stores, magazines and digging tools. These are links out, not sources we copy from: nearly all of them are privately owned and reserve their rights. The exception is Discogs, whose release data is openly licensed even though its images are not.

Browse the directory

How we handle rights

  • Images are only kept when the source states a public-domain or Creative Commons licence.
  • Licence, artist and credit are stored with every image and shown on the card.
  • Fair-use or all-rights-reserved files are removed automatically during our regular sweeps.
  • Every card links back to the archive that holds the original record.
  • A share of what the site earns goes to the Wikimedia Foundation.
Against the dead internet

Bots wrote the feed. Models ate the web. Wikipedia and our open archives online are the last human-made commons left — support the real internet.

Donate to Wikipedia →