Project Common Screens is a corpus of web screenshots and metadata covering more than 70 million websites, published as open data on AWS and updated monthly.
Convert any host name into the flat S3 object key and its public URL.
Documented S3 prefixes for images and derived data, bucket common-screens in us-west-2.
Available through the Registry of Open Data on AWS, free to use under CC BY 4.0.
Step-by-step guides for deriving screenshot URLs and working with the corpus.
Screenshots come with metadata for analysis, research and machine learning.
Related Dosvak projects cover AI-built software, compute, AI assets, genealogy and more.