What it does
The Medialake crawler visits the public pages of your website, finds the images and other media on them, and files everything into your Medialake library — each asset tagged with the page it came from, its caption or alt text, and the original source URL.
The goal is a complete, current picture of the creative actually live on your site: campaign banners, editorial photography, product imagery, and the pages they appear on. Nothing more.
How it works, in three steps
1. It builds a list of pages.
Rather than clicking blindly through every link on your site, the crawler reads your website's own sitemap — the same index your site publishes for Google and other search engines. This tells it exactly which pages exist, so it visits the right pages once instead of wandering through the site guessing. Pages that hold no creative — checkout, account, search, legal boilerplate — are filtered out before we start.
2. It visits each page the way a person does.
The crawler drives a real Chrome browser. It opens the page, waits for it to finish loading, dismisses the cookie banner, and scrolls down through the page so that images which only load as you scroll come into view — exactly what happens when one of your visitors reads the page. Because it sees the finished, rendered page rather than raw code, it picks up what a person actually sees: hero banners, background images, carousels, and gallery photography.
3. It collects the media.
For every image on the page, the crawler takes the highest resolution version your site offers, downloads it once, and files it in your library alongside the page title, the image's caption or alt text, and its original address. Duplicates are recognized and stored once, no matter how many pages they appear on.
Impact on your website
There is no noticeable impact on your site's performance or availability. The crawler is deliberately built to be a quiet visitor, not a stress test:
What we control | How it behaves |
Pages open at once | Two or three — the same footprint as two or three people browsing your site simultaneously |
Pace | A deliberate pause between every page; the crawler is never in a hurry |
Reads only | Every request is an ordinary page view. Nothing is submitted, changed, or written to your site |
Your rules respected | We follow your |
No repeat work | Anything already collected in the last 30 days is skipped, and each image is downloaded only once |
A full pass over a large multi-language site works out to a few thousand page views spread over several hours — a rounding error next to normal traffic, and a fraction of what search engines request from the same site every day.
If your infrastructure team would prefer a slower pace, a specific time window, or a particular set of sections, we can set that up. It is a configuration change on our side, not a rebuild.
What we collect, and what we don't
We collect
Images: campaign and editorial photography, banners, hero and background imagery, product photography
Video files hosted on your own site, where present
References to embedded video (for example a YouTube player on one of your pages), so you have a record of which video appears where
Light page context: page title, description, and — on product pages — the product name, reference numbers, and listed price
We never
Log in, create accounts, fill in forms, or place anything in a basket
Collect personal data, customer data, or anything behind a login
Change, add to, or remove anything from your website
Visit sections your
robots.txtasks crawlers to avoid
A note on embedded video: most brands host video on YouTube or Vimeo rather than on their own servers. Where that's the case, the video file itself sits with that platform, not with your website, so the crawler records which video appears on which page and keeps the preview image — it can't retrieve a file your site doesn't hold. If you want those masters in Medialake, the cleanest route is to supply them directly.
Identifying us
If your security or infrastructure team wants the crawler positively identified — so it can be allow-listed and excluded from your analytics — we can run it under a dedicated Medialake user agent from a fixed IP address, and we'll share both with your team. Several clients run it this way. Just let us know it's your preference.
Keeping it current
Sites change, so the crawl is repeatable. On each run the crawler picks up anything new or updated and skips what it already has, which keeps both the collection time and the load on your site low. Runs can be scheduled at whatever cadence suits your release rhythm, or triggered on demand after a campaign goes live.
In short
A real browser opens your public pages, scrolls through them like an ordinary visitor, and files the creative it finds into your Medialake library — reading only, at the pace of two or three people browsing, within the rules your site already publishes for search engines.
Questions, or a specific configuration you'd like? Talk to your Medialake contact and we'll set it up.
