curatebot — the ZETO crawler

ZETO shows people the sale items, in their size, from the online stores they follow. To do that it reads those stores' public product pages with a crawler that identifies itself, follows robots.txt, and goes slowly.

Identity

Product token
curatebot
User-Agent
curatebot/<version> (+https://shopcur8.com/crawler)
Operator
ZETO
Contact
tcmprog@gmail.com
Egress IPs
Not published yet. Write to the address above and we will confirm whether a specific request was ours.

One User-Agent, always over HTTPS, never varied per site. We never rotate addresses or disguise curatebot as a browser to get around a block.

What it fetches

What it stores

Opting out

Add this to your robots.txt and the crawler stops within 24 hours (it re-reads robots.txt at least that often):

User-agent: curatebot
Disallow: /

Disallow rules for specific paths are honoured the same way, as is Crawl-delay. For an immediate stop, or to have already-collected product data removed, email tcmprog@gmail.com.

One exception, stated here so it is never a surprise: before ZETO reads anything else from a store, and about once a month after that, it reads the store's public terms page (on Shopify, /policies/terms-of-service) even where robots.txt disallows that path, because ZETO cannot honour terms it is not allowed to read. What that page says is read, recorded and reviewed by a person; what the crawler obeys is robots.txt. A robots rule, a 403, a challenge, or an email to the address above stops it — even that one page. The crawl is disclosed here because a crawl found later is the exposure, not the read itself.

For retailers

What a store can do to be listed on ZETO — a product feed, or a robots.txt allowance plus a written yes — and what ZETO reads first and honours, is on the retailers page.

How fast it goes