About our crawler
If you are a site operator and found this address in your access logs, this page is for you. It says who we are, what we request, and how to make us stop.
Who is calling
Normdiff tracks published changes in security standards and regulations — new versions, amendments, effective dates — and tells subscribers what changed. To do that we read a small, hand-curated list of pages and feeds published by standards bodies, regulators, and national authorities.
Every request we make carries this User-Agent:
Normdiff/0.1 (+https://normdiff.securityvp.ai/crawler; [email protected])
We never send any other identity. We do not pretend to be a browser, and we do not pretend to be another company's crawler. If a request in your logs claims to be Googlebot, it is not us.
What we actually do
- We read, we do not crawl. There is no link-following and no site discovery. We request specific URLs that a human put on a list, and nothing else on your domain.
- Rarely. Most sources are polled a few times a day. The
fastest is every 30 minutes, and where a site declares a
Crawl-delaywe keep to it. - One request at a time per domain, with a deliberate pause
between them. We use conditional requests (
ETag,If-Modified-Since) wherever a server supports them, so most of our calls cost you a 304 and no body at all. - We take published text, not accounts or paid content. We do not log in, do not submit forms, do not buy or borrow credentials, and do not fetch anything sold behind a paywall.
What we do about robots.txt
We respect it. Of the sources on our list, all but seven are fetched from hosts
whose robots.txt permits what we request, and where a host declares
a Crawl-delay we keep to it.
There are documented exceptions on seven hosts, and we would rather name them here than let you find them in your logs:
zakon.rada.gov.uaanddata.rada.gov.ua, two hosts of the Parliament of Ukraine, whoserobots.txtdisallows everything while the pages themselves publish the official texts of laws and open data.www.healthnz.govt.nz, which allows Googlebot, Bingbot and DuckDuckBot and disallows everyone else. The single page we read is the public home of the Health Information Security Framework, the security standard New Zealand health providers and their suppliers must meet.www.bclaws.gov.bc.ca, which allows Googlebot and Bingbot and blocks everything else. The single page we read is the Table of Legislative Changes for British Columbia's Personal Information Protection Act.www.bma.bm, the Bermuda Monetary Authority, which blocks all crawlers by default and names a few search engines as exceptions. The single page we read is the public list of insurance policy and guidance documents. We do not touch the two paths itsrobots.txtsingles out,/viewPDF/documents/and/pdfview/.www.ismap.go.jp, the Japanese government cloud assessment programme, which disallows everything but/csm. We read its announcements API instead, because/csmis an empty shell that renders no content at all without a browser.leginfo.legislature.ca.gov, California Legislative Information, which disallows every crawler on every path and asks for ten seconds between requests. We read two pages there, each the published text of one section of the California Civil Code: 1798.91.04 on the security of connected devices, and 1798.82 on notifying people after a data breach. This is the one host on this list that names a delay, and we take thirty seconds rather than the ten it asks for.
On every one of these we fetch the listed URLs, no more than once a week each, and never less than thirty seconds apart — slower than we fetch anywhere else — under this same identity and this same contact address. We do not pretend to be Googlebot or any other agent to get in, and we never will: ignoring a request is one thing, lying about who is asking is another. Each of these exceptions was a decision a person took by name, not something our code can do on its own.
How to make us stop
Email [email protected] and say which domain or paths you want left alone. There is no form and no appeal process. We remove the source and it stays removed — this is a list maintained by a person, not a policy engine, so a one-line email is genuinely enough.
If you would rather act first and talk later, blocking our
User-Agent works and we will not route around it. We do not rotate
addresses or change our identity to get back in.
If we are costing you money
Tell us and we will fix it, even if you do not want us gone. We can poll less often, switch to a feed or API endpoint you would rather we used, or take a mirror if you publish one. A source that is expensive for you to serve is a source we would rather read a better way.
Contact: [email protected]