normdiff Compliance change feed
Feed Frameworks Deadlines Pricing Account EN УКР

About our crawler

If you are a site operator and found this address in your access logs, this page is for you. It says who we are, what we request, and how to make us stop.

Who is calling

Normdiff tracks published changes in security standards and regulations — new versions, amendments, effective dates — and tells subscribers what changed. To do that we read a small, hand-curated list of pages and feeds published by standards bodies, regulators, and national authorities.

Every request we make carries this User-Agent:

Normdiff/0.1 (+https://normdiff.securityvp.ai/crawler; [email protected])

We never send any other identity. We do not pretend to be a browser, and we do not pretend to be another company's crawler. If a request in your logs claims to be Googlebot, it is not us.

What we actually do

  • We read, we do not crawl. There is no link-following and no site discovery. We request specific URLs that a human put on a list, and nothing else on your domain.
  • Rarely. Most sources are polled a few times a day. The fastest is every 30 minutes, and where a site declares a Crawl-delay we keep to it.
  • One request at a time per domain, with a deliberate pause between them. We use conditional requests (ETag, If-Modified-Since) wherever a server supports them, so most of our calls cost you a 304 and no body at all.
  • We take published text, not accounts or paid content. We do not log in, do not submit forms, do not buy or borrow credentials, and do not fetch anything sold behind a paywall.

What we do about robots.txt

We respect it. Of the sources on our list, all but seven are fetched from hosts whose robots.txt permits what we request, and where a host declares a Crawl-delay we keep to it.

There are documented exceptions on seven hosts, and we would rather name them here than let you find them in your logs:

  • zakon.rada.gov.ua and data.rada.gov.ua, two hosts of the Parliament of Ukraine, whose robots.txt disallows everything while the pages themselves publish the official texts of laws and open data.
  • www.healthnz.govt.nz, which allows Googlebot, Bingbot and DuckDuckBot and disallows everyone else. The single page we read is the public home of the Health Information Security Framework, the security standard New Zealand health providers and their suppliers must meet.
  • www.bclaws.gov.bc.ca, which allows Googlebot and Bingbot and blocks everything else. The single page we read is the Table of Legislative Changes for British Columbia's Personal Information Protection Act.
  • www.bma.bm, the Bermuda Monetary Authority, which blocks all crawlers by default and names a few search engines as exceptions. The single page we read is the public list of insurance policy and guidance documents. We do not touch the two paths its robots.txt singles out, /viewPDF/documents/ and /pdfview/.
  • www.ismap.go.jp, the Japanese government cloud assessment programme, which disallows everything but /csm. We read its announcements API instead, because /csm is an empty shell that renders no content at all without a browser.
  • leginfo.legislature.ca.gov, California Legislative Information, which disallows every crawler on every path and asks for ten seconds between requests. We read two pages there, each the published text of one section of the California Civil Code: 1798.91.04 on the security of connected devices, and 1798.82 on notifying people after a data breach. This is the one host on this list that names a delay, and we take thirty seconds rather than the ten it asks for.

On every one of these we fetch the listed URLs, no more than once a week each, and never less than thirty seconds apart — slower than we fetch anywhere else — under this same identity and this same contact address. We do not pretend to be Googlebot or any other agent to get in, and we never will: ignoring a request is one thing, lying about who is asking is another. Each of these exceptions was a decision a person took by name, not something our code can do on its own.

How to make us stop

Email [email protected] and say which domain or paths you want left alone. There is no form and no appeal process. We remove the source and it stays removed — this is a list maintained by a person, not a policy engine, so a one-line email is genuinely enough.

If you would rather act first and talk later, blocking our User-Agent works and we will not route around it. We do not rotate addresses or change our identity to get back in.

If we are costing you money

Tell us and we will fix it, even if you do not want us gone. We can poll less often, switch to a feed or API endpoint you would rather we used, or take a mirror if you publish one. A source that is expensive for you to serve is a source we would rather read a better way.

Contact: [email protected]

Feed Frameworks Deadlines Pricing Privacy Terms Cookies Crawler

Normdiff tracks published changes in standards and regulations. It is not legal advice.