Our crawler
When someone asks for a company’s Digital Reality Score, we read the company’s own website to see what it states about itself: its name, what it does, who leads it, where it is. This page says who we are, what we fetch and how to allow or block us.
How to recognise us
User agent:
DigitalRealityScore/0.1 (+https://digitalrealityscore.com/crawler/) robots.txt name: digitalrealityscore (case does not matter, so DigitalRealityScore works too).
We identify ourselves honestly and never pose as another crawler, a browser or a person.
When we visit
- Only when someone asks for a Score of that domain on this site, or we re-run one to test the method. Never on a schedule, and never to crawl the web.
- A published Score is reused for 30 days, so repeat requests for the same company do not bring us back.
What we fetch, per visit
/robots.txt, first. We follow it (below).- The homepage.
- One About page: a link on the homepage that looks like one, otherwise a few usual paths such as
/aboutor/om-os, one at a time, until one exists. /llms.txt, to report whether it exists (for information; Google Search does not use it).
We fetch from Cloudflare’s network, where this site runs. When a site cannot be reached from there because of a network or certificate error, the same requests go through a second fetch path outside Cloudflare’s network: same user agent, same robots.txt rules.
One request at a time, each with a short timeout. We do not run scripts, fill in forms, log in, or follow other links. We keep what the pages state about the company (for example its name or its CEO) with the page address, not copies of the pages.
When a company lists another company as its partner or the brands it distributes, we may also read that partner’s own list (a brand’s distributor page, a vendor’s partner directory) to see whether it names the company. The same rules apply there.
How we follow robots.txt
- A group naming
digitalrealityscoreapplies to us; otherwise the*group does. - No robots.txt (the site answers “not found”): we may read the pages above.
- A robots.txt we cannot fetch because of a server or network error: we read nothing until we can, as the robots.txt standard (RFC 9309) asks.
To let us read your site
User-agent: DigitalRealityScore
Allow: / Your Score then uses what your own site states. When we cannot read a site, what it states is taken from its own pages as search results show them, and the profile says so.
To block us
User-agent: DigitalRealityScore
Disallow: / Blocking us does not remove a company from the Digital Reality Score. Its Score is then based on what other sources show, and its profile says the site could not be read.
What we do not test
Our crawler access check reports what your robots.txt declares for search engines and AI crawlers, and what happened when we visited as ourselves. A server or firewall can still turn away a named crawler whatever robots.txt says. We do not test that, because it would mean posing as that crawler.
Contact
Questions or problems with our visits: hello@digitalrealityscore.com. How Scores are made is on the methodology page.