Automated requests

You found our crawler in your logs

PharosHub monitors hotel websites for intrusion and for copies published elsewhere. If a request from us reached your server, one of the three reasons below applies. If none of them fits, we want to know — that is a bug on our side, not yours.

Which one visited you

PharosHubCloneCheck

A page that loads our customer's beacon reported your domain as the host serving it. We fetch the page once to check whether it is genuinely a copy of their site before telling them anything.

PharosHubCTWatch

A certificate for a domain resembling a customer's was published to a public Certificate Transparency log. We fetch the site to see whether it serves their content.

PharosHubPosture

You are our customer, or your site is registered to one. We check from outside that the protections you believe are switched on actually are — HTTPS redirects, HSTS, framing rules.

What it does

  • Requests pages over HTTP, the same way a browser would
  • Reads the response and its headers
  • Fetches a small number of pages, then stops
  • Identifies itself honestly in every request

What it will never do

  • Probe for vulnerabilities, or request paths to see what breaks
  • Attempt a login, submit a form, or send anything you did not publish
  • Crawl at a rate intended to load your server
  • Disguise itself as a browser or rotate through addresses

If you would rather we did not

Block it. We will not treat that as evasion, and we do not retry around a block — a site that refuses us is simply a site we report nothing about.

User-agent: PharosHubCloneCheck
User-agent: PharosHubCTWatch
User-agent: PharosHubPosture
Disallow: /

Add that to your robots.txt, or block the user agent at your edge. Either works.

One exception, stated plainly: PharosHubCloneCheck does not consult robots.txt. It runs only when a page has already claimed, through our customer’s own beacon, to be serving their site from your domain. If that is a mistake we want to find out by looking; if it is not, the operator would simply add a robots.txt. The other two honour it.

Something looks wrong, or a request did not match anything above? Tell us — include the timestamp and the user agent and a person will read it.