What a crawler sees when it opens your site
A search engine, a messenger and an AI crawler read something entirely different from a human. The most expensive mistake in that group is one line in a file most companies do not know about.
A human opens a website and sees images, layout and colour. A search engine crawler sees text and a few files. A messenger sees four tags. An AI crawler sees whatever it has been allowed to.
Those three views can drift a long way from what you see.
One line that removes a site from search
There is a robots.txt file in the root directory. If it contains this:
User-agent: *
Disallow: /
— the site is closed to every search engine. It looks normal, opens normally, works normally. It simply does not exist in the results.
That line usually gets there while a new version is being built, so the test version does not enter the index. And it stays, because nobody remembers to take it out.
Checking takes fifteen seconds: type your own address in the browser and add /robots.txt to it.
A sitemap nobody was told about
The sitemap.xml file is the list of addresses you want to see in search. Sometimes it exists but is not mentioned in robots.txt — and then the search engine may never find it.
Sometimes it is the other way round: robots.txt points at a sitemap that returns a 404. Both situations look identical from a browser, which is to say invisible.
Four tags that decide what the link looks like
You paste your address into WhatsApp or LinkedIn. What appears — title, description, image — comes from four tags in the code: og:title, og:description, og:image, twitter:card.
If they are missing, the link looks like a bare address. If og:image points at a file that does not exist, the result is the same. Our own site had the second case until August 2026 — the tag was there, the file was not.
AI crawlers are a separate decision
GPTBot, ClaudeBot and their kind read pages for models. They can be blocked in robots.txt, and that is an entirely legitimate choice for an owner to make.
It is only worth knowing that the choice was made. It happens that a contractor entered the block without asking — and the company finds out only when it notices that an AI assistant describes its competitors and knows nothing about it.
The same goes for the llms.txt file, which tells models what the company does. It is not mandatory and it is not a standard — it is simply cheap.
What you can check yourself, in fifteen minutes
Three addresses in a browser: /robots.txt, /sitemap.xml, /llms.txt. Then paste a link to your site into any messenger and see what comes up.
The rest — which structured data is in the code, whether the site loads resources over http on an https page, what exactly a crawler sees — needs a tool. Our scan checks all of it at once and shows the result on screen, without registration.
Start with robots.txt. It is the one field where a mistake costs the most and the fix takes a minute.
Do you have the same problem?
Start with the free scan — it will show whether what is described above applies to your site. Or book a call and we will go through it together.