Five times we pointed the same checklist at a different population of the Polish web: 440 organisations across five samples, each drawn from a register somebody else maintains. This page is not a sixth measurement — not one page was fetched for it. It asks the only question none of the five studies could ask alone: which of these results belong to an industry, and which belong to the Polish web.
The answer is shorter than we expected. On 3 of 7 findings the five populations land within 10 points of one another — a listed company, a shop, a software house, a city hall and a hotel come out the same. One stands apart. And one item returned no verdict about anyone at all, which is a finding about our instrument rather than about anybody measured.
Almost everything published about visibility in AI answers is measured on English-language sites. This is a whole non-English market, measured five times with one instrument, and the five registers are not ours — every sample is defined by an institution with no interest in the result:
The basis is comparable, not sample size, and that is not a detail. The first study dropped domains it could not fetch; the four that followed keep them in the sample with a “cannot verify” verdict. Lining the raw sample sizes up would hand the younger populations a difference manufactured by that mismatch rather than by the market — so every column counts the organisations its own study treats as comparable.
| Finding | companies in WIG20 and mWIG40 | e-commerce domains | software houses | city governments | five-star hotels |
|---|---|---|---|---|---|
| whether anyone cites them — undecidable from the page source | 55 of 55 (100%) | 128 of 128 (100%) | 94 of 94 (100%) | 59 of 59 (100%) | 59 of 59 (100%) |
| carries no statement by a named person | 52 of 55 (95%) | 123 of 128 (96%) | 89 of 94 (95%) | 57 of 59 (97%) | 58 of 59 (98%) |
| shows no change history | 47 of 55 (85%) | 111 of 128 (87%) | 86 of 94 (91%) | 54 of 59 (92%) | 57 of 59 (97%) |
| links no credible external sources | 53 of 55 (96%) | 121 of 128 (95%) | 85 of 94 (90%) | 20 of 59 (34%) | 59 of 59 (100%) |
| declares the Article type in structured data | 0 of 55 (0%) | 7 of 128 (5%) | 4 of 94 (4%) | 1 of 59 (2%) | 0 of 59 (0%) |
| has a complete Organization entity | 4 of 55 (7%) | 19 of 128 (15%) | 16 of 94 (17%) | 0 of 59 (0%) | 1 of 59 (2%) |
| turns AI crawlers away | 3 of 55 (5%) | 24 of 128 (19%) | 8 of 94 (9%) | 6 of 59 (10%) | 5 of 59 (8%) |
Wikidata is absent from this table although all five studies measure it. The probe asks for
an entity under the name the instrument read off the page — and among city governments almost
none publish a machine-readable name for the office, so the name became the first segment of
the <title>, which for some of them is literally “Home page”. A column
where one cell means something different from the rest is worse than no column.
The sharpest repeat is no named human attached to the content: 52 of 55 listed companies, 123 of 128 shops, 89 of 94 software houses, 57 of 59 cities and 58 of 59 hotels carry not one statement signed with a first and last name. The spread across five independent samples is 3 points. Just outside the threshold sits showing no change history — 85–97%, meaning no trace anywhere that the content was ever updated. Its 12-point spread is 2 points past the line at which we are willing to call a finding repeated, so we do not call it that: the gap between populations is too wide to speak of one property of the Polish web.
Both are cheap to fix and both are exactly what an engine looks for when deciding whether to cite a page: who wrote this, and when. A result that repeats across five populations stops being a trait of an industry and starts being a trait of how company websites get built in this country. Which is why a threshold decides it and not an impression: the measured spread says at which of those two findings we are allowed to say that, and at which we are not yet.
One finding breaks the pattern, and it breaks it the good way. On “links no credible external sources” four populations sit between 90–100%, and city governments sit at 34%. The distance to the nearest of the other four is 56 points, which is more than the spread of every other finding in this table put together (51 points).
The explanation is structural, not meritorious: a city government writes about things with a legal basis, so it links statutes, regulations and official journals. For reasons that have nothing to do with visibility in AI, it does precisely what that visibility requires. A company wanting to repeat the result does not need a new idea — it needs to stop writing about its own field without pointing at anything outside its own website.
The item “whether anyone cites them — undecidable from the page source” reads 100% in every column, and that does not mean nobody cites them. It means that across all 395 organisations not one verdict was reached — because it is not visible in the page source. Citation happens on the engine's side; we measure the preconditions.
We publish this rather than omit it, for two reasons. First, a site whose claim is “we measure visibility in AI” owes the reader a statement of which part of it we do not measure. Second, 100% in a results column looks identical to a finding about the measured — and a reader taking it that way would leave convinced that nobody in Poland is cited. Those are opposite statements under the same notation.
Last updated:
All five runs are published in full under CC BY 4.0, row by row, together with the sampling rules — in Polish, with machine-readable data in any language:
The comparison itself — seven findings on an evened basis, the spread of each and the thresholds deciding what counts as repeated — is a separate file, because it appears in none of those five: cross-section data (JSON, CC BY 4.0). To cite a single number about one population, reach for that population's own study above.