Five Polish populations, one instrument

Five times we pointed the same checklist at a different population of the Polish web: 440 organisations across five samples, each drawn from a register somebody else maintains. This page is not a sixth measurement — not one page was fetched for it. It asks the only question none of the five studies could ask alone: which of these results belong to an industry, and which belong to the Polish web.

The answer is shorter than we expected. On 3 of 7 findings the five populations land within 10 points of one another — a listed company, a shop, a software house, a city hall and a hotel come out the same. One stands apart. And one item returned no verdict about anyone at all, which is a finding about our instrument rather than about anybody measured.

Why a Polish sample is worth reading elsewhere

Almost everything published about visibility in AI answers is measured on English-language sites. This is a whole non-English market, measured five times with one instrument, and the five registers are not ours — every sample is defined by an institution with no interest in the result:

Five samples side by side

The basis is comparable, not sample size, and that is not a detail. The first study dropped domains it could not fetch; the four that followed keep them in the sample with a “cannot verify” verdict. Lining the raw sample sizes up would hand the younger populations a difference manufactured by that mismatch rather than by the market — so every column counts the organisations its own study treats as comparable.

One checklist, five samples. Each column counts the organisations its own study treats as comparable — 395 of 440 measured in total.
Findingcompanies in WIG20 and mWIG40e-commerce domainssoftware housescity governmentsfive-star hotels
whether anyone cites them — undecidable from the page source55 of 55 (100%)128 of 128 (100%)94 of 94 (100%)59 of 59 (100%)59 of 59 (100%)
carries no statement by a named person52 of 55 (95%)123 of 128 (96%)89 of 94 (95%)57 of 59 (97%)58 of 59 (98%)
shows no change history47 of 55 (85%)111 of 128 (87%)86 of 94 (91%)54 of 59 (92%)57 of 59 (97%)
links no credible external sources53 of 55 (96%)121 of 128 (95%)85 of 94 (90%)20 of 59 (34%)59 of 59 (100%)
declares the Article type in structured data0 of 55 (0%)7 of 128 (5%)4 of 94 (4%)1 of 59 (2%)0 of 59 (0%)
has a complete Organization entity4 of 55 (7%)19 of 128 (15%)16 of 94 (17%)0 of 59 (0%)1 of 59 (2%)
turns AI crawlers away3 of 55 (5%)24 of 128 (19%)8 of 94 (9%)6 of 59 (10%)5 of 59 (8%)

Wikidata is absent from this table although all five studies measure it. The probe asks for an entity under the name the instrument read off the page — and among city governments almost none publish a machine-readable name for the office, so the name became the first segment of the <title>, which for some of them is literally “Home page”. A column where one cell means something different from the rest is worse than no column.

What repeats

The sharpest repeat is no named human attached to the content: 52 of 55 listed companies, 123 of 128 shops, 89 of 94 software houses, 57 of 59 cities and 58 of 59 hotels carry not one statement signed with a first and last name. The spread across five independent samples is 3 points. Just outside the threshold sits showing no change history — 85–97%, meaning no trace anywhere that the content was ever updated. Its 12-point spread is 2 points past the line at which we are willing to call a finding repeated, so we do not call it that: the gap between populations is too wide to speak of one property of the Polish web.

Both are cheap to fix and both are exactly what an engine looks for when deciding whether to cite a page: who wrote this, and when. A result that repeats across five populations stops being a trait of an industry and starts being a trait of how company websites get built in this country. Which is why a threshold decides it and not an impression: the measured spread says at which of those two findings we are allowed to say that, and at which we are not yet.

What stands apart: city governments

One finding breaks the pattern, and it breaks it the good way. On “links no credible external sources” four populations sit between 90–100%, and city governments sit at 34%. The distance to the nearest of the other four is 56 points, which is more than the spread of every other finding in this table put together (51 points).

The explanation is structural, not meritorious: a city government writes about things with a legal basis, so it links statutes, regulations and official journals. For reasons that have nothing to do with visibility in AI, it does precisely what that visibility requires. A company wanting to repeat the result does not need a new idea — it needs to stop writing about its own field without pointing at anything outside its own website.

What our instrument decided about nobody

The item “whether anyone cites them — undecidable from the page source” reads 100% in every column, and that does not mean nobody cites them. It means that across all 395 organisations not one verdict was reached — because it is not visible in the page source. Citation happens on the engine's side; we measure the preconditions.

We publish this rather than omit it, for two reasons. First, a site whose claim is “we measure visibility in AI” owes the reader a statement of which part of it we do not measure. Second, 100% in a results column looks identical to a finding about the measured — and a reader taking it that way would leave convinced that nobody in Poland is cited. Those are opposite statements under the same notation.

What this comparison does not settle

Last updated:

All five runs are published in full under CC BY 4.0, row by row, together with the sampling rules — in Polish, with machine-readable data in any language:

The comparison itself — seven findings on an evened basis, the spread of each and the thresholds deciding what counts as repeated — is a separate file, because it appears in none of those five: cross-section data (JSON, CC BY 4.0). To cite a single number about one population, reach for that population's own study above.

Related