Skip to content

[ Definition ]

Entity clarity: getting recognised before getting cited

[ In short ]

Before deciding whether to cite you, an engine has to work out who you are and tell you apart from namesakes. It cross-checks your structured data, your canonical identifiers, and how third-party sources describe you. When those three layers disagree the system does not pick the likeliest one: it stays uncertain and moves to a clearer source.

Published
Reading
8 min

What does "entity" mean to an engine?

A resolvable subject, not a string of text. That is the whole difference. "Rossi Consulting" as a string is a sequence of characters appearing on many pages; as an entity it is one specific subject, with an address, a sector, a founder, and a set of facts distinguishing it from any similar name.

A generative engine works at the second level, because it has to answer questions about subjects rather than about strings. When you ask who does a certain thing in a certain city, the system is not looking for a textual match: it is looking for subjects it can attribute that property to with enough confidence.

This inverts the order of priorities relative to SEO habit. You do not first get found for a keyword and then build reputation. You first become an identifiable subject, because until that point there is nothing for the reputation you are building to attach to.

Why does ambiguity cost more than a missing fact?

Because a missing fact produces a gap, while ambiguity produces distrust of everything else. If your page does not say which city you work in, the system does not know and moves on. If it says so in three places with three different answers, the system records an inconsistency, and from then on treats the correct fields with caution too.

It is the rational behaviour of any system deciding how much trust to give a source. Absent information is neutral. Contradictory information is a signal about the quality of whoever provided it.

In practice the most expensive ambiguity is never deliberate. It is the old profile never updated, the address from two moves ago still sitting in a directory, the name written two ways across site and invoices, the bio carrying a different number than the About page. None of these is a decision: they are all residue.

Which signals does an engine use to resolve the entity?

Three layers, weighted differently and broken differently. You fully control the first, partly the second, almost not at all the third, and it is no accident that the order of control is the inverse of the order of authority.

The third layer outweighs the first, and it is the one you control least. Hence the value of third-party sources.
LayerWhat it declaresHow it breaks
Structured dataWho you are, according to you: name, type, address, contact, founderSchema contradicting the visible page
Canonical identifiersWhich already-known subject you correspond tosameAs pointing at dead profiles or a namesake
Third-party mentionsWho you are, according to someone with no interest in saying itDiverging descriptions across directories and profiles
Internal consistencyThat your own pages say the same thingBio, footer, and privacy notice that do not agree

What has to match character for character?

A small and boring set, which is exactly why almost nobody fixes it. It takes no creativity, it takes a list and an afternoon.

The criterion is literal: not similar enough, identical. "Via Lattuada 26" and "Via Lattuada Serviliano 26" are two different addresses to a system comparing strings, even though to you they are the same place. The same goes for a legal name with and without its legal form, and for a brand name with and without unusual capitalisation.

  • Business name, in the same form across site, public profiles, and commercial material.
  • Full address, with the same punctuation and the same abbreviations.
  • Contact: the same email and the same number, with no format variants.
  • Short description: one canonical version, reused rather than rewritten each time.
  • Declared category: if the site says independent consultancy and a profile says agency, you have two subjects.
  • Founder and role, linked to the company subject rather than merely named.

What is a canonical identifier for?

Removing the namesake problem. An identifier in a public register is not a name, it is a key: two companies can share a name, they cannot share a key. When one exists, the system stops having to guess which of the two subjects you mean.

In `schema.org` this is expressed with `sameAs`, linking your subject to its profiles and its entries in external registers. It is the most underused field in the markup, and also the most abused: many fill it with any link at all, and a `sameAs` pointing at a page that does not describe the same subject is worse than absence, because it is a false claim made by you.

One clarification that matters: `sameAs` is a claim, not proof. It says "I assert I am this", and it gets verified by checking whether the linked profile says the same thing. If your public profile carries a different description than your site, the link does not confirm identity, it questions it. The markup is worth exactly as much as the consistency of what it points at.

How do you check whether an engine has resolved you?

By asking it, and reading the answer for the type of error rather than for the presence of your name. The useful question is not "does it talk about me", it is "does it talk about me or about someone else with my name".

Do it in a clean session, with no history, and several times, because a single answer cannot separate stable behaviour from an accident. The conditions to declare are the same as for any other measurement.

  • Ask "who is [name]" and check whether the answer describes you or a namesake.
  • Ask which city it operates in and in what category: those are the two fields that go wrong first.
  • Look for merged facts: an answer mixing your history with another company's is the clearest symptom of an unresolved entity.
  • Check whether it cites stale information. If so, a dated version exists somewhere and outweighs the current one, and it needs finding.
  • Compare the answer against your canonical description. The divergences tell you which sources the system is preferring over yours.

What does entity clarity not do?

It does not make you authoritative. It makes you identifiable, which is the prerequisite. A perfectly resolved subject with nothing relevant to say is still a subject nobody has reason to cite.

It is not a ranking factor you switch on. Structured data communicates information about a page; it is not a lever that raises visibility. Anyone presenting it as one is selling a reading no official documentation supports.

It does not create demand. If nobody asks questions you are a plausible answer to, being identifiable produces no citations: it only produces the certainty that, when those questions arrive, you will not be discarded for ambiguity. It is defensive work, and should be judged as such: cheap, invisible when it works, very visible when missing.

[ What to take away ]

  • Before being citable you have to be identifiable: two pieces of work in sequence, not alternatives.
  • Ambiguity costs more than a missing fact, because it casts doubt on the correct fields too.
  • Identical means identical: name, address, contact, and description in the same form, everywhere.
  • sameAs is a claim, not proof. It is worth as much as the consistency of what it points at.
  • Check by asking engines who you are, in a clean session and several times, and read the type of error.

How visibility inside a generative engine gets measured, written out in full. Read the article

[ Author ]

Nicola Dussin

Founder of Creaitivo. Every analysis is run directly by me.

LinkedIn profile

Want the same reading on your case?

The automated preview gives a first signal in seconds. I prepare the useful reading myself, and it arrives within 48 hours.

Send me your case
All insights