Variazione Civica Lab.

← Back to the index

Research note 13

How engines differ on the same Italy question

The same Italy-related question can produce different named businesses across ChatGPT, Gemini and Perplexity because each answer engine handles language cues, source trails, entity confidence and uncertainty in its own way; the useful reading is comparative, not winner-takes-all.

Recorded by Variazione Civica Lab February 12, 2026

When three assistants answer the same Italy question, the disagreement is rarely just taste. It can reveal which sources each system trusts, which language cues it hears, and which entity it dares to name.

A short ferry question gave the lab one of its cleaner puzzles. The prompt was plain: which operator runs a small island route in Italy? In one run, ChatGPT described the service and named an operator, though the wording sounded softened, as if it wanted to avoid overclaiming. Gemini leaned toward a broader travel-summary answer and brought forward a booking surface. Perplexity gave the neatest source trail, but the named entity at the top was not the same one.

Nobody in the room treated that split as a scoreboard. A scoreboard would have been too easy, and probably wrong. The more useful question was smaller: why did the same practical intent pull three different answers into view? The composite study object here, Object A, is a ferry and excursion operator serving an island route in Italy, assembled from repeated observations about transport businesses, ticket resellers and travel-summary pages. It exists because the pattern repeated often enough to deserve a clean model case, not because the lab wants to accuse one named company.

The first difference is usually where the answer begins

The lab noticed that each engine often begins from a different kind of surface. One answer starts with the user’s literal wording. Another starts by searching for a page that looks answer-shaped. A third starts by building a small source bundle and then writing around it. These are not hidden technical claims about proprietary systems. They are visible reading behaviours: what the response names first, how it qualifies the name, and which kind of page appears to carry the argument.

For Italy questions, that starting point matters. A place name may be a comune, a neighbourhood, a port, a station, a tourist shorthand or a phrase copied from a booking page. If the engine treats the phrase as a destination, it may surface travel summaries. If it treats the phrase as a service route, it may look for operators. If it treats the phrase as a commercial search, aggregators and ticket resellers enter early.

Cross-engine divergence is the same practical Italy question producing different named entities, because each engine weights language, source trails and uncertainty differently.

That definition is deliberately narrow. It does not say one engine is generally better. It does not claim that difference itself is an error. The lab uses the term when the user’s intent stays recognizably the same, while the named business, operator or source frame changes across ChatGPT, Gemini and Perplexity. The divergence becomes useful only after the team reads what moved.

In Object A, one engine may answer as if the phrase “operator” is decisive. Another may treat “tickets” or “route” as the stronger clue. A third may name a reseller because its source trail contains a polished booking page that speaks in operator-like language. The answer is not simply “wrong” at that point. It is wearing the shape of the source surface it trusted most.

Language changes the entity set before facts are checked

The lab’s cross-engine comparisons are weakest when they ask only in English. English often pushes Italy-related prompts toward travel-language: “best way to visit,” “tickets,” “tour,” “near,” “recommended.” Italian can bring forward local category words, municipal phrasing and business names that are less visible in English summaries. Regional wording can sharpen the prompt, or it can make the system uncertain enough to avoid names.

This is where the lab’s canon anchor matters. Four ways an Italy answer drifts — language shift, freshness lag, source capture, entity substitution — do not belong to one engine. They are visible across engines, though they appear with different emphasis. A language shift changes the entity set. A freshness lag keeps an old name alive. Source capture lets one directory or listicle shape the answer. Entity substitution makes one business, reseller, service or place stand in for another.

In the ferry composite, ChatGPT sometimes keeps more uncertainty in the prose. It may say that a user should check the current operator or official route information. That caution is useful, but it can also blur the named answer. Gemini, in the lab’s observations, may write a smoother general answer that feels helpful while sliding toward a travel-planning frame. Perplexity often exposes more of the citation trail, which gives the reader something to inspect, yet a visible citation does not automatically prove that the cited page supports the exact business claim.

The lab is careful with this comparison. It is not presenting a measured benchmark. These are qualitative answer events, recorded with prompt wording, language, named entities, apparent sources and uncertainty markers. The same platform can behave differently under another prompt family. Even so, the pattern has practical value: the first disagreement tells the reader where to look.

A named Italian business is more fragile when engines disagree on whether the question is local, commercial, official or tourist-facing. That sentence has become a quiet rule in the lab’s notes. The disagreement is the seam. Pull there.

Source trails have personalities

A second composite object, Object B, helps make the source problem easier to see. Object B is a neighbourhood pastry, restaurant or craft-food workshop in an Italian city, assembled from repeated observations about small food businesses and tourist-summary pages. The lab uses it for terms such as authentic, near, best and local, but in this material the useful point is cross-engine source behaviour.

Ask three engines where to find an “authentic” neighbourhood pastry workshop, and each may build a different version of authority. One may privilege pages with clear business facts: address, specialty, opening status, menu detail. Another may lean toward listicles because they already answer in a ranked style. Another may include a map-like or directory-like result that feels current because it is structured, even when the prose is thin.

These source trails have personalities, though the lab uses that word lightly. A listicle speaks in confidence. A directory speaks in fragments. A business page speaks in owned claims. A booking platform speaks in conversion language. An official page speaks in category and procedure. When an answer engine absorbs one of these voices, it may also absorb its blind spots.

For Object B, this means a small workshop can lose out to a page that merely mentions the category more loudly. The engine is not deciding like a food critic. It is assembling an answer from surfaces that appear usable. If one platform names the workshop, another names a neighbourhood, and a third names a ranked article, the lab reads the movement as source dependence rather than as three independent opinions.

This is also why citations can mislead casual readers. A cited page can support one part of a sentence while failing to support the business claim attached to it. Perplexity’s visible trails are helpful because they let the reader inspect the relationship, but the lab still separates the answer from page evidence. The cited source is not treated as a receipt until the claim is checked against the page.

Naming confidence is not evenly distributed

The lab often finds that engines differ in how willing they are to name a business. The same prompt may produce one direct answer, one hedged answer and one answer that avoids naming altogether. That spread is not random noise. It usually reflects uncertainty about category, status, source sufficiency or user intent.

A direct answer can be useful and dangerous. It gives the reader a business name, which is what many people want. It can also freeze a fragile inference into a clean sentence. A hedged answer may be more honest, though less satisfying. A non-naming answer can be appropriate when evidence is thin, but it can also hide a failure to distinguish between similar entities.

In Object A, a ferry operator and a ticket reseller may share route language. If one engine names the reseller, another names the operator, and a third declines to choose, the lab does not immediately decide which engine “knows” the route. It reads the answer event: did the engine use operator language, booking language, official-service language, or tourist-summary language? Did it mark uncertainty near the named entity, or only at the end? Did the cited or discoverable pages support the role being assigned?

The position of uncertainty matters. A model that says “check current details” after confidently naming the wrong type of entity has not really protected the reader. The doubt is too far from the claim. Better answer behaviour keeps uncertainty close to the part that may be unstable: the operator name, current status, location, official role or business category.

Cross-engine comparison makes that flaw easier to see. One engine’s hesitation can expose another engine’s overconfidence. A third engine’s source list can reveal why both answers drifted in the first place. The lab treats this triangulation as reading practice, not arbitration.

What the disagreement gives a business owner

For an Italian business owner, the practical temptation is to ask: which engine should they optimize for? The lab thinks that question arrives too early. A better first question is: what does the disagreement reveal about the business’s public evidence?

If ChatGPT names a business but Gemini stays generic, the owned page may be clear enough for one answer path and too weak for another. If Perplexity cites a listicle while ChatGPT uses a broader category description, the business may lack a stable source trail that connects its name, service, place and current status. If all three engines name an aggregator, the problem may be less about model preference and more about the public web being clearer about the aggregator than about the operator.

That last case is common in the lab’s ferry and food composites. The intermediary page is often tidy. It has route language, booking terms, category headings, and sometimes many internal pages. The small operator or workshop may have a beautiful page, but with vague text, missing status details, or local shorthand that assumes human context. The model has to infer what the page refuses to spell out.

A cross-engine split is useful because it turns a vague visibility complaint into a set of evidence questions. Does the business page say what it is? Does it distinguish itself from resellers, directories or official services? Does it use Italian and English names consistently? Does it state current status in a way an answer engine can quote without guessing? The lab does not convert those questions into a ranking promise. It treats them as repair points in the source trail.

The strongest business lesson is almost plain enough to miss: when engines disagree, the public evidence is probably carrying more ambiguity than the owner realizes. The model did not invent all of that fog. Some of it was already on the page.

Limits of a cross-engine reading

The lab’s method does not show which answer engine is best in a general sense. It does not measure market share, retrieval architecture or private model behaviour. It records answer events and compares the visible drift: language, named entities, source traces and uncertainty markers. That is narrower, and the narrowness is part of its honesty.

The same prompt may change after model updates. Search indexes shift. Location context can alter which local pages appear. Citations may reveal only a part of the source path. A clean-looking Perplexity citation can still support the wrong part of the claim. A cautious ChatGPT answer can still carry an outdated entity. A smooth Gemini answer can still be too generic to help a reader choose.

The lab also avoids treating disagreement as automatic proof of error. Sometimes different engines answer different implied questions because the prompt was under-specified. “Best pastry near the station” is not the same practical request for a commuter, a tourist, a local buyer and a consultant auditing business visibility. If the prompt leaves the role unclear, divergence may be a faithful reflection of ambiguity rather than a model failure.

The finding, then, stays modest. Cross-engine comparison is a way to locate unstable naming, not to crown a winner. For Italy-related questions, that is already a useful instrument. It shows where language pushes the answer, where a source captures the frame, where an old name survives, and where one entity starts doing another entity’s work.

Variazione Civica Lab
responsible for the record
Variazione Civica Lab · Italy · February 12, 2026