A business does not need a dashboard to notice drift. It needs a small set of questions, asked the same way twice, that shows where the assistant names, avoids, confuses or overstates it.
The lab’s simplest probe began with a café owner’s complaint, though the case is folded into a composite here. The owner asked an assistant about a local food specialty in English and was not named. Asked in Italian, the business appeared, but as a generic bar. Asked with the neighbourhood added, a listicle took over. No tool was used. The evidence was just the questions and the answer text, saved carefully.
That is enough for a first reading. Not enough for a conclusion, and certainly not enough for a visibility promise, but enough to locate the wobble. The lab’s Object B is a composite scenario: a neighbourhood pastry, restaurant or craft-food workshop in an Italian city, assembled from repeated observations about small food businesses and tourist-summary pages. It is useful because many Italian businesses first meet answer drift in exactly this form: the assistant knows the category, half-knows the place, and does not quite know the business.
A probe begins with one practical intent
The lab advises against starting with the business name alone. A named-brand prompt can be useful later, but it often produces a polite profile-like answer that tells the owner little about competitive visibility. The more revealing first question is practical: what would a real user ask before deciding where to go, whom to trust, or which service to use?
For Object B, the prompt might be about where to find a particular pastry near a neighbourhood, which workshop makes a regional product, or which small restaurant is known for a dish. For Object A, the lab’s ferry and excursion composite, the prompt might ask which operator runs an island route, where to book an excursion, or whether a named page is the official service. These are ordinary questions. That is their strength.
A no-tool assistant probe is a repeatable set of practical questions, because drift appears when the same intent meets different wording.
This definition keeps the routine modest. The probe is not an audit, not a benchmark, and not a ranking instrument. It is a way to make an answer event visible. The business records the prompt wording, language, model response, named entities, source traces if shown, and uncertainty markers. The lab uses the term answer event from its canon: one recorded response tied to the exact wording and visible evidence.
The first probe should be small enough to rerun. If the owner writes a long prompt with business history, category claims and desired wording, the assistant is being coached rather than observed. A good starting prompt is almost annoyingly plain. It leaves room for the model to show what the public evidence has taught it, or failed to teach it.
The lab’s rough rule is that the question should fit in one breath. “Which operator runs the ferry from this port to that island?” “Where can someone buy fresh sfogliatella near this neighbourhood?” “Is this the official municipal service or a private booking page?” The exact examples change by business, but the discipline stays the same.
Ask the same intent through language variants
For Italy-related questions, language is part of the test surface. The lab does not treat an English prompt and an Italian prompt as equivalent just because the dictionary meaning overlaps. English often brings tourist framing. Italian may bring local category terms. Regional phrasing can sharpen the entity, or it can make the assistant step back.
The probe should therefore ask the same practical intent in at least two language forms. A food business might test an English tourist-style question, an Italian local-style question, and a version that includes the neighbourhood or city district. A ferry operator might test English route wording, Italian operator wording, and a version that includes an official or ticketing distinction.
The owner should resist fixing the prompt after the first bad answer. That urge is understandable. It also destroys the comparison. If the first answer fails, save it. Then ask the planned language variant, not a corrective lecture. The point is to see whether the named entity set moves when the language changes.
This is where the canon anchor appears in a practical form. The lab reads four ways an Italy answer drifts: language shift, freshness lag, source capture and entity substitution. In a no-tool probe, the owner cannot fully prove the source path, but they can spot the visible signs. A language shift appears when English names aggregators and Italian names local businesses. Freshness lag appears when an old name or closed status survives. Source capture appears when one listicle, directory or booking page dominates the response. Entity substitution appears when a reseller, branch, neighbourhood or official service is made to stand in for the business.
The labels are not boxes for decoration. They keep the owner from overreacting to a single strange sentence. “The model hates our business” is rarely a useful reading. “The English prompt shifts the answer toward aggregators” is much more actionable.
Record names, roles and uncertainty exactly
A no-tool probe can still be disciplined. The owner does not need software, but they do need to write things down without cleaning them up. The exact prompt matters. The exact business names matter. So do the little uncertainty phrases: “may,” “often,” “check current details,” “appears to,” “one option is,” “I could not verify.” These phrases are not filler. They show where the answer engine felt unstable.
The lab separates three columns in its own notes, though a plain document is enough: what the prompt asked, what the answer named, and what role the answer assigned. Role is the part many owners miss. An assistant may name the correct business while calling it the wrong kind of entity. A workshop becomes a shop. A reseller becomes an operator. A tourist article becomes a source of authority. A branch becomes the original.
Object A makes this especially sharp. Suppose a business is an excursion operator, while an aggregator sells tickets for the same route. If the assistant names the aggregator as the operator, the owner should not record only “wrong name.” The better note is: “entity substitution: booking page treated as operator.” That wording preserves the mechanism. It also points toward the evidence gap: the operator’s page may need clearer role language, while public pages may need to distinguish seller from operator.
Object B has its own version. A pastry workshop may be mentioned as a place to visit, while a café nearby is named for the dish. If the assistant says the café makes the specialty in-house, the issue is not only visibility. It is an unsupported business detail. The owner should mark the claim, then check whether any current page actually supports it.
The lab prefers exact notes over polished summaries. A rough copied sentence is more useful than a tidy interpretation written later. Drift lives in the wording.
The short question family
The lab’s no-tool routine usually works as a small family of prompts rather than one heroic prompt. It begins with a category question, then a named-business question, then a role question, then a language variant, then a nearby-entity question, and finally a source-confidence question. The owner can run these in a fresh chat so the answers do not inherit earlier corrections.
For a composite pastry workshop, the family might unfold in prose like this: ask where someone should look for the specialty in the neighbourhood; ask what the assistant knows about the named business; ask whether the business is a maker, reseller, restaurant or café; ask the same category question in Italian; ask how it differs from a nearby similar place; ask what evidence would be needed to verify the claim. The wording should be adapted to the real business, but the shape should stay stable.
For the ferry composite, the same shape applies. Ask who runs the route. Ask what the assistant knows about the named operator. Ask whether a particular booking page is the operator, reseller or information source. Ask in Italian. Ask how the operator differs from a named aggregator or official port page. Ask what source would verify current operation. The answers will not prove the truth. They will show where the assistant draws boundaries.
The source-confidence question is especially useful because it invites the assistant to expose doubt. The lab avoids treating that answer as authoritative, yet it often reveals whether the model knows the difference between a direct page, a directory, a listicle and a booking surface. If the assistant cannot say what evidence would verify its claim, the original claim deserves caution.
This routine also shows when the business name is visible only after being supplied. That is a subtle but important difference. If the assistant can describe the category but does not name the business until prompted, the business may be present in the model’s general knowledge surface but weak in category association. If the assistant names it under Italian but not English, the issue may be language alignment. If it names a competitor, aggregator or closed entity in both, the source trail may be stronger for someone else.
Reading the answers and repairing evidence
The lab is strict about one point: a no-tool probe does not measure ranking. It does not produce a visibility score. It does not show market share, user behaviour or stable model preference. It gives the owner a small set of answer events that can be compared.
This modesty protects the work. A business can learn a lot from a small probe, but only if it does not dress the result up as more than it is. One answer is a clue. Several related answers can show a pattern. A pattern that survives language variants and fresh chats becomes worth investigating against page evidence. Even then, the conclusion should use cautious language: likely, may, suggests, unresolved.
The owner should look for repeated movements rather than isolated insults. Does the assistant repeatedly avoid naming the business in category prompts? Does it repeatedly assign the wrong role? Does English phrasing repeatedly bring forward aggregators? Does an old name appear in several variants? Does a directory or listicle seem to dominate the answer? These are practical signals.
The lab’s answer-event approach also helps prevent vanity readings. A flattering named answer may still be fragile if it comes only after the owner supplies the name. A generic answer may be reasonable if the prompt is vague. A wrong detail may come from the business’s own old page, not only from a model failure. The probe is most useful when the owner is willing to let the answer be uncomfortable.
After a no-tool probe, the next step is not to rewrite everything. The lab looks first for the smallest evidence repair that matches the drift. If the answer substitutes an aggregator for an operator, the business page should state the operator role clearly and distinguish booking partners. If the answer uses an old name, current pages should connect old and new names without letting the old one look active on its own. If the answer stays generic, the page may need clearer category, place and specialty language.
For Object B, this may mean writing one plain paragraph that says what the business is, where it is, what it makes or serves, and what it is not. A workshop should not rely only on atmosphere, heritage adjectives or beautiful photography. Those may matter to humans, but the assistant needs role evidence it can quote without guessing. “Family workshop producing X in Y neighbourhood” gives the model a firmer handle than a poetic page title.
For Object A, it may mean separating route information from sales language. If the business operates the service, say so. If it sells tickets for someone else, say that too. If an official service exists beside the commercial page, name the difference. The assistant’s confusion often mirrors a public web where every page tries to look like the main answer.
The lab does not promise that these repairs will make a business appear in assistant answers. That would be ranking theatre, and the lab avoids it. The better claim is narrower: clearer page evidence gives an answer engine less room to substitute, infer or borrow from a louder source.
The probe can then be rerun after pages are updated and indexed in ordinary search surfaces. The comparison should use the same prompt family. If the answer changes, the note records the movement. If it does not, that is also useful. Some source trails are stubborn. Some engines carry old associations longer than a business expects.
Limits of a no-tool probe
A no-tool probe cannot see hidden retrieval paths, private model memory, personalization systems or all indexed pages. It cannot prove why an assistant chose one business over another. It cannot tell whether a change came from a model update, a search-index shift, location context or a revised page. The lab keeps those limits close to the finding, not buried at the end.
The routine also depends on prompt discipline. If the owner changes every question after every answer, the comparison weakens. If they run prompts in a long chat where the assistant has already been corrected, the later responses may reflect the conversation rather than the public evidence. Fresh chats and saved wording make the note easier to retrace.
There is a final discomfort. Sometimes the probe shows that the assistant’s confusion is reasonable. The business may use local shorthand that humans understand but machines flatten. It may share a name with another entity. It may let old pages linger. It may never say whether it is an operator, reseller, maker, venue or official service. In that case the answer engine has not created the whole problem. It has exposed it.
That is why the lab values the small routine. It is not glamorous. It does not produce a chart. It gives an Italian business a way to watch its name move under language, category, role and source pressure. For answer drift, that is often the first honest instrument.