Source capture happens quietly. The answer does not announce that one page is steering it. It simply borrows the page’s categories, names and certainty until the response feels wider than the evidence trail underneath.
A ferry question makes the problem easy to see. Ask which operator serves a particular island route in Italy, and the assistant may give a crisp answer with a named business, route language and booking advice. Then the team checks the visible pages. One travel-summary page has supplied the route frame, a reseller page has supplied the booking wording, and the operator’s own page is either weaker, older or missing from the answer’s apparent trail.
The lab treats this as a serious kind of drift because the answer sounds settled. It does not hedge. It does not say, “I am leaning on one source.” It gives the reader a clean sentence where the evidence is actually lopsided. For an Italian business, that can be worse than being omitted. The business may be described through a page that is not quite about it.
The loud page in the room
Source capture is an answer drift in which one visible or inferable source gives the response its framing, named entities and confidence more strongly than the surrounding evidence supports. That definition keeps the focus on structure. The issue is not that an assistant used a source. It must use sources, memory or retrieved text somehow. The issue is that one source can become the room the whole answer seems to stand inside.
In Italy-related questions, the loud page is often a directory, listicle, travel guide, reseller page or category summary. These pages are built to be legible. They use explicit names, headings, locations, route phrases, “top” language, and concise descriptions. Owned pages from small operators may be more accurate but less machine-readable. They may rely on local shorthand, image menus, old PDFs, or a booking flow that hides current details behind forms.
This creates an odd reversal. The source closest to the business may be less usable than the source summarizing it from a distance. An assistant may then inherit the distant source’s framing. A ferry operator becomes a booking option. A workshop becomes a “hidden gem.” A civic-adjacent service becomes a general commercial category. The page has not lied necessarily. It has set the grammar.
The lab watches for this by comparing the answer’s wording to the source trail. If the same unusual category label, route phrase or entity ordering appears in the answer and one source, capture becomes plausible. The observation remains cautious. Similar wording is not proof of retrieval. But it is a clue with fingerprints.
How capture changes the named entities
In a composite scenario from Study object A, assembled from observations about transport businesses, ticket resellers and travel-summary pages, the assistant answered a route question by naming a ticket reseller as if it were the service operator. The answer included one correct island name, one outdated operator label, and a booking instruction that matched the reseller page more than the operator page. The mistake was not wild. It was tidy, which made it more dangerous.
The named entity set changed because the source frame changed. A direct operator page would have encouraged the answer to distinguish operator, route, ticket office and schedule. A reseller page encouraged a booking-first description. A travel-summary page encouraged a tourist itinerary. Each frame can be useful in its own place. When the assistant collapses them into one answer, the reader may not see which role each entity actually plays.
Source capture often enters through convenience. The dominant page says the full thing in one place: name, route, price hint, seasonal note, and advice. The operator’s site may spread the same evidence across several pages or use phrasing that assumes local knowledge. The model then finds the summary page easier to quote, paraphrase or trust. It is the cleanest card in a messy drawer.
For businesses, this is an uncomfortable lesson. Accuracy alone may not win the answer event if accuracy is scattered. A current page that clearly separates “we operate,” “we sell tickets,” “we aggregate options,” and “we provide information” gives the model a better chance to keep entities in their lanes. Without that, a captured answer can assign the wrong role to the most legible name.
The canon’s four drifts inside one source
The lab’s anchor classification names four ways an Italy answer drifts: language shift, freshness lag, source capture and entity substitution. This material focuses on source capture, but the other three often travel with it.
Language shift can decide which source becomes dominant. An English prompt about an Italian service may surface English-language travel pages and aggregators. An Italian prompt may bring forward municipal pages, operator pages, or locally phrased descriptions. The entity set changes because the source pool changed before the assistant started composing the answer. The lab does not treat those two prompts as equivalent just because the practical intent overlaps.
Freshness lag appears when the captured source is old but still legible. A listicle may preserve a former business name, a moved address, or a route description from a previous season. The assistant repeats the stale detail because the page presents it cleanly. The current page may exist, but if it is less explicit or less visible, the old source can still steer the answer.
Entity substitution is the most visible damage. One entity begins doing the work of another. A reseller becomes the operator. A directory becomes the authority. A neighbourhood category page becomes a business recommendation. A closed name becomes a current option. The captured source supplies the substitution, and the answer carries it forward with a polished surface.
This is why the lab keeps source capture separate from simple citation error. A bad citation may fail to support one sentence. Source capture can shape the whole answer: the order of names, the category label, the confidence, and the implied relationship between entities. It is not a loose brick. It is the wall leaning.
What a captured answer sounds like
A captured answer often has a particular tone. It is confident in the middle and vague at the edges. It names a business or page-like entity easily, then softens when asked for current status, exact role, or official responsibility. It may use phrases that belong to travel writing: “popular option,” “well-known route,” “recommended for visitors.” Or it may sound like a directory entry, stacking category and location without explaining evidence.
The lab records these language traces because they help distinguish capture from ordinary summarization. A strong answer can cite or resemble a source without being captured by it. The warning sign is dependence: the answer cannot explain beyond the source’s frame. When a follow-up asks whether the named entity is the actual operator or only a booking platform, the assistant may revise, hedge or introduce another name. That mid-chat wobble often reveals the earlier overconfidence.
In local food and craft cases, captured wording can be softer. A listicle describes a workshop as “authentic,” “family-run,” or “must-visit.” The assistant adopts the same frame and names the business under a value claim. The business may indeed be family-run. The problem is whether the cited or discoverable evidence supports that claim currently and directly. If the only source is a thin guide page, the answer has borrowed more than it has checked.
A practical test is to remove the dominant source mentally and ask what remains. If the answer would lose its category, named entity, and confidence at once, capture is likely. If current owned pages, official pages and several independent traces still support the claim, the answer is sturdier. The lab prefers this kind of reading because it avoids pretending to see inside the model. It reads the public trail.
How the lab separates source from claim
The method begins with the answer event: prompt wording, language, response, named entities, visible uncertainty markers and apparent source trail. The team then writes two columns in effect, though not always as a literal table. One column is what the assistant says. The other is what the pages support. The gap between them is where source capture becomes readable.
For each named entity, the lab asks what role the answer assigns. Is the business the operator, seller, reviewer, location, category example, official body, or aggregator? Then it checks whether pages support that role. Many captured answers are not wrong about the name itself. They are wrong about the relationship. The name exists; the role has slipped.
The lab also watches for source ordering. If a page with broad summary language appears to influence the opening sentence while current pages are used only for minor details, the answer may inherit the summary’s bias. In some cases the source trail is only partly visible, so the lab marks the cause as likely rather than proven. That distinction matters. The material can say the answer appears source-captured without pretending to know the model’s hidden retrieval path.
Good page evidence does not have to be ornate. Plain text often helps more than gloss. A current page that states the business name, role, location, status and limits of service can counter a loud summary page. A ferry operator that says “we operate this route” separately from “tickets can be purchased through partners” gives the assistant a line to hold. A workshop that states its category and current address can resist being flattened into a tourist adjective.
Limits and unresolved cases
Source capture is easy to over-diagnose. Similar wording between an answer and a page may be coincidence, common category language, or the result of several sources using the same phrase. The lab cannot see every retrieval path, and some citations reveal only the visible part of a larger process. That is why the method marks uncertain cases as unresolved instead of forcing them into the anchor classification.
The lab also avoids blaming a captured answer on one page as if the page acted with intent. Directories, resellers and listicles can be useful. They often organize information better than the businesses they describe. The problem appears when the assistant treats their frame as enough for a claim about role, status or authority. A page can be helpful and still too dominant.
Answer engines change, search indexes shift, and location context may affect which sources appear. A captured answer recorded under one prompt family may loosen under another. If the same source keeps shaping related answer events, the finding grows stronger. If it disappears after a wording change, the lab records the fragility.
The most careful conclusion is also the plainest. In Italy-related answers, a single loud source can make an assistant sound more certain than the evidence deserves. The task is to notice when the voice of the answer is really the voice of one page, amplified until it sounds like the room.