New-ZZZ
RU / EN
29 September 2026

Getting the Source Right, Not Just the Fact: Source-Aware Verification for MCP Agents

N
New-ZZZ desk
Hugging Face Blog · 20 hours ago

The paper introduces ProvenanceGuard, a sophisticated and critical post-generation verification layer designed specifically to address a major vulnerability in advanced Large Language Model (LLM) agents, particularly those operating in complex, data-sensitive environments. This vulnerability is termed 'cross-source conflation,' which represents a failure mode where an agent generates a claim that is factually correct, but incorrectly attributes that fact to a source that did not provide it. Traditional, 'source-blind' verifiers are insufficient because they only check for the mere existence of a fact within the pooled evidence, thereby passing misleading answers. ProvenanceGuard, however, is built to maintain and verify the precise connection between every generated claim and its originating source, making it indispensable for high-stakes applications.

Consider the practical implications of source attribution failure. In a customer support setting, an agent might correctly state that a '30-day refund window' exists, but if the answer implies this fact comes from the customer's specific 'account record' when it actually originates from a general 'policy document,' the attribution is wrong. In such a data-sensitive context, this wrong attribution can be as damaging, if not more so, than simply stating an incorrect fact. A similar, highly critical pattern emerges in clinical settings: a patient-specific medication detail, correctly retrieved from a patient-history tool, becomes misleading if the agent presents it as a finding derived from general medical literature. ProvenanceGuard directly tackles this by ensuring that the supporting source explicitly matches the source the answer claims or implies.

Because of this critical need for source fidelity, the authors argue that existing 'faithfulness scores' are inadequate for Multi-Component Process (MCP) agents. An answer generated by an MCP agent inherently carries 'provenance'—this is the record of where the information came from, which can be explicitly stated (e.g., 'according to the account record') or implicitly understood. ProvenanceGuard’s core function is to keep this vital connection between the claim and its source available for rigorous inspection throughout the entire AI pipeline. It does this by acting as a verification layer that sits atop a black-box MCP agent, meaning it does not require retraining of the underlying agent model.

When an agent produces an answer, ProvenanceGuard intercepts the process. Crucially, it never collapses the evidence into one anonymous context; instead, it meticulously carries the source identity through every stage. It processes the captured MCP trace, which includes detailed outputs from various tools and their corresponding unique source IDs. The verification process is highly structured and sequential, involving five distinct steps: First, the system breaks the complex answer down into discrete, verifiable claims. Second, it identifies the source material most relevant to each individual claim. Third, it performs a support check to verify whether that identified source actually supports the claim. Fourth, and most critically, it compares the supporting source against the source that the answer explicitly names or implicitly suggests. Finally, it emits two types of verdicts: a per-claim source verdict, and a global, answer-level allow or block decision. This structured flow ensures that source identity is preserved through decomposition, routing, support scoring, attribution checking, and potential repair, rather than being lost or pooled.

For its experimental setup, the researchers utilized local models to ensure a controlled, offline environment for processing the captured traces. This setup involved MiniLM for the task of finding the most relevant source, a DeBERTa Natural Language Inference (NLI) verifier model to check if the source supported the claim, and a local language model to perform the initial claim decomposition. Furthermore, the verifier incorporates a strict check for literal values—meaning any number, date, or identifier mentioned in the answer must be physically present in the source material, preventing plausible-sounding but unsupported claims. The entire system culminates in a calibrated decision step that combines all these signals. If an answer is blocked, a sophisticated RARR-style repair mechanism can attempt a source-grounded revision or a safe fallback, which is then subjected to the verifier's checks once more, creating a robust feedback loop.

While the named models were used for the initial evaluation, the design philosophy of ProvenanceGuard is modular. The core claim, source, and decision steps can be adapted for cloud-hosted models, provided a new setup undergoes its own rigorous testing and calibration. The reported results, however, stem from the local configuration, which is noted for its conservative decision policy. This caution is highly desirable in data-sensitive review settings, where the priority is ensuring the source is correct, even if it means sacrificing the speed of the generated answer. The system was tested using a medical agent that accessed diverse tools, including patient records and general research articles, providing 281 real traces for study. Medicine serves as an ideal test case because a fact derived from a patient's unique record cannot be treated identically to a fact derived from general academic research. The main test involved human experts reviewing 361 claims drawn from 40 answers that were entirely set aside from the system's development data. The results were highly compelling: the experts determined that 139 claims should have been blocked, and ProvenanceGuard successfully caught 138 of them. Moreover, the system demonstrated a high degree of caution by flagging 67 claims that human experts considered supported, sending them for mandatory review or repair. This performance underscores the system's commitment to safety. For claims where a source was identifiable, ProvenanceGuard correctly identified the source approximately 86% of the time in this test. When compared against four other support checkers, ProvenanceGuard achieved the highest score on the paper's metric for catching claims that should have been blocked while maintaining a high rate of accuracy for supported claims.

Why it matters

  • —(демо) Почему это важно: ключевой эффект для рынка ИИ.
  • —(демо) На кого и как это повлияет в ближайшее время.

Key facts

  • (демо) Первый ключевой факт из новости.
  • (демо) Второй важный факт.
  • (демо) Третья деталь, влияющая на выводы.
Read the original →

The full text is in the original source. Here we provide a brief summary and key facts.