Observability Shows You What Your AI Did. Verification Proves It Was Right.
Enterprise AI has an observability boom. Every serious AI programme now traces its prompts, tracks token spend, watches latency and runs evaluations on model output. Dashboards light up when something drifts. For engineering and security teams, that visibility is a genuine step forward.
But there is a question no dashboard answers. When an AI tells a customer, a regulator or a court something, was it true? And if someone asks you to prove it a year from now, can you?
Observability tells you what your AI did. It was never designed to prove that what it said was right. For regulated industries, that second job is the one that carries the liability. This article explains what AI observability covers, where it stops, and how VarsaAI's verification layer complements it to close the gap.
What observability is
Observability is the practice of understanding a system from the data it emits. In software, that data comes in three main forms, usually called the three pillars (Elastic):
- Metrics are numeric measurements over time: request counts, error rates, response times.
- Logs are timestamped records of individual events.
- Traces follow a single request as it moves through a distributed system, step by step.
OpenTelemetry, the open standard developed under the Cloud Native Computing Foundation, has become the common way to collect and route these signals to whatever monitoring platform a team prefers (Elastic).
LLM observability applies the same discipline to AI. On top of the usual signals, it captures what is specific to language models: the prompt that went in, the response that came out, the tokens consumed, the cost per call, which model answered, and how long it took. Most platforms add evaluations, automated checks that score responses for quality, safety or relevance. The result is a detailed, searchable record of how an AI application behaves in production.
Where LLM observability stops
The best LLM observability platforms now go beyond telemetry and try to catch hallucinations. Datadog's is a good example, because it is well built and honest about its scope.
Datadog's hallucination evaluation flags any output that disagrees with the context provided to the LLM. It separates two failure types: contradictions, where the response goes directly against the provided context, and unsupported claims, which aren't grounded in it (Datadog docs). It uses an LLM-as-a-judge approach and can flag hallucinations within minutes in production (Datadog).
That is useful. It is also bounded by design, and Datadog says so. Its own engineering blog notes that this kind of faithfulness check assumes the context is correct and accepts it as ground truth, and that verifying the context is an independent problem (Datadog engineering).
Three limits follow from that:
- It checks consistency, not truth. If the retrieved document is out of date or wrong, a response that faithfully repeats it passes the check.
- It only covers instrumented applications. Detection relies on developers annotating their LLM spans with the user query and retrieved context, and does not run when that context is empty (Datadog docs). The workforce using ChatGPT, Claude or Copilot directly is outside its view.
- It produces telemetry, not evidence. A flag on a dashboard helps an engineer fix a pipeline. It is not something you can hand to an auditor as independent proof that a specific claim was checked.
The question observability can't answer
Picture the moment that matters. An auditor, a regulator or opposing counsel points at a sentence your AI generated eighteen months ago and asks two things: was this true, and how do you know?
Your observability stack can tell them a lot. Which model produced the sentence, when, for which user, at what cost and latency. Possibly that an evaluation scored the response as consistent with whatever document was retrieved at the time.
What it can't tell them is whether the claim matched an authoritative source, which source, and which passage. It can't produce a record that a third party can check independently without trusting your own systems. In a regulated industry, that is the answer that counts.
This is the core distinction. Observability measures the system. Verification proves the claim. One is an engineering signal about how your AI behaves. The other is evidence about whether what it said was right. A mature AI programme needs both, and most have only the first.
How VarsaAI complements observability
VarsaAI is AI verification infrastructure. It does not replace your observability stack. It adds the layer observability was never built to provide: independent proof that each claim is true.
- It checks against the outside world, not just your own context. Where faithfulness checks compare a response with the document your pipeline retrieved, VarsaAI verifies claims against authoritative sources: peer-reviewed literature, regulator records, financial filings, case law, legislation and patents, as well as your own approved reference repository. If the retrieved context is wrong, VarsaAI catches what a consistency check would pass.
- It delivers a verdict on every claim. Each claim is marked supported, unsupported or contradicted, and where a source disagrees, the conflicting passage is surfaced. Your teams see exactly which sentence is the problem and why.
- It attributes every sentence to its source. VeriRef™ links each sentence to the exact passage behind it, so a reviewer can check the evidence in seconds rather than re-researching the claim.
- It produces proof, not just telemetry. Every verification carries a cryptographically signed receipt that any third party can verify independently. That is the artefact an auditor can accept, because it doesn't depend on trusting your internal dashboards.
- It covers the workforce, not just the apps you instrument. VarsaAI connects through MCP into the LLMs people already use, including Claude, ChatGPT and enterprise AI gateways. Verification happens at the moment of generation, wherever the output is produced.
- It feeds your existing stack. Verification results export in JSON and CSV, so they can sit alongside your traces and logs. Verification becomes one more signal your teams monitor, and the evidence behind it is kept for audit.
| LLM observability | VarsaAI verification | |
|---|---|---|
| Core question | How is the AI system behaving? | Is this claim true, and can we prove it? |
| Checks output against | The context the app supplied | Authoritative external sources and your approved repository |
| Coverage | Instrumented applications | Workforce LLMs and agents, via MCP |
| Unit of analysis | Requests, spans, responses | Individual claims and sentences |
| Output | Metrics, traces, evaluation scores | Per-claim verdicts, sources and signed receipts |
| Primary user | Engineering and operations | Security, compliance, legal and audit |
One response, both layers
Here is how the two work together on a single piece of AI output. A compliance analyst asks an AI assistant to draft a client briefing on a recent regulatory change.
- The request is observed. The prompt passes through the enterprise AI gateway. Observability records the trace: the user, the model, the tokens, the latency and the documents retrieved.
- The model drafts a response containing five factual claims, grounded in an internal guidance note the pipeline retrieved.
- The observability evaluation passes. All five claims are consistent with the retrieved note, so the faithfulness check raises no flag.
- VarsaAI verifies each claim against authoritative sources. Four are supported, each linked to its source passage. One is contradicted: the regulator updated its guidance after the internal note was written. VarsaAI surfaces the regulator's current wording next to the flagged sentence.
- The analyst fixes the claim before the briefing goes out. A signed receipt records every claim, its verdict and its source.
- Months later, an auditor asks about the briefing. Observability shows who generated it, when and with which model. The VarsaAI receipt shows what each claim said, what it was checked against and the verdict, and the auditor can verify the receipt independently.
Observability confirmed the system worked as designed. VarsaAI caught the one thing the design couldn't: the source itself was out of date.
Measure the system. Prove the claim.
Observability is now table stakes for enterprise AI, and it should be. You can't run what you can't see. But visibility into how a system behaves is not the same as proof that what it said was true, and in regulated industries, liability attaches to what the AI said.
The strongest AI programmes run both layers. Observability keeps the system healthy. Verification makes every output defensible. Together they answer both questions a board, a regulator or a court will ask: what did your AI do, and was it right?
VarsaAI already does this in production at regulated global enterprises. Regeneron's Global IT Lead reports that 80%+ of generated medical claims now carry traceable evidence. Ipsen's Chief AI Officer reports a 75% reduction in repetitive manual checking.
Find out where your AI output is unverified. Take the VarsaAI AI Risk Assessment or book a demo to see verification running inside the LLMs your teams already use.
Sources
- Elastic, What is OpenTelemetry?
- Datadog, Hallucination evaluation documentation
- Datadog, Detect hallucinations in your RAG LLM applications with Datadog LLM Observability
- Datadog engineering, How we built LLM hallucination detection
- Customer results (Regeneron, Ipsen): VarsaAI customer testimonials
See where your AI output is unverified
Take the VarsaAI AI Risk Assessment, or book a demo to see verification running inside the LLMs your teams already use.
Take the AI Risk Assessment →