The Shadow AI Gap: Why Shadow AI Tools Leave Regulated Enterprises Exposed
Executive summary
Shadow AI tools make AI usage visible. They do not make it true. Every shadow AI product on the market answers two questions: who is using AI and what data is going in. Regulated enterprises are liable for a third question none of them answer: is what came out true, and can we prove it?
That third question is where the real exposure sits. When Deloitte Australia refunded part of a A$440,000 government report full of fabricated references, the AI involved was Azure OpenAI, a sanctioned enterprise tool, not shadow AI. Courts and tribunals are now holding companies liable for what their AI says, regardless of which tool said it.
The gap gets bigger as shadow AI programmes succeed. Moving the workforce from unsanctioned tools onto approved ones multiplies the volume of AI output going out under the company's name, and none of it is checked.
VarsaAI is the output layer. It verifies every claim at the moment of generation, attributes each sentence to an authoritative source, and issues signed, third-party-verifiable proof. It sits alongside the existing shadow AI stack and replaces nothing in it.
The positioning: shadow AI tools govern the input. VarsaAI proves the output. A shadow AI programme without output verification controls access and leakage, but leaves the enterprise's largest liability — what its AI actually says — completely ungoverned.
The confusion in the market
Buyers say "we've already deployed shadow AI tooling, so we're covered." The belief is reasonable, which is why it is dangerous. They have bought a real control for a real risk, and the vendor's own language ("AI governance", "secure AI", "AI visibility") sounds like it covers everything.
The risk they bought for is real. Gartner predicts that by 2030 more than 40% of organisations will suffer security and compliance incidents from unauthorised AI tools, and found 69% of organisations already have evidence or suspicion of it. Netskope estimates more than half of all current AI app adoption is shadow AI.
Where the belief breaks is in what "covered" means. Shadow AI tools were built by the data security market, so they solve data security problems: leakage, unauthorised apps, sensitive prompts. The question a regulator, court or auditor asks is different: "Your AI said this. Was it true, and where did it come from?"
The reframe: "Your shadow AI tooling tells you who used AI. Who tells you whether it was right?"
What shadow AI tools actually do
The market splits into four categories. All four are useful, and all four stop at the same line: they see, block or log AI output, but none judge whether it is true against authoritative sources.
| Category | What it does | Where it stops |
|---|---|---|
| Security service edge / DLP | Discovers AI apps in use, analyses usage, coaches users in real time about data exposure | Watches traffic and data leaving; does not read whether a response is accurate |
| Data security posture for AI | Discovers shadow AI, applies DLP to prompts, logs prompts and responses for investigation | Logs the response; the log is a record of what was said, not a judgement on whether it was right |
| Secure AI gateway | Governed access to GPT, Claude, Gemini; RBAC; PII redaction; prompt monitoring | Monitoring is based on the prompt submitted, not the AI's response |
| Developer guardrails | Checks whether an app's responses match the source documents the developer supplies | Only as good as the documents fed in; built into custom apps, not the workforce's chat tools; no signed evidence |
| Capability | SSE / DLP | DSPM for AI | Secure AI gateway | Developer guardrails | VarsaAI |
|---|---|---|---|---|---|
| Discover which AI tools staff use | Yes | Yes | Partial (own platform) | No | No |
| Block or redact sensitive data in prompts | Yes | Yes | Yes | No | No |
| Log prompts and responses | Partial | Yes | Yes | No | Yes |
| Verdict on whether each claim is true | No | No | No | Partial (vs supplied docs) | Yes |
| Check claims against external authoritative sources | No | No | No | No | Yes |
| Attribute each sentence to its source | No | No | No | No | Yes |
| Signed, third-party-verifiable proof per claim | No | No | No | No | Yes |
| Org-wide library of verified claims and sources | No | No | No | No | Yes |
The top three rows are where shadow AI tools compete with each other. The bottom five are empty across the category, and that empty space is VarsaAI's market.
Why the gap is a liability, not a nice-to-have
The precedent is set: organisations are liable for what their AI says, and "the AI made it up" is not a defence. In each case below, shadow AI tooling would not have changed the outcome, because the tool was known, approved or the company's own.
| Case | What happened | What it proves |
|---|---|---|
| Moffatt v. Air Canada, 2024 BCCRT 149 (Feb 2024) | Airline's own chatbot gave a customer wrong refund policy; tribunal rejected the argument that the chatbot was a separate entity and found the airline did not take reasonable care to ensure its accuracy | The company owns its AI's statements. The duty is reasonable care over accuracy, not over access |
| Deloitte Australia government report (Oct 2025) | Partial refund on a A$440,000 report containing a fabricated quote from a federal court judgment and references to non-existent papers; the revised report disclosed Azure OpenAI was used | Sanctioned enterprise AI hallucinates too. The financial and reputational hit landed on the firm, not the model |
| Court sanctions for AI-fabricated citations (to June 2026) | Charlotin's database recorded 1,058 US cases of improper GenAI use in court filings by the start of June 2026 | Hallucination is not an edge case. It is now routine enough to be tracked as its own category of misconduct |
The pattern: every one of these is an output failure. None was a data leak, an unsanctioned app or a risky prompt. A fully deployed shadow AI stack would have logged each one perfectly and prevented none of them.
Why output verification is necessary, not complementary
"Complementary" means nice to add. "Necessary" means the control framework is incomplete without it. The argument holds up step by step:
- Liability attaches to output. The duty in Moffatt was reasonable care to ensure the AI was accurate. Access controls and DLP do not discharge that duty.
- Sanctioned AI still produces false output. Deloitte's errors came from Azure OpenAI, an approved enterprise tool. Moving staff from shadow to sanctioned AI changes who hosts the model, not whether it hallucinates.
- Success makes the gap bigger. Every shadow AI programme aims to move usage onto approved tools at scale. That multiplies the volume of AI output carrying the company's name, all of it unverified.
- A log is not a control. Purview, gateways and SIEMs record every response. When a response is false, that record proves the company published it and did nothing to check it. Logging without verification turns the audit trail into evidence against you.
- No other layer can close it. The matrix above shows the output rows are empty across every shadow AI category. If the control doesn't exist in the stack, the stack doesn't meet the duty.
The conclusion: a shadow AI programme without output verification is an incomplete control. It governs who used AI and what went in, and leaves ungoverned the one thing the company is legally accountable for.
How VarsaAI plugs in
VarsaAI sits after the model and before the outside world, connected by MCP into the LLMs already in use. Nothing in the existing shadow AI stack is removed or reconfigured. The flow runs top to bottom:
| Layer | What it answers | Components |
|---|---|---|
| Employees and agents | Where AI use starts | Workforce chat, apps, agentic workflows |
| Input controls (your existing shadow AI stack) | Who used AI, and what data went in | Discover AI apps (SSE / DLP, e.g. Netskope); block sensitive data (DSPM for AI, e.g. Purview); govern model access (secure gateway, e.g. AI CTRL) |
| LLMs | Generates the output | GPT, Claude, Gemini, Copilot |
| Output verification (VarsaAI, connected via MCP) | Is each claim true, and can we prove it | Verdict on every claim (supported or contradicted); source for every sentence (linked to the exact passage); signed, verifiable proof (third-party verifiable receipt) |
| Customers, regulators, auditors, courts | Where liability lands | Receives only verified, evidenced output |
Everything above the model governs who uses AI and what goes in. Everything that reaches a customer, regulator, auditor or court passes through the output layer first.
The business case
The case rests on two lines of value: labour recovered from manual checking today, and incident exposure avoided tomorrow. The labour line alone usually pays for the platform; the incident line is the reason the board cares.
Line 1: labour recovered. Regulated firms already pay people to check AI output by hand, or they don't check it at all. Ipsen's Chief AI Officer reports a 75% reduction in repetitive manual checking with VarsaAI as the LLM trust layer.
Worked example, using illustrative assumptions:
| Input | Illustrative value | Basis |
|---|---|---|
| Staff producing external-facing AI-assisted content | 500 | Assumption: ~2.5% of a 20,000-person enterprise |
| AI-assisted outputs checked per person per week | 5 | Assumption |
| Manual check time per output | 15 min | Assumption |
| Reduction in manual checking | 75% | Ipsen customer result |
| Loaded cost per reviewer hour | £75 | Assumption |
| Hours recovered per year | ~24,400 | 500 × 5 × 0.25 h × 75% × 52 |
| Labour value recovered per year | ~£1.8M | 24,400 h × £75 |
Line 2: incident exposure avoided. This is your own number, built from your worst plausible output failure: a fabricated claim in a regulatory submission, a client deliverable, a court filing or a customer-facing answer. Deloitte shows the shape of the cost: a refund, a public correction and a senator calling for the full fee back. What would one of those cost you?
Line 3: what you can now prove. Regeneron's Global IT Lead reports 80%+ of generated medical claims now carry traceable evidence. That converts an unknowable liability into a measurable control the CISO can report to the board and the auditor.
Common objections, answered
| They say | We say |
|---|---|
| "We've deployed Netskope / Purview / a secure gateway. We're covered." | You're covered for who uses AI and what data goes in. Which tool in that stack tells you whether a response was true? |
| "Purview logs every prompt and response." | It does, and that's the risk. A log of a false claim proves you published it without checking. We turn the log into proof you did check. |
| "We have groundedness checks in our apps." | Those check answers against the documents you feed in, inside the apps you build. They don't cover the workforce's chat tools, don't check regulators or literature, and don't leave signed proof. |
| "We only allow approved AI now." | Deloitte's errors came from Azure OpenAI, an approved tool. Approved changes the host, not the hallucination rate. |
| "Our people check AI output before it goes out." | Then you're paying for it. Ipsen cut repetitive manual checking 75%. And how do you prove the check happened? |
| "It's another tool to deploy." | It's an MCP connector into the LLMs you already run. No re-platforming, no new interface for staff. |
| "This is a nice-to-have for next year." | The duty of reasonable care over accuracy exists today (Moffatt, Feb 2024). Every unverified output between now and then is exposure you've chosen to carry. |
Owning the category: AI Output Verification
The category name: AI Output Verification. The enemy is not shadow AI vendors; it is the belief that input controls are enough.
Positioning line: Shadow AI tools govern the input. VarsaAI proves the output.
Category definition: AI Output Verification is the control layer that checks every AI-generated claim against authoritative sources at the moment of generation, attributes it to its source, and produces independently verifiable proof.
| Pillar | Message | Proof point |
|---|---|---|
| The shadow AI gap | Every shadow AI tool stops at the response. Liability starts there. | Capability matrix: output rows empty across SSE, DSPM, gateways |
| Approved isn't accurate | Sanctioned AI hallucinates too. | Deloitte refund; Azure OpenAI disclosed in the revised report |
| You own what your AI says | Courts don't accept "the AI did it". | Moffatt v. Air Canada; 1,058 US court cases to June 2026 |
| Proof, not logs | A log records what was said. A signed receipt proves it was checked. | Ed25519-signed, third-party-verifiable receipts |
| Plugs in, replaces nothing | One MCP connector across every LLM already in the stack. | Native MCP fit with Expedient AI CTRL, Claude, ChatGPT |
| Already proven | Live in regulated global enterprises. | Regeneron 80%+ traceable claims; Ipsen 75% less manual checking |
Sources
- Gartner shadow AI prediction: Infosecurity Magazine; Freevacy
- Netskope shadow AI research: Netskope press release; Netskope generative AI controls
- Microsoft Purview DSPM for AI: Microsoft Tech Community
- Azure AI Content Safety groundedness: Microsoft Learn
- Expedient AI CTRL: service overview; prompt monitoring; MCP integration
- Moffatt v. Air Canada: McCarthy Tétrault; American Bar Association
- Deloitte Australia refund: AP via CTV News
- Court hallucination cases: UNC Law Library
- Customer results (Regeneron, Ipsen): VarsaAI customer testimonials
Close the shadow AI gap
See how VarsaAI verifies every claim your AI makes — and proves it — without replacing anything in your existing stack.
Take the AI Risk Assessment →