Shadow AI Risk Assessment: A Practical Framework to Discover, Score, and Govern Unauthorized AI Use at Scale

Key takeaways
- A shadow AI risk assessment evaluates the business use case alongside the tool. The same model can carry low or critical risk depending on the data and decision involved.
- Discovery should combine network, browser, endpoint, OAuth, and SaaS telemetry with procurement records, expense data, and employee reporting, because no single source reveals the full estate.
- Scoring multiplies impact by probability on a 1 to 5 scale. Data sensitivity, autonomy, permissions, and risk velocity then appear as separate overlays instead of folding into one number.
- Treatment falls into four dispositions: ban, sandbox, monitor, or formally adopt, each with an accountable owner, control evidence, and a reassessment trigger.
- Approved alternatives and role-based training reduce exposure more reliably than prohibition, because employees return to unmonitored accounts when a policy blocks useful work.
A shadow AI risk assessment gives security teams a structured way to find unauthorized AI use, measure its exposure, and protect sensitive data without blocking productive innovation. It evaluates both the technology and the business purpose behind public chatbots, embedded features, AI agents, APIs, open-source models, and locally hosted systems.
This guide shows security, IT, privacy, compliance, procurement, and business leaders how to build a living inventory, review data and vendor controls, and assign accountable owners. It also connects discovery to proportionate decisions, including whether to contain, sandbox, monitor, or formally adopt a tool.
The framework applies a 1 to 5 scale to probability and impact, then separates inherent risk from residual risk. Risk velocity, user count, autonomy, and regulatory exposure sharpen the resulting priorities.
The sections below cover practical detection methods, evidence requirements, governance roles, employee guardrails, and measures that show whether controls reduce exposure while approved AI use creates value. Applied consistently, the process makes hidden AI activity visible and keeps risk ratings current as tools, models, and use cases change.
Security leaders who want to see how behavioral signals and targeted training support this work can book a demo of Adaptive Security.

What Is a Shadow AI Risk Assessment?
A shadow AI risk assessment is a structured review of unsanctioned or insufficiently governed artificial intelligence use across an organization. It identifies the tools, models, data flows, users and business purposes involved, then determines whether each use creates unacceptable security, privacy, compliance, operational or trust exposure.
The assessment avoids judging employees for adopting new technology. Its purpose is to make useful AI activity visible, governed and accountable.
What Does a Shadow AI Risk Assessment Cover?
Shadow AI includes any generative AI tool, AI agent, embedded AI feature, application programming interface, open-source model or locally hosted model used without sufficient organizational oversight. A public chatbot used to rewrite a customer email belongs in scope, as does an AI feature quietly enabled inside an approved software platform.
Risk depends on the data entered, the decision influenced and the permissions granted rather than on the hosting arrangement.
A practical assessment examines both the tool and the business use case. A public chatbot used to brainstorm an internal agenda presents a different exposure from the same chatbot receiving unreleased financial results.
An AI résumé screener creates different concerns when it ranks candidates than when it summarizes applications for a human reviewer. A customer-service bot carries a different risk when it answers general questions than when it can issue refunds, retrieve account data or change a customer record.
The assessment should document these dimensions for every discovered use:
- Tool and model: Identify the provider, model, version, account owner, browser extension, API or local deployment.
- Data exposure: Record whether users submit personal information, intellectual property, credentials, regulated records, source code or confidential business plans.
- Business action: Determine whether the system drafts content, recommends decisions, makes decisions, accesses systems or takes autonomous action.
- Access and retention: Review permissions, connected applications, logging, training-data practices, retention periods and administrative ownership.
- Human oversight: Establish who verifies outputs, handles errors, approves consequential decisions and reports misuse.
- Residual risk: Decide whether to approve, restrict, redesign, replace or retire the use after controls are applied.
An application inventory alone cannot show what the organization is risking. Two departments can use the same model with completely different consequences. Security leaders need a use-case register that connects each AI activity to its data, workflow, business owner and human decision point.
The assessment lifecycle moves from discovery to validation, classification, treatment and residual-risk review. Discovery uses approved software inventories, identity records, browser and API telemetry, procurement data, expense reports, developer repositories and employee interviews.
Open-source intelligence (OSINT), meaning information gathered from publicly available sources, can add context about exposed employee roles, public code repositories, published prompts or vendor relationships. It should support governance rather than employee surveillance.
Validation confirms that a detected application is genuinely used, who controls it and what information enters or leaves it. Classification rates the use according to data sensitivity, model behavior, autonomy, regulatory impact, third-party dependency and potential harm. Treatment can include an approved enterprise alternative, access restrictions, data-loss controls, contractual review, prompt-handling rules, human approval or targeted training.
Residual-risk review asks whether the remaining exposure is acceptable for the specific business purpose. The NIST Generative Artificial Intelligence Profile, published in 2024, frames AI risk management around identifying and managing risks across the system lifecycle. That approach prevents a common governance failure: approving a tool once while ignoring later changes to its model, integrations, permissions, data practices or business role.
How Is Shadow AI Different From Shadow IT?
Shadow IT is the broader practice of using unauthorized hardware, software, cloud services or systems outside formal IT processes. Shadow AI is a specialized category within shadow IT because AI can generate content, infer sensitive information, make recommendations, automate decisions or act through connected tools.
An unauthorized project-management application might expose task data through weak access controls. An unauthorized AI assistant can expose the same data while also summarizing it, retaining it, using it to produce new content or transmitting it to a third-party model. An AI agent adds another layer because it can interpret instructions and take actions across connected systems.
The assessment must therefore examine more than which applications are present and who can access them. It must also track the prompts, files, records, outputs, recommendations and automated actions moving through each tool. Model behavior, data provenance, hallucination risk, bias, explainability, vendor terms and the consequences of an incorrect output all belong in the review.
The human layer connects both categories. Employees adopt tools to complete legitimate work under time pressure, support customers, meet sales deadlines or solve technical problems. Treating that behavior as misconduct drives use underground and removes the signals needed to reduce risk.
A constructive program provides approved tools, clear data-handling rules, practical examples and a reporting channel for uncertain cases. Training builds judgment around what employees can enter, what they must verify and when a person must stop an automated action.
What Are Common Workplace Examples and Adoption Drivers?
Shadow AI appears wherever employees can access a capable model faster than they can obtain an approved workflow. Common examples include public chatbots used for drafting or summarization, AI meeting assistants that record confidential discussions and résumé-screening systems that rank applicants.
The list also covers customer-service bots, marketing automation, predictive analytics embedded in spreadsheets and AI-generated code copied into production systems.
The same pattern extends through less visible channels. Employees install browser extensions that summarize webpages, activate vendor-provided AI features inside customer relationship or productivity platforms and call external APIs from scripts.
Others download open-source models, run models locally on workstations or purchase subscriptions through corporate cards without consulting procurement. A decentralized purchase by one team can become a business-critical dependency before security, privacy, legal or compliance teams know it exists.
IBM's 2025 survey of 1,000 American office workers found that 80% used AI in their roles, while only 22% relied exclusively on employer-provided tools. The same research found that 97% believed AI improved productivity, which explains why prohibition-only policies fail.
Employees who experience a direct benefit will seek alternatives when approved tools are unavailable, difficult to use or less capable. These figures appear in IBM's 2025 analysis of rising AI adoption and shadow risk.
That productivity signal should guide the response. Security leaders should identify the need an unauthorized tool is meeting and address it without normalizing uncontrolled data use. If employees need transcription, provide a reviewed meeting assistant. If developers need code support, establish an approved environment with repository controls and review gates. If marketing teams need rapid content creation, define permitted data classes, human approval and brand safeguards.
Behavioral visibility completes the assessment. A team that repeatedly pastes customer records into a public chatbot needs targeted instruction, an approved workflow and a way to report near misses without fear of blame.
Human-layer signals such as risky browser behavior, unauthorized AI use and repeated policy exceptions should inform role-specific training and risk review alongside phishing, vishing and smishing behavior. Human risk management connects those behaviors to accountable remediation and treats employees as trainable defenders.
A shadow AI risk assessment succeeds when it turns an invisible productivity workaround into a documented business decision. That decision becomes actionable only when security, privacy, compliance and business owners agree on the exposure they will accept and the controls that must remain in place.
What Are the Main Shadow AI Risks for Security, Privacy, Compliance, and Trust?
Shadow AI risk assessment identifies what happens when employees use artificial intelligence tools without approved data controls, model reviews, or accountable owners. That gap creates an unmanaged path from internal work to external models, storage systems, vendors, and downstream decisions.
The 2025 ISACA analysis of shadow AI in the enterprise describes shadow AI as unauthorized technology use that bypasses formal IT governance. A practical response discovers usage, classifies risk, and gives employees a safe approved path.
How Does Shadow AI Expose Data and Create Privacy Risk?
Data exposure ranks first among the shadow AI risks to test, because employees often paste information into a prompt before considering where it will be stored, processed, or reused.
Sensitive material can include customer records, personal information, health details, payment data, legal advice, source code, product plans, trade secrets, credentials, and intellectual property. A prompt can become an input record, prompt history, support ticket, telemetry event, or material available to a third-party provider under terms the organization has not reviewed.
A proper shadow AI risk assessment maps the complete data path rather than inspecting only the visible chatbot. Follow information from the browser prompt to the model provider, plug-ins, application programming interfaces, retrieval-augmented generation systems, embeddings, vector databases, logs, backups, model-training processes, and human reviewers. A document that appears free of sensitive data can become identifying when combined with an employee name, customer account number, project code, or internal URL.
Require data classification before submission, block regulated and confidential fields in unapproved tools, and approve only providers with contractual limits on retention, training use, access, deletion, and cross-border transfers. Security leaders can extend this human risk management approach to AI use by treating risky prompts, unauthorized applications, and personal accounts as behavioral signals that require targeted action.
Retrieval-augmented generation creates another exposure point because a model can retrieve internal documents that the requesting user should not see. Weak access controls turn a helpful assistant into a search interface that bypasses existing permissions.
Enforce identity-aware retrieval, document-level authorization, tenant isolation, encryption, retention limits, and logging of both the question and retrieved sources. Test whether a user can extract another department's records through carefully worded prompts before deploying an internal AI assistant.
Third-party sharing expands the blast radius. An employee might connect a browser extension, upload a spreadsheet to a free service, or authorize an AI plug-in to read cloud files without understanding the permissions granted.
Maintain an inventory of AI applications, review OAuth scopes and extensions, and prohibit personal accounts for company work. Give employees an approved catalog that is easier to use than an unauthorized alternative. An AI ban without a usable replacement drives activity into personal accounts, unmanaged browsers, and harder-to-observe channels.
What Operational and Technical Failures Can Shadow AI Cause?
Operational risk begins when an AI tool enters a business process without an owner, service expectation, validation method, or exit plan. An internal experiment that summarizes meeting notes has a limited blast radius when a person checks every output.
The same model embedded in invoice processing, customer support, recruiting, underwriting, clinical administration, or legal workflows carries materially higher risk. An error there can trigger a financial, employment, health, or legal consequence.
The risk assessment should separate three use cases:
- Internal experimentation: Require data boundaries, approved accounts, and human review of every output.
- Internal automation: Require testing, access controls, rollback procedures, monitoring, and a named process owner because the model can act repeatedly at machine speed.
- Customer-facing or high-impact use: Require documented validation, decision-appropriate explainability, human appeal, bias testing, incident escalation, and periodic reapproval.
Hiring, lending, insurance, health care, and legal advice deserve the strictest controls because a plausible but wrong answer can affect a person’s livelihood, access to care, coverage, or rights.
AI-generated code creates related technical and legal risk. Code assistants can produce insecure authentication logic, vulnerable dependencies, hard-coded secrets, or license-incompatible material.
Developers should treat generated code as untrusted third-party code, require peer review, and run static and dynamic testing. They should also scan dependencies and licenses, then record which model produced material that enters a production system. Coding assistants belong inside the same review pipeline used for every other code contribution.
Technical failure also appears inside the model itself. Hallucinations create confident but false outputs. Bias can reproduce unequal treatment from training data or skewed business records. Prompt injection can cause a model to ignore system instructions when malicious text appears in an uploaded file, webpage, email, or retrieved document. Model manipulation and data poisoning can corrupt outputs over time.
Apply input filtering, isolated execution, least-privilege tool access, adversarial testing, source citations, output validation, rate limits, and human approval for consequential actions.
Business continuity deserves equal attention. An unplanned AI rollout can generate unexpected usage charges, overload data pipelines, or consume application programming interface quotas. It can also create a single point of failure when a provider changes its model, pricing, availability, or safety behavior.
Establish budgets and alerts, apply quotas, and maintain fallback procedures. Version prompts and models, test provider changes before release, and ensure employees can complete critical work without one model or vendor.
A practical control set includes:
- Discover: Monitor approved and unapproved AI applications, browser extensions, API connections, and data transfers.
- Classify: Assign each use case a risk level based on data sensitivity, autonomy, affected people, and business consequence.
- Control: Apply identity, access, retention, network, browser, and provider-policy controls that match the risk.
- Validate: Test accuracy, bias, prompt injection, data leakage, security, licensing, and failure recovery before release.
- Review: Reassess models, prompts, vendors, permissions, costs, and outputs throughout the system lifecycle.
Weak model transparency makes these controls harder to enforce. If a team cannot identify the model version, training basis, retrieval sources, data retention period, or decision owner, it cannot investigate an incident or explain an outcome.
Require model cards or equivalent documentation, change records, test evidence, and an explicit retirement trigger. Poor lifecycle management leaves abandoned prompts, stale embeddings, excessive permissions, and forgotten integrations active long after the original experiment ends.
How Do Shadow AI Risks Affect Compliance, Reputation, and Human Trust?
Compliance risk follows when an organization cannot prove where data went, how a decision was made, who approved the system, or whether affected people received meaningful review. Privacy obligations, sector rules, records requirements, intellectual property protections, and contractual commitments can all apply when an employee uses a consumer AI tool for convenience.
Maintain an AI register, document lawful purpose and data flows, and retain approval and testing evidence. Route high-impact uses through privacy, legal, security, and business review before deployment.
Customer trust can fail through misleading advice or marketing generated without adequate review. A chatbot that invents a product feature, a sales page that makes an unsupported claim, or a support assistant that gives unsafe instructions turns model uncertainty into a public promise.
Require approved source material, retrieval citations, restricted response boundaries, escalation to trained staff, and sampling of live outputs. Marketing and customer-facing teams remain accountable for every published claim, regardless of whether a model drafted it.
Employee trust also changes when AI is introduced without transparency. Workers who believe an undisclosed system ranks performance, screens applicants, monitors conversations, or replaces judgment will withhold information and work around controls. Explain what the system does, what it does not do, which data it uses, how humans review results, and how employees can challenge an outcome. Training should build practical judgment rather than assign blame for an unclear policy.
The strongest shadow AI risk assessment measures more than application count. It records data sensitivity, decision impact, model autonomy, vendor exposure, control maturity, user behavior, incident history, and residual risk.
Permit low-risk experimentation inside clear guardrails, apply deeper review to automation, and prohibit only uses that cannot meet the organization's privacy, safety, security, or accountability requirements. That approach keeps innovation visible while giving employees a trusted way to use AI without moving sensitive work into the shadows.
How Should Organizations Build a Shadow AI Inventory?
A shadow AI inventory starts with discovery rather than policy enforcement. Identify every AI tool, embedded feature, connection and use case before deciding whether it is approved. Combine technical telemetry with procurement, finance, vendor and employee input, then document each finding in a central register with a named owner and risk rating.
Treat the inventory as a living control within the shadow AI risk assessment. Update it when a model, provider, data flow, permission or business purpose changes, and connect it to a broader shadow AI management program.
1. Establish the Inventory Scope and Accountability
Define shadow AI broadly enough to capture more than standalone chatbots. Include generative AI websites, browser extensions, coding assistants, meeting transcription tools, automated decision features and AI embedded in approved applications such as Microsoft Teams, Salesforce and Tableau. An approved application is not automatically an approved AI use case because an embedded model can introduce new data flows, retention terms, permissions or fourth parties.
Assign one accountable team, such as security governance, privacy, enterprise architecture or an AI governance committee, to maintain the register. Give business units responsibility for validating their records, but do not make employees personally responsible for judging vendor risk. Employees provide valuable evidence about how work happens, while governance teams convert that evidence into controls.
Set a written scope that covers tools introduced by vendors, consultants, contractors, temporary workers and external partners. Include decentralized purchasing, personal devices and home networks when people perform company work. A corporate network-only review misses AI use through a contractor’s laptop, a personal browser profile or a home connection.
Use a consistent status model such as discovered, under review, approved, restricted, prohibited, retired or unknown. Record the discovery date, confidence level and evidence source for every entry. Ownership and risk-based review then become part of the operating model instead of an annual exercise.
2. Combine Technical and Human Discovery Sources
No single data source reveals the full estate. Start with network traffic, DNS and secure web gateway records to identify AI domains, API endpoints and unusual upload patterns. Add browser extension inventories, endpoint logs and EDR/XDR telemetry to find locally installed assistants, plug-ins, command-line clients and AI-enabled developer tools.
These sources show that a tool was accessed, but they often cannot explain the business purpose or information sensitivity. Use SaaS discovery and CASB data to identify unsanctioned applications, dormant accounts, shadow integrations and departmental usage.
Review OAuth connections in Microsoft 365, Google Workspace and identity platforms for AI applications with access to mailboxes, cloud storage, calendars, contacts or internal documents. DLP events can reveal sensitive data transfers, while SIEM correlation can connect a user, device, application, destination and time. Treat DLP, CASB, SIEM and EDR/XDR as complementary signals rather than interchangeable inventories.
Technical evidence still leaves important gaps. Reconcile it with procurement records, purchase orders, corporate-card and expense data, software renewals, legal reviews and vendor questionnaires. Ask accounts payable to flag unfamiliar AI-related merchants, and require procurement to collect AI-use disclosures during vendor onboarding.
Question vendors, consultants, contractors and partners about models used to process company information, including tools embedded in deliverables or service workflows. Create a safe employee reporting channel and explain that its purpose is visibility and risk reduction rather than punishment.
Use self-reporting forms, anonymous employee surveys and manager interviews to uncover occasional use, personal accounts and work performed outside managed devices. Ask which AI tools employees use, what tasks they perform, what information they enter, whether outputs enter company systems and whether another party introduced the tool. Accept “unknown” rather than encouraging guessed answers.
After collection, normalize names and domains, deduplicate aliases and separate a product from a use case. “Chat assistant” is not enough. Record whether the use involves drafting public copy, summarizing confidential contracts, writing code, analyzing customer data or making recommendations about people.
The same provider can represent low, moderate or high risk depending on the data and decision involved.
3. Capture the Fields Needed for a Defensible Record
Each inventory record should answer four questions: who is accountable, what happens to the data, how quickly can the risk change and what action follows. Use one record for each distinct tool-and-use-case combination when different departments, data types or permissions create different exposure.
Capture these fields:
- Ownership and use: Owner, accountable executive, department, users, user count, business purpose, use case, discovery source, date identified, last-verified date and inventory status.
- Model and provider: Provider, product, model and version, hosting arrangement, API or browser access, whether the model is embedded in an approved application, and any fourth parties or subprocessors.
- Data flow: Data inputs and outputs, data classification, examples of personal, financial, health, confidential or regulated information, prompt and output handling, retention period, deletion terms, residency and transfer locations.
- Access and connections: Permissions, roles, authentication method, OAuth scopes, integrations, connected repositories, downstream systems, export paths and whether use occurs on managed devices, personal devices or home networks.
- Provider assurance: Vendor certifications, independent assessment reports, security documentation, incident-notification terms, training use, opt-out controls, contract coverage, data-processing terms and restrictions on secondary use.
- Risk and accountability: Regulatory exposure, business impact, likelihood, risk velocity, control gaps, mitigation owner, approval authority, review frequency, exception expiry date, evidence links and responsibility for monitoring or retirement.
Define risk velocity as how quickly exposure can change. A tool that receives weekly model updates, gains new plugins, expands permissions or reaches hundreds of workers needs a shorter review cycle than a stable internal assistant handling public information.
Record current risk alongside change triggers, including a provider acquisition, new fourth party, altered training policy, new integration or sharp increase in user count. Do not mark an item “low risk” merely because the provider holds a certification.
A certification can support assurance, but it does not establish that the contract covers the use case, prompts are excluded from training or data remains in the required residency. Require evidence for every conclusion and record unresolved questions as open actions.
4. Handle Occasional and Indirect Use Without Losing Visibility
Occasional use belongs in the inventory when it touches company work, even when no employee has an account. A worker who pastes a customer email into a public chatbot creates a traceable use case. The same applies to a consultant who uses an AI transcription service during a project meeting or a contractor who generates code through a personal assistant.
Record the tool as intermittent, estimate frequency, identify the data involved and assign a business owner who can confirm whether the practice continues.
Indirect use requires a separate review path. A vendor can use AI to summarize support tickets, and a recruiting agency can screen applications with a third-party model. A software provider can also add an AI feature to an existing platform without a new procurement event.
Ask every critical vendor and partner to disclose AI systems, model providers, fourth parties, training use, retention, residency, deletion and incident procedures. Update the parent application record when an embedded feature appears, but create a linked use-case record when its data or permissions differ.
5. Turn the Register Into a Living Control
Review high-velocity or high-impact records monthly, medium-risk records quarterly and low-risk records at least annually. Trigger an out-of-cycle review after a DLP event, new OAuth connection, material permission change, provider policy update, new regulatory requirement, unusual usage spike or reported data exposure.
Reconcile the inventory against technical telemetry and purchasing data during every cycle so stale records do not create false assurance. Use the register to route action. Security can restrict risky permissions, privacy can assess data handling, and procurement can close contract gaps. Legal can evaluate regulatory exposure, and managers can replace unsafe workflows with approved options.
A focused shadow AI risk management program should connect these findings to human behavior because employees often discover useful tools before governance teams do. The objective is to make every AI use case visible, attributable and proportionate to the risk it creates, while allowing experimentation to continue inside guardrails.

How Should Organizations Score Shadow AI Risk With Probability, Impact, and Risk Velocity?
A shadow AI risk assessment becomes actionable when every tool receives the same repeatable score for probability and impact. Inventory the tool, rate its exposure, calculate inherent risk, overlay business context such as user count and risk velocity, and reassess after controls. Keep the arithmetic simple, but never let one number conceal sensitive data, excessive permissions, weak vendor assurance, or regulatory exposure.
1. Apply a Reproducible 1-to-5 Scoring Model
Start with the tool, its users, and the business process it supports. Record the application name, owner, purpose, account type, departments using it, data entered, integrations, permissions, vendor, contract status, and last review date.
Treat an employee using a public chatbot to draft harmless text differently from a finance team uploading customer records. An AI platform that can retain prompts for model improvement changes the exposure entirely.
Score impact from 1 to 5 based on the most serious credible consequence if the tool mishandles data or produces harmful output.
| Impact | Definition | Typical Consequence |
|---|---|---|
| 1 | Negligible | No sensitive data, no meaningful operational effect, and no difficult replacement |
| 2 | Limited | Internal information exposure or minor workflow disruption |
| 3 | Material | Confidential business information, repeated process disruption, or a contained compliance issue |
| 4 | Major | Personal, regulated, customer, financial, or strategic data exposure with significant legal or operational consequences |
| 5 | Critical | Large-scale regulated-data exposure, material fraud risk, safety impact, major customer harm, or existential business consequences |
Use the highest applicable impact rather than averaging categories downward. A productivity tool handling one sensitive payroll export deserves an impact rating based on that export rather than on the majority of harmless prompts.
Score probability from 1 to 5 using observable evidence rather than instinct.
| Probability | Definition | Evidence Pattern |
|---|---|---|
| 1 | Rare | No active users or access path, strong controls, and no credible misuse scenario |
| 2 | Unlikely | Limited use, low-value data, restricted permissions, and reliable oversight |
| 3 | Possible | Regular use, incomplete inventory, moderate data access, or inconsistent employee practice |
| 4 | Likely | Broad adoption, sensitive prompts, weak vendor terms, excessive permissions, or known policy violations |
| 5 | Almost certain | Active data leakage, public accounts, unmanaged integrations, repeated risky behavior, or an imminent high-impact use case |
Calculate the base score with one formula:
Inherent risk = impact × probability
A marketing employee using an unapproved AI writing tool for public copy might receive an impact score of 2 and probability score of 3. That combination produces an inherent risk score of 6. A recruiting platform receiving applicant resumes could receive an impact score of 4 and probability score of 4, producing 16. The second tool receives priority because both the consequence and the exposure are substantial.
Treat the score as a starting point for review rather than a final verdict. Record the factors that drove each rating so another reviewer can reproduce the decision. A security leader should be able to explain why a tool scored 12 in March and 6 in June without relying on undocumented judgment.
Map assessment fields to the NIST AI Risk Management Framework and its Generative AI Profile. The NIST AI RMF Generative AI Profile, published in 2024, organizes AI risk work around governing, mapping, measuring, and managing risk. That structure supports reassessment when a tool, use case, data class, or permission set changes.
2. Add Context Without Distorting the Core Score
A simple 1-to-25 score becomes misleading when two tools receive the same result for entirely different reasons. Keep impact and probability as the heat-map axes, then display the following context as separate fields, badges, or overlays:
- Data sensitivity: Classify inputs as public, internal, confidential, personal, regulated, or restricted. A tool handling health information, payment data, credentials, source code, or unreleased strategy should carry a visible sensitivity tag.
- Number of users: Record active users and affected departments. A score of 9 for 10 users demands a different response from a score of 9 for 10,000 users.
- Regulatory exposure: Identify relevant obligations, including GDPR, HIPAA, PCI DSS, contractual privacy terms, records-retention rules, and sector requirements. Show which obligation creates the exposure rather than marking the tool as merely “regulated.”
- Vendor assurance: Review data-use terms, retention, deletion, subprocessors, breach notification, model-training commitments, independent assessments, and contract ownership. An approved vendor is not evidence that every use case is approved.
- Autonomy: Rate whether the system drafts content, recommends actions, executes tasks with approval, or acts independently. Autonomous behavior increases the consequence of inaccurate output and unauthorized action.
- Permissions: Record access to email, files, source repositories, customer systems, financial workflows, APIs, browser sessions, and identity providers. A low-data tool with write access to a production system remains high risk.
- Customer exposure: Mark whether outputs reach customers, regulators, patients, applicants, investors, or the public. External delivery raises trust and reputational consequences even when the input data is not regulated.
- Risk velocity: Measure how quickly exposure is increasing. Track new users, new data classes, new integrations, policy violations, and changes in model capability over a defined period.
- Control maturity: Record whether controls are absent, documented, enforced, monitored, tested, or independently reviewed.
Do not add these factors mechanically to the impact-probability score unless the organization has validated a weighted model. An additive formula can hide a critical permission behind several low-risk attributes. Display a 12 with a red “restricted data” badge and a “high velocity” marker rather than converting every nuance into false precision.
Risk velocity deserves its own signal because a moderate tool can become unacceptable before the next quarterly review. Define velocity as the change in exposure over time, such as active users added per month, sensitive-data events per week, or the percentage increase in connected permissions. A tool moving from 40 to 400 users in 30 days should trigger review even if its current probability and impact ratings remain unchanged.
3. Prioritize With a Heat Map and Exposure Overlays
Build the heat map with probability on the horizontal axis and impact on the vertical axis. Each cell represents an inherent risk score from 1 through 25.
Place one circle per tool, size the circle by user count, and color its border by risk velocity. This design shows a committee three priorities at once: how serious the outcome could be, how likely it is, and how widely or quickly the exposure is spreading.
Use example categories that fit the organization’s risk appetite:
| Category | Example Score | Default Decision |
|---|---|---|
| Unacceptable | 20 to 25 | Ban or suspend use immediately unless the committee approves an exceptional, time-limited remediation plan |
| High | 12 to 19 | Sandbox, restrict, or formally adopt only after mandatory controls and accountable ownership |
| Moderate | 6 to 11 | Monitor, document, and remediate defined gaps within a set deadline |
| Low | 1 to 5 | Permit within policy, retain an owner, and reassess on a scheduled cycle |
These thresholds serve as governance examples rather than universal regulatory limits. A hospital, bank, defense contractor, or organization handling children's data should lower its unacceptable threshold for tools that process regulated information. An internal brainstorming tool with no sensitive data or integrations can remain low risk under a less restrictive appetite.
Distinguish inherent risk from residual risk on every heat map. Inherent risk is the exposure before controls. Residual risk is the exposure that remains after controls operate and have been tested.
If a customer-support AI tool starts at impact 4 and probability 4, its inherent score is 16. After restricting it to de-identified data, removing write permissions, enforcing single sign-on, blocking model training on prompts, and reviewing outputs, probability might fall to 2. Its residual score becomes 8, while the impact remains 4 because a future control failure would still affect customers.
Show both points or connect them with an arrow. A tool moving from the upper-right quadrant to the center demonstrates control effectiveness. A tool whose score remains unchanged after a policy announcement shows that the organization changed paperwork without changing exposure.
4. Treat Residual Risk and Decide the Tool’s Fate
Risk treatment begins with the residual score rather than the original approval request. The AI steering committee should review the score, context overlays, control evidence, business value, and accountable owner before selecting one of four dispositions.
Ban the tool when it creates unacceptable exposure, lacks a viable control path, processes prohibited data, or has permissions that the business cannot constrain. Blocking access is complete only when the organization also removes accounts, revokes tokens, identifies exported data, and offers an approved alternative for the underlying work.
Sandbox the tool when the use case is valuable but the evidence is incomplete. Use test accounts, synthetic data, isolated workspaces, restricted browser access, read-only permissions, and a short expiration date. Require the owner to demonstrate safe prompts, output review, retention behavior, and incident reporting before granting broader access.
Monitor the tool when residual risk is moderate and controls are adequate for the current use case. Set thresholds for new users, sensitive-data events, permission changes, customer-facing output, and risk velocity. A monitoring decision must include an owner, review date, alert route, and escalation trigger.
Formally adopt the tool when the business need is documented, the vendor passes assurance review, and data handling is contractually clear. Permissions should follow least privilege, training should be complete, and residual risk should fit the organization's appetite. Adoption does not mean permanent approval. Reassess after a major model update, new integration, material user growth, policy violation, security incident, or change in data classification.
Controls reduce risk only when they change a scored input. Data-loss prevention that blocks restricted prompts can lower probability. Removing API write access can lower probability or impact, depending on the scenario. Human approval before customer-facing output can lower probability.
Targeted training can reduce risky behavior when completion and observed use are measured rather than merely recorded. Human risk management and risk scoring can connect employee behavior, policy violations, and remediation to the same review process.
Set a reassessment date for every treatment action. At review, preserve the original inherent score, document the control implemented, test whether it operates, assign the residual score, and record the next trigger.
If a tool remains high after remediation, escalate rather than repeatedly extending the deadline. The heat map should keep unresolved exposure visible until the committee bans the tool, contains it, or accepts the residual risk with a named executive accountable for that decision.

How Should Organizations Assess AI Data Flows, Vendors, and Embedded Features in a Shadow AI Risk Assessment?
A shadow AI risk assessment compares approved SaaS AI features with separately purchased tools across data exposure, accountability, and contractual control. An embedded feature sits inside an existing vendor relationship, while a separately purchased tool introduces a new provider, model, API, contract, and fourth-party chain. Both categories require evidence that the organization understands the model, controls the data, and can shut the use case down safely.
Embedded functionality often appears lower risk because procurement approved the parent application. A newly added AI feature can change where data travels and turn that application into shadow AI. Separately purchased tools require explicit ownership and due diligence before employees submit company information or rely on generated outputs.
How Should Teams Review the AI Data Flow?
Data-flow review must establish what enters the system, where it goes, what the model produces, and which parties retain access. Document the provider, model version, application, API endpoint, hosting region, input fields, output destinations, user roles, integrations, logs, backups, and deletion path. Include consultants, resellers, plug-ins, embedded models, and subprocessors rather than stopping at the visible application.
The assessment should determine whether prompts and uploaded files train a shared model and how long inputs and outputs remain available. It should also establish whether administrators can delete them and whether deletion reaches backups and derived data. Require encryption in transit and at rest, logical tenant isolation, physical security controls, least-privilege access, privileged-session monitoring, and a defined retention schedule.
For sensitive use cases, prohibit production secrets, regulated records, source code, credentials, and personal data until the provider documents each control.
Third-party technology, data governance, testing, and lifecycle management operate as connected activities across the full workflow. Connect the use case to the organization's human risk management program so risky AI behavior becomes an actionable signal rather than an invisible policy violation.
What Vendor and Fourth-Party Checks Belong in the Assessment?
Vendor due diligence must test the provider's operating model rather than its marketing description. Request model documentation, intended and prohibited uses, training-data provenance, validation methods, accuracy testing, bias testing, adversarial testing, change-management procedures, and the process for retiring or replacing a model. Confirm how the provider handles hallucinated outputs, unsafe recommendations, prompt injection, model abuse, and material changes to model behavior.
Trace the fourth-party chain behind the use case. Identify the foundation-model provider, cloud host, API gateway, data processor, support contractor, analytics service, plug-in, and consultant that can view prompts or outputs. Review subprocessor locations, notification periods, objection rights, access controls, incident history, business continuity, disaster recovery, and exit assistance.
A vendor that cannot identify who receives customer data cannot support a defensible approval decision. Require the provider to map every data recipient to a specific control, contract, and accountable owner before the use case moves into production.
What Contract and Audit Evidence Should Procurement Require?
Contract review should establish whether the AI feature is covered by the existing SaaS agreement. Examine the order form, data-processing addendum, acceptable-use policy, privacy notice, service description, security exhibit, subprocessor terms, and product-specific AI terms. A general SaaS agreement does not automatically govern a later AI assistant, new API, or beta feature. Those terms often reserve broad rights to retain prompts or use them for model improvement.
Require written commitments covering data ownership, no-training restrictions where needed, retention and deletion, data residency, confidentiality, and encryption. Those commitments should also address logical isolation, breach notification, vulnerability management, insurance, service availability, model-change notice, and termination support. Contract terms should define incident cooperation, regulator support, forensic access, audit rights, independent assessment reports, penetration-test summaries, and remediation deadlines.
Direct audits are often impractical, so procurement can accept current vendor certifications and independent assurance reports when their scope includes the AI service, its hosting environment, and relevant subprocessors. Documents that exclude the AI feature do not provide evidence for approving that feature.
What Approval Evidence Should Authorize an AI Tool?
Approval must create an accountability record that survives employee turnover and vendor changes. Procurement should not authorize a tool based on a completed questionnaire alone. Store the model and version, intended function, business owner, technical owner, approved users, input categories, and expected outputs. Record the human review requirement, prohibited data, control mapping, vendor assessment, contract review, renewal date, monitoring plan, and decommissioning trigger.
Use this practical questionnaire before approval:
- Use case: What business task does the AI perform, and what decision remains with a qualified employee?
- Data: What inputs and outputs exist, and do they contain confidential, regulated, personal, or proprietary information?
- Model: Which model, API, application, and version process the data, and how are outputs validated?
- Chain: Which provider, consultant, subprocessor, cloud host, plug-in, or fourth party can access the workflow?
- Controls: Are retention, training use, deletion, residency, encryption, isolation, access, logging, and human review documented?
- Accountability: Who owns the use case, approves exceptions, monitors changes, reviews incidents, and retires the tool?
- Evidence: Are the security documents, contract terms, subprocessor list, audit report, model documentation, and approval record stored centrally?
Reassess the tool when its model, data use, permissions, subprocessors, or embedded AI features change. That review prevents an approved application from quietly becoming shadow AI through an update employees adopt before security evaluates its consequences. It also gives teams a clear record of what responsible use requires.
How Can Organizations Detect and Monitor Shadow AI Usage?
A shadow AI risk assessment compares monitoring layers to show where AI use becomes visible and where it remains hidden. Network and API monitoring identify centralized traffic and service calls, while endpoint, browser, SaaS, OAuth and identity telemetry connect activity to a person, device and permission context.
No single control sees the complete picture, so organizations should combine signals according to data sensitivity, workforce model and response capacity. Reliable shadow AI detection depends on that combination.
Which Telemetry Sources Reveal Shadow AI Activity?
Network monitoring identifies connections to known AI domains, unusual upload volumes, new destinations and traffic that bypasses approved gateways. API monitoring adds provider, token, model and application details when teams use managed AI services, but it cannot see employees working through consumer websites or personal API keys. Both methods create useful perimeter signals without necessarily identifying the prompt author or the sensitive information submitted.
Browser extensions and endpoint telemetry close that identity gap by recording the active account, browser session, device posture, domain, upload event and policy result. They can identify an employee pasting source code into a public chatbot, downloading generated files or switching to an unsanctioned model version.
EDR and XDR contribute process, identity and device context, but they are not designed to interpret prompt semantics. CASB and SaaS discovery map cloud applications and usage patterns, while OAuth review exposes risky third-party grants and unusual permissions.
DLP detects sensitive data movement, and SIEM correlates these events with authentication, ticketing and incident records.
The coverage gap is largest outside managed infrastructure. Personal devices, unmanaged browsers and home networks can hide AI use from corporate DNS, proxy, API and EDR controls. Close that gap with conditional access, identity-based SaaS policies, managed browser sessions for sensitive roles, approved remote-work devices and clear rules for personal accounts. Unobserved activity still carries risk, so record missing telemetry and route high-risk access for review.
How Can Organizations Detect Sensitive Data Without Retaining Prompts?
Privacy-aware content detection should classify prompt data without creating a searchable archive of employee conversations. Inspect content transiently at the browser or policy-enforcement point, then match it against detectors for credentials, payment data, health information, source code, confidential project terms and regulated identifiers. Retain only the category, confidence, destination, user role, timestamp, policy action and cryptographic event identifier.
Hashing known secrets, tokenizing identifiers and redacting content before inspection limits exposure during investigation. Access to any quarantined sample should require documented approval, narrow role-based permissions and a short retention period.
Security teams should publish what they inspect, why they inspect it and how long metadata remains available. The 2025 NIST Cyber AI Profile recommends minimizing sensitive data in AI prompts and using runtime redaction and guardrail.
Permissions require equal attention. Review OAuth grants for excessive scopes, service accounts with broad model access, newly added plugins and vendors that can read corporate files. Alert when a user grants a new application access to mail, drives or repositories. Alert again when a model or vendor version changes, or when a low-privilege role attempts a high-impact action. These signals indicate unauthorized access even when prompt content remains unavailable.
How Should Organizations Validate Shadow AI Controls?
Control validation must use known activity rather than dashboard confidence. Create a controlled test matrix covering an unsanctioned AI domain, an approved API, a personal account, a managed browser, an unmanaged device and a home network.
Extend the matrix to a sensitive-data prompt, a large file upload, an OAuth grant, a prompt injection string and an attempt to manipulate model instructions. Use synthetic secrets and fictional records so testing proves detection without exposing real information.
Record whether each control observed the event, identified the user and device, classified the data category, blocked or warned, generated a ticket and preserved only the approved metadata. Test correlation across browser or endpoint telemetry, CASB, DLP, SIEM and identity systems.
A control that blocks a domain but cannot link the event to a user creates an investigation gap. A control that detects a prompt but retains its full contents creates a privacy risk.
Repeat the matrix after browser, model, vendor, policy and agent updates. Monitor for prompt injection indicators such as instructions to ignore system rules, requests to reveal hidden context, encoded commands, tool-call anomalies and sudden changes in output behavior. Assign each signal an owner, severity and escalation path, and set a defined monitoring cadence so control drift does not create silent exposure.
High-confidence events should trigger access restriction, token revocation or data-loss containment. Ambiguous events should create a review queue rather than punish the employee. Link approved monitoring outcomes to role-based training and human risk reporting so employees build safer AI habits while security leaders retain an auditable view of exposure. A repeatable shadow AI risk assessment turns hidden activity into accountable decisions about access, data and human behavior.

How Should Organizations Treat Shadow AI Risk by Tier?
A shadow AI risk assessment should convert discovery findings into proportionate treatment plans rather than a blanket ban on every unsanctioned tool. Classify each tool, match controls to its exposure, assign an accountable owner, and record the decision. Reassess whenever the tool adds features, changes its terms, or begins handling more sensitive information.
1. Apply Controls by Risk Tier
Treatment depends on the consequences of misuse, the data involved, the provider's controls, and the business need. A tool that processes public marketing copy does not require the same response as one receiving patient records, source code, or payment data.
| Risk Tier | Required Treatment | Minimum Controls |
|---|---|---|
| Unacceptable | Ban or replace | Remove access, preserve evidence, investigate data exposure, assess notification duties, and provide an approved alternative |
| High | Sandbox or tightly restrict | Limit data, require human approval, enforce stronger authentication, retain logs, review contracts, and assign executive ownership |
| Moderate | Monitor and govern | Define approved uses, restrict permissions, review activity periodically, and train users on handling rules |
| Low | Formally adopt with light controls | Register the tool, accept standard terms, document its owner, and continue monitoring for changed risk |
Unacceptable tools create exposure the organization cannot accept. Examples include providers that retain prompts for model training without approval, tools lacking contractual privacy commitments, and services that receive regulated data without authorization. Containment comes first. Disable accounts, revoke tokens, block access where appropriate, and preserve browser, identity, proxy, and application logs before remediation changes overwrite evidence.
Investigate what was submitted, when, by whom, under which account, and whether the provider stored, shared, or used the material. Replacement prevents employees from returning to the same unsafe workflow. Provide an approved tool that meets the business need, and explain the restriction without blaming the employee who surfaced the gap.
High-risk tools remain useful only inside a controlled boundary. Place them in a sandbox, prohibit regulated or confidential data, require human approval for generated outputs that influence customers or transactions, and enforce phishing-resistant MFA where available. Log prompts, users, outputs, approvals, and administrative changes.
Procurement and legal teams should review data-processing terms, retention, subprocessors, intellectual-property rights, model-training practices, and breach obligations. An executive owner must accept residual risk and fund the controls required to keep the use case active.
Moderate-risk tools need enforceable operating rules rather than silent tolerance. Permit defined use cases, limit permissions to the smallest workable scope, and review activity and provider changes periodically. Train employees to remove personal, financial, health, customer, and confidential business information before submission. Connect those rules to security awareness training for AI-era behavior so employees practice the decision instead of memorizing a policy.
Low-risk tools still require lightweight registration. Record the tool name, owner, purpose, data types, users, provider terms, and review date. Low risk reflects a current judgment rather than a permanent label. A summarization tool can become high risk after adding file uploads, plugins, agentic actions, or a new retention policy.
2. Govern the Decision and Assign Accountability
A documented decision process prevents risk ratings from becoming isolated security opinions. Security should assess technical exposure, privacy should evaluate personal-data implications, and legal should review contractual and regulatory duties. Procurement should validate the provider, and the business owner should justify continued use. High-risk approvals should expire unless the owner renews them with current evidence.
Use four decisions consistently:
- Ban when exposure is unacceptable or no defensible control exists.
- Sandbox when business value justifies restricted experimentation.
- Monitor when controlled use is acceptable but evidence remains incomplete.
- Formally adopt when the organization has approved the tool, owner, contract, data boundaries, and review cadence.
The 2024 NIST Generative AI Profile directs organizations to align incident response with breach-reporting, data-protection, and privacy requirements. That guidance makes AI governance part of operational risk management rather than a one-time procurement exercise.
Record the rationale, rejected alternatives, compensating controls, approval date, and trigger for reassessment. Report exceptions to an executive risk committee, especially when a department continues using a banned tool or business pressure encourages teams to bypass review. Clear ownership turns a risk register into an operating control.
3. Respond to a Confirmed Regulated-Data Submission
A confirmed submission of regulated data to an unauthorized AI service counts as a potential data incident rather than a simple policy violation. Immediately stop further access, preserve evidence, and identify the affected data subjects and fields.
Determine whether the provider retained or disclosed the content, then rotate exposed credentials or secrets. Do not ask the employee to delete evidence or contact the provider independently before legal and incident-response teams establish a record.
The incident lead should involve privacy, legal, security, compliance, the affected business owner, and communications as appropriate. Assess whether contractual notice, regulator notification, customer communication, law-enforcement reporting, or insurer notification applies in each relevant jurisdiction.
For organizations subject to UK GDPR, the Information Commissioner's Office 2025 personal-data breach guidance states that a likely risk to people's rights requires notification as soon as possible. Where feasible, that notification should occur within 72 hours. The incident record should document the decision, the evidence reviewed, and the reason for any notification or non-notification.
Close the response by documenting root cause and changing the control that failed. Update the approved-tool register, block the unauthorized service or data path, revise training, and test whether the new control stops recurrence. Treat the employee’s report as a valuable signal. Transparent reporting shortens containment time and reveals where safer workflows must become easier to use.
Which Teams Should Participate in Shadow AI Governance?
Security should own the control architecture, cyberthreat scenarios, monitoring signals and incident-response connection. IT should maintain identity, device, browser, network and SaaS visibility, then enforce access, isolation and approved integration patterns. Privacy determines whether prompts, outputs or telemetry contain personal data and whether a data protection impact assessment is required.
Legal interprets contractual, intellectual-property, employment and regulatory exposure. Compliance maps controls and evidence to the organization’s obligations. Procurement conducts vendor reviews with security, privacy and legal input, including data-use terms, retention, subprocessors, breach notification and model-training provisions. Data governance classifies information, defines permitted data uses and sets retention rules.
HR handles workforce communication, acceptable-use expectations, employee consultation and role-based AI literacy. Finance approves business cases and controls spending on unapproved tools. Internal audit independently tests whether the operating model works as documented. Business owners remain accountable for each AI use case because they understand its process, users and consequences better than a central committee.
AI developers and application engineers own technical implementation. They should work with governance teams to isolate development environments, encrypt data in transit and at rest, and enforce least-privilege access. They should also validate models and prompts, monitor drift and abuse, and retire systems without a defensible purpose. Employees who discover useful AI tools need a safe route to disclose them, so visibility replaces concealment.
How Should Ownership and Risk Appetite Be Assigned?
Create a central AI register, but assign an accountable owner to every record. The register owner maintains the inventory, confirms the system's purpose, records the provider and model, identifies data flows, documents users and schedules reassessment.
The business owner accepts operational accountability for the use case. Security owns technical risk treatment, privacy owns personal-data analysis, procurement owns vendor due diligence, and legal or compliance owns regulatory interpretation.
Policy exceptions require a named approver, a business justification, compensating controls, an expiration date and a review trigger. Incident response should have one coordinator who can convene security, IT, privacy, legal, communications, HR and the affected business owner. Residual-risk acceptance belongs to an executive with authority matching the potential impact rather than to the person requesting the exception.
An AI steering committee should meet on a defined cadence under a written charter. Its mandate should cover inventory completeness, risk-tier decisions, approved-use standards, exception review, vendor escalation, incident lessons and investment priorities. Set risk appetite in business terms. An organization can prohibit unapproved disclosure of regulated data, require human review for consequential decisions and allow low-impact experimentation only inside an isolated environment.
Escalation should be predictable. A policy breach goes to the system owner and security lead. Personal-data exposure adds privacy and legal review. Material financial, regulatory, safety or reputational risk escalates to executive risk governance.
Preserve assessment, approval, exception, testing, monitoring, incident and retirement records in a controlled repository. Immutable timestamps, version history and access logs turn governance decisions into evidence. A human risk management program can add behavioral signals when employees use unapproved AI tools or transfer sensitive information through them.
How Should the Assessment Align With Major Frameworks?
Map the lifecycle rather than claiming that one framework creates automatic compliance. Use the NIST AI Risk Management Framework's Govern, Map, Measure and Manage functions to structure accountability, context analysis, testing and remediation. ISO/IEC 42001 provides management-system discipline for documented policies, leadership oversight, continual improvement and evidence.
Enterprise risk management should connect each AI risk to business objectives, risk owners, treatment plans, controls and board reporting. A documented AI governance framework keeps those connections consistent across departments.
The 2024 EU AI Act adds binding obligations that depend on the system's role, purpose and risk classification. Those obligations involve AI literacy, risk management, data governance, technical documentation, logging, human oversight, monitoring and incident reporting for applicable systems.
Treat these frameworks as alignment points rather than a compliance shortcut. For each use case, retain the classification rationale, data assessment, validation results, access design, monitoring plan, human-oversight procedure, vendor evidence, exception decision and retirement record.
What Belongs in a Shadow AI Risk Assessment Checklist?
1. Build the Shadow AI Risk Assessment Checklist
Define the assessment scope before discovery begins. Include public generative AI tools, employee-created accounts, browser extensions, AI features embedded in approved SaaS applications, APIs, internally developed models, AI agents, and vendor services that process organizational data. Set the organization’s risk appetite at the same stage, including prohibited data types, acceptable business uses, approval thresholds, retention requirements, and ownership responsibilities.
Run discovery across identity logs, browser telemetry, expense records, procurement systems, SaaS catalogs, API repositories, help desk tickets, and employee declarations. For every tool or feature, record the business owner, users, purpose, model or version, provider, deployment location, data processed, connected systems, and access permissions. Note whether outputs influence customers, employees, financial decisions, or regulated processes.
Use one shadow AI risk assessment checklist to drive consistent decisions:
- Scope and appetite: Define systems, business units, jurisdictions, prohibited uses, approval levels, and escalation owners.
- Use-case classification: Rate each use as administrative, internal decision support, customer-facing, high-impact, or prohibited. Record whether the model generates, summarizes, recommends, executes, or makes decisions.
- Data-flow mapping: Trace inputs from an employee or system to the model provider, subprocessors, storage locations, logs, training use, outputs, and downstream applications. Attach classifications for personal, confidential, regulated, proprietary, and public information.
- Vendor and contract review: Retain security questionnaires, privacy terms, data-processing agreements, subprocessors, breach-notification terms, retention settings, model-training provisions, regional processing commitments, audit rights, and deletion procedures.
- Scoring and treatment: Score inherent risk using data sensitivity, exposure, privilege, business impact, model behavior, vendor dependency, and likelihood of misuse. Document controls such as access restrictions, masking, human review, logging, output validation, approved prompts, and data-loss controls.
- Approval and exception: Capture the accountable owner, security and privacy approvals, legal review where required, exception rationale, expiration date, compensating controls, and risk-acceptance statement.
- Monitoring and response: Define telemetry, control tests, incident ownership, reporting channels, containment steps, user retraining, vendor notification, and reassessment dates.
Small businesses with limited resources should begin with a minimum viable assessment. Create one inventory spreadsheet, require owner and purpose fields, prohibit sensitive data in unapproved tools, review the 10 highest-use applications, and document approval or rejection. That baseline creates visibility without waiting for a full governance platform.
2. Assemble the Evidence Package
Evidence must show what the organization discovered and how it acted. Retain dated inventory records, screenshots or telemetry summaries, data classifications, completed vendor questionnaires, contracts, and data-processing agreements. Keep architecture or data-flow diagrams, access reviews, control-test results, training records, approval tickets, exception decisions, risk-acceptance statements, incident tickets, remediation records, and the next review date.
Store evidence with a unique assessment ID and link every record to the specific tool, use case, owner, risk score, and decision. A control that exists only in policy is not an operating control.
Test whether access restrictions, logging, deletion, human review, and approved-use instructions work, then record the test method, result, tester, and corrective action. A documented human risk management workflow can connect employee behavior signals to assigned training, control ownership, and review decisions.
Run a tabletop exercise for a shadow AI breach or compliance event. Participants should determine who receives the alert and which logs establish what data was submitted. They should also establish whether the provider retained or trained on it, which regulators or customers require notification, how access is contained, and how affected employees receive non-punitive coaching. Record decisions, delays, missing evidence, and assigned corrective actions.
3. Set the Review Cadence and Reassessment Triggers
Set a routine review at least monthly for high-risk tools and annually for low-risk tools, with ownership confirmed at every review. Recalculate residual risk after controls operate, compare it with the approved appetite, and either continue, remediate, restrict, retire, or formally accept the remaining exposure. A closed assessment should include the reviewer, decision date, unresolved actions, and next checkpoint.
Trigger an out-of-cycle review after a new vendor, model version, ownership change, data-processing location change, or embedded AI feature. Rapid user growth, a security or privacy incident, a material policy change, or increased risk velocity should trigger the same response. Increased risk velocity includes a sharp rise in usage, new data types entering prompts, faster model releases, expanded permissions, or repeated policy violations.
Reassessment is complete only when the organization can answer three questions: What changed? What evidence proves the current controls work? Who accepted any residual risk? That discipline turns shadow AI from invisible employee behavior into a governed, reviewable business process, while unresolved ownership and data-flow gaps remain signals for immediate action.
How Should a Shadow AI Risk Assessment Handle AI Agents, Code, Open-Source Models, and High-Impact Uses?
A shadow AI risk assessment must distinguish ordinary chatbot experimentation from AI systems that can act inside business processes. Agency creates the central difference. A chatbot produces content for a person, while an AI agent can use tools to change systems, communicate externally, spend money or make decisions.
AI-generated code creates development and intellectual-property risk, while employee experimentation creates data-handling and confidentiality risk. Internal automation adds operational and access risk, and customer-facing systems add accuracy, disclosure, consumer-trust and regulatory exposure.
These uses share a need for inventory and ownership, but controls must scale with autonomy, permissions, reversibility and the harm an inaccurate output can cause.
How Should an Assessment Measure Agent Autonomy and Permissions?
Agent autonomy is the primary escalation point because an incorrect answer becomes an incident when the system can act without review. Record every connected tool, API, data source, identity, permission scope, action limit, environment and business owner.
An agent that drafts a purchase order is lower risk than one that submits it. An agent that submits it is lower risk than one that can approve payment. Broader AI access governance practices apply the same logic to non-human identities.
A practical assessment should require:
- Permission boundaries: Use least-privilege service accounts, separate read and write access, restrict production systems and block sensitive actions by default.
- Action limits: Set transaction ceilings, recipient allowlists, rate limits, time windows and prohibitions on bulk deletion, external posting or unsupervised purchasing.
- Approval gates: Require a named employee to approve payments, customer commitments, code merges, destructive changes and messages sent in an executive’s name.
- Rollback: Test cancellation, version restoration, credential revocation and quarantine procedures before deployment.
- Monitoring: Log prompts, retrieved data, tool calls, outputs, approvals, overrides and final actions, then alert on unusual volume, destinations or privilege use.
- Accountability: Assign one accountable owner who can suspend the agent, explain its decisions and coordinate incident response.
Treat agent risk as a lifecycle process. Reassess agents after model, tool, prompt, data or permission changes instead of approving them once.
How Do AI-Generated Code and Open-Source Models Compare?
AI-generated code carries a different risk profile from employee experimentation because it can enter production, inherit hidden vulnerabilities or expose proprietary logic. Assess whether generated code is a learning aid, committed to an internal repository, deployed into production or embedded in customer-facing software. Require human code review, automated testing, dependency scanning, secret detection, provenance records and a documented decision on who accepts the residual risk.
Open-source and locally hosted models remove some external data-transfer concerns but do not remove governance obligations. Record the model name, version, checksum, source repository, maintainer, license, fine-tuning data, known limitations and security history. Patch the runtime, libraries, inference server and host infrastructure, then isolate model-serving environments and protect model weights. Restrict administrative access and test for prompt injection, data poisoning, model manipulation and unauthorized extraction.
Training-data provenance requires equal attention. Identify whether data was licensed, public, purchased, synthetic or supplied by employees, and whether it contains personal information, trade secrets, copyrighted works or regulated records.
A license that permits code use does not automatically permit model training or commercial redistribution. The European Union Artificial Intelligence Act, adopted in 2024, addresses model documentation, training-data summaries and copyright policies, making provenance and licensing records essential for downstream review.
How Should High-Impact or Customer-Facing Decisions Be Controlled?
High-impact use cases require stronger controls because inaccurate outputs can deny access, misprice services, discriminate, mislead customers or create legal exposure. Classify systems by affected population, decision authority, reversibility, data sensitivity, scale and dependence on the output. A model that summarizes an internal meeting is materially different from one that ranks job candidates, evaluates creditworthiness, triages healthcare cases or recommends account restrictions.
For customer-facing systems, disclose when a person is interacting with AI, label synthetic content and route consequential disputes to trained employees. Test factual accuracy, bias, refusal behavior, prompt injection and misleading content using representative scenarios before launch and after material changes. Keep an evidence trail showing the input, model version, output, reviewer, decision and correction path.
Human review must be substantive rather than a formality. Give reviewers enough context, authority and time to reject outputs, override recommendations and stop the system safely.
Human oversight, accuracy, robustness, cybersecurity, logging, and risk management form the regulatory baseline the EU AI Act sets for high risk AI systems, alongside transparency requirements for certain AI interactions and generated content. Treat those controls as a practical standard even when a use case falls outside a formal regulatory scope.
Which AI Uses Require the Strictest Assessment?
The strictest tier combines autonomous action, sensitive data, external impact and limited reversibility. Rank an AI agent that can move funds or alter production systems above an employee testing a public chatbot. Rank a customer-facing decision engine above an internal drafting assistant. Reassess the ranking whenever the system gains a new tool, broader data access, a larger user population or authority to act without approval.
This comparison keeps shadow AI risk assessment focused on consequences rather than product labels. The decisive question concerns what the system can access, change, publish or decide, and whether a responsible person can detect, reverse and explain the outcome.
Open-source status, local hosting and assistant branding matter far less than those boundaries. Examine them before deployment and monitor them continuously as the system and its surrounding workflow evolve.
How Can Policies, Training, and Approved Alternatives Reduce Shadow AI Risk During a Shadow AI Risk Assessment?
A shadow AI risk assessment should produce controlled adoption rather than a blanket ban that pushes legitimate work into unmonitored accounts. Build the program around clear data rules, approved tools, isolated testing environments, role-based permissions, and fast reporting. Review usage signals and employee feedback continuously, because a policy that blocks useful work without a practical alternative moves risk out of sight instead of removing it.
1. Set Policy Guardrails Employees Can Apply
A workable AI acceptable-use policy must define permitted activity, prohibited data, and situations requiring human review. A shadow AI policy template can accelerate that drafting work. Allow low-risk tasks such as drafting generic text, summarizing public information, brainstorming campaign ideas, or transforming nonconfidential content.
Prohibit employees from entering credentials, authentication codes, trade secrets, source code, customer records, protected health information, payment data, unpublished financial information, or regulated case details into public AI tools.
Require human review before AI-generated content reaches a customer, regulator, employee, investor, production environment, or legal workflow. Finance teams should verify payment instructions and generated analyses against trusted records. HR should exclude identifiable employee information from prompts.
Developers should review generated code for vulnerabilities, licensing conflicts, exposed secrets, and unsafe dependencies. The 2025 San Francisco Generative AI Guidelines also emphasize approved tools, privacy safeguards, human oversight, and accountability for public-sector use.
The policy must cover personal accounts, browser extensions, mobile apps, and AI features embedded in ordinary software. Require employees to report accidental data disclosure, suspicious output, unexpected tool behavior, or unapproved service use without fear of punishment for an honest mistake. Reserve disciplinary consequences for deliberate concealment, repeated violations after training, or intentional misuse, and apply them consistently across roles and seniority.
2. Provide Approved Alternatives and Controlled Sandboxes
A ban fails when employees still need faster research, writing, coding, analysis, or customer support. Give each high-use function an approved AI workspace with enterprise data controls, retention settings, access logs, contractual privacy protections, and administrator ownership. Use role-based permissions so marketing employees can access public campaign templates while finance analysts work in restricted environments designed for approved financial workflows.
Create isolated AI sandboxes for experimentation. Developers can test prompts and models with synthetic data, masked records, and nonproduction code before requesting access to a controlled API.
Security teams should require rate limits, logging, model allowlists, and review gates for applications that connect AI to internal systems. Approved APIs should use service identities instead of personal accounts, while templates should show employees how to remove sensitive fields, specify the desired output, and verify results.
A help desk or AI governance channel should answer questions before employees resort to an unapproved tool. Publish a current catalog that identifies approved tools, prohibited use cases, data classifications, exception owners, and escalation paths. Connect those rules to security awareness training for AI-era cyberthreats so employees practice decisions instead of reading policy language once.
3. Train Each Role to Change Behavior
Role-based education turns policy into repeatable judgment. Developers need prompt hygiene, code verification, dependency review, secret detection, and sandbox procedures. Executives need training on confidential board material, deepfake impersonation, approval fraud, and personal-account exposure. Finance teams should rehearse invoice verification and business email compromise (BEC) controls, while HR teams practice protecting personnel records and sensitive investigations.
Customer support and marketing teams need guidance on redacting customer data, checking factual claims, protecting brand voice, and obtaining approval before publishing AI-generated content. Occasional users need a short module covering shadow AI, data classification, safe prompts, output verification, and incident reporting. Every employee should understand that a polished answer is not evidence of accuracy, authorization, or safety.
Measure policy performance through approved-tool adoption, exception requests, personal-account access, blocked uploads, help-desk questions, incident reports, training completion, and repeat violations by role. If personal-account use rises after a ban, treat that signal as a design failure and identify the missing approved capability.
Hold monthly reviews with security, legal, privacy, IT, HR, and representative employees, then update rules, templates, permissions, and training based on the findings. That feedback loop keeps the organization's AI risk boundary visible while turning employee judgment into a measurable control.
How Should Executives Measure Shadow AI Risk Reduction and Business Value?
A shadow AI risk assessment should compare measurable exposure with business outcomes rather than count how many employees completed a policy module. Completion metrics show that an activity occurred. Behavioral and exposure metrics show whether employees use AI safely and whether organizational risk is changing.
Risk reduction requires evidence that high-risk tools were remediated, residual risk declined, and containment became faster. Business value connects those changes to productive AI use, avoided incident costs, and the opportunity cost of restricting legitimate work.
What Risk and Control Metrics Should Executives Track?
Risk and control metrics show whether governance is finding the activity that creates exposure and changing the conditions behind it. Begin with inventory coverage, calculated as discovered AI tools divided by estimated tools in use across managed browsers, identity systems, procurement records, and endpoint telemetry. Track the percentage of discovered tools with a named owner, sanctioned versus unsanctioned usage, users per tool, sensitive-data events, and high-risk tools remediated.
Control effectiveness requires more than a policy document. Measure detection rates for prohibited uploads, unauthorized account use, risky browser behavior, and access to unapproved AI services. Pair those results with policy exceptions, exception age, remediation completion, and time to contain an event. A control that detects 95% of risky activity but leaves exceptions unresolved for 90 days is not operating effectively.
Use a single operating view that includes:
- Exposure: Inventory coverage, users per tool, sensitive-data events, sanctioned versus unsanctioned usage, and high-risk tools.
- Response: High-risk tools remediated, control detection rates, policy exceptions, exception age, and time to contain.
- Behavior: Training completion, simulation results, reporting rates, repeat events, approved-tool adoption, and risky actions by role.
- Movement: Residual-risk change, risk velocity, and the rate at which new tools or users enter the environment.
Risk velocity measures how quickly exposure changes. Calculate it as the change in risk-weighted events or users over a defined period, such as 30 days. A falling risk score alongside accelerating discovery of new unapproved tools shows that controls are reducing existing exposure while adoption continues unchecked.
The NIST 2024 Generative AI Risk Management Profile treats measurement, monitoring, and governance as continuing activities rather than one-time approval steps. A human risk management program can connect those signals to role and department-level action.
How Should Executives Measure AI Business Value?
Business-value measurement starts with a use-case baseline. For each approved AI workflow, record the number of users, usage frequency, task duration before and after adoption, output-review time, implementation cost, oversight cost, incident cost, and opportunity cost.
Net business value = productivity or revenue gain − implementation cost − oversight cost − usage cost − incident cost − opportunity cost.
Implementation cost includes integration, licensing, configuration, and change management. Oversight cost includes review time, privacy assessments, model testing, legal review, and security operations. Incident cost includes investigation, containment, notification, recovery, regulatory response, and lost productivity. Opportunity cost captures value lost when a team cannot use an approved workflow or spends excessive time navigating unnecessary restrictions.
Executives should compare sanctioned-tool adoption with unsanctioned usage rather than treat lower AI activity as success. If a marketing team moves from a public chatbot to an approved enterprise tool, sensitive-data events fall and output-review time improves. The organization has reduced risk while preserving value.
Track the same use case over time by department, role, data sensitivity, and risk tier. This prevents productivity gains from concealing concentrated exposure in finance, legal, engineering, or executive teams.
What Belongs in Board-Ready Shadow AI Reporting?
Board reporting should translate operational signals into trend lines, decisions, and accountability. Show total discovered tools, inventory coverage, the share with accountable owners, sanctioned-tool adoption, and high-risk tools remediated. Add residual-risk movement, risk velocity, sensitive-data events, control detection rates, policy exceptions, and median time to contain.
Report each metric against the prior period and a defined threshold. A standalone number cannot show whether governance is improving, but a trend tied to an executive decision can show where investment or policy changes are required.
Segment trends by role, department, use case, and risk tier. A companywide risk score can remain flat while executive users generate more sensitive-data events or a product team accumulates new high-risk tools. Include training completion separately from behavior results, then connect targeted education to subsequent reporting rates, repeat-event reduction, and safer tool selection.
Employees provide the signals that make shadow AI visible. Reporting should identify where policy, approved tooling, or role-specific guidance is failing, rather than assign blame. That distinction turns employee behavior into an operating signal leaders can act on.
How Does Shadow AI Measurement Fit Human Risk Management?
Shadow AI belongs in the broader human-risk discipline because tool adoption reflects role context, workflow pressure, and perceived usefulness. An employee who uses an unauthorized tool often reveals an unmet business need, such as slow procurement, missing functionality, or unclear data-handling rules. Use those signals to prioritize targeted education, clarify approved alternatives, and route high-risk behavior into focused coaching rather than blanket restrictions.
A mature assessment links AI-use events with phishing reports, simulation behavior, training completion, data-handling decisions, and role exposure. That combined view shows whether a finance employee faces invoice-fraud and sensitive-data risks or whether an engineer needs guidance on source-code handling and model-output review.
The executive outcome consists of faster discovery, safer behavior, controlled innovation, and a measurable reduction in residual human risk. When those outcomes move together, shadow AI governance becomes a business control rather than another annual compliance exercise.
Shadow AI Risk Assessment FAQs
What Is a Shadow AI Risk Assessment?
A shadow AI risk assessment identifies unauthorized or insufficiently governed AI tools, evaluates how employees use them, and prioritizes controls for data, security, privacy, compliance, and business impact. The assessment covers public chatbots, AI meeting assistants, browser extensions, APIs, embedded features, open-source models, and locally hosted systems. It records each tool's owner, users, data flows, permissions, provider, retention terms, and business purpose.
Security, IT, privacy, legal, procurement, and business teams should combine telemetry, procurement records, employee reporting, and open-source intelligence (OSINT) to build the inventory. The lifecycle runs from discovery and scoring through treatment, monitoring, and residual-risk review.
How Is a Shadow AI Risk Score Calculated?
Calculate a baseline shadow AI risk score by multiplying probability, rated from 1 to 5, by impact, also rated from 1 to 5. A score of 20, for example, represents high priority when a tool has high data exposure and a likely event path.
Rate probability using access, user count, permissions, autonomy, vendor assurance, and control maturity. Rate impact using data sensitivity, regulatory exposure, customer reach, operational dependence, and potential harm. Display risk velocity separately, because rapidly growing use can outrank a static score.
Recalculate inherent risk before controls and residual risk after controls, documenting the owner, treatment decision, and review date.
What Should Be Included in a Shadow AI Risk Assessment Checklist?
A shadow AI risk assessment checklist should cover scope, risk appetite, discovery, inventory, use-case classification, and data-flow mapping. It should also cover vendor due diligence, contract review, scoring, treatment, approval, monitoring, incident response, and reassessment.
Retain evidence for each decision, including telemetry summaries, data classifications, vendor questionnaires, retention terms, contracts, approvals, exceptions, control tests, training records, incident tickets, and risk-acceptance statements. Confirm the tool owner, users, model, inputs, outputs, permissions, integrations, residency, subprocessors, and deletion process.
How Often Should Organizations Reassess Shadow AI Risk?
Organizations should reassess shadow AI risk at least monthly for active tools and immediately after a material change or incident. Trigger an additional review when a provider, model version, owner, data type, processing location, permission, embedded feature, user population, contract, or business purpose changes.
Review high-risk or autonomous use cases more frequently, such as monthly, when risk velocity or customer exposure is high. Track discovery coverage, sensitive-data events, control-detection rates, exceptions, residual-risk movement, and time to contain.
A scheduled review prevents an approved use case from retaining an outdated risk rating after its capabilities, integrations, or audience expand.
How Should an Organization Respond if an Employee Submits Sensitive Data to an Unauthorized AI Tool?
An organization should treat the submission as a potential data incident, contain further exposure, preserve evidence, and assess notification obligations without blaming the employee. Disable or restrict the account, revoke connected tokens, and stop repeat submissions.
Determine what data, model, provider, recipients, retention settings, and jurisdictions are involved. Ask the provider to delete the content and confirm whether it was retained, accessed, shared, or used for training.
Involve security, privacy, legal, and the data owner, document the decision, and notify affected parties or regulators when required. OAIC guidance advises safeguards before entering personal or confidential information into commercial AI tools. Build the response into a broader human-risk program that gives employees clear reporting paths and safer approved alternatives.
Build Clearer Visibility Into Human Risk From Shadow AI
Unauthorized AI use can expose sensitive data while leaving security teams without clear ownership or behavioral context. Adaptive Security connects continuous employee behavior signals with targeted training so teams can identify risky patterns and guide safer decisions. Take the Self-Guided Tour to see how it supports a broader human-risk program.
As experts in cybersecurity insights and AI threat analysis, the Adaptive Security Team is sharing its expertise with organizations.
Get started with Adaptive Security
Related articles

AI Governance Strategy: The Complete Guide to Frameworks, Implementation, and Best Practices for Enterprise Leaders

AI Governance Maturity Model: The 5 Stages, 7 Key Dimensions, and How to Build and Advance a Framework

Shadow AI Best Practices: How to Detect, Govern, and Mitigate Unsanctioned AI Tools Without Stifling Innovation
Get started