Back to Blog
Guides

AI Service Identity Monitoring for SMBs After the Azure AI Foundry Auth Bypass

Managed AI platforms run on non-human identities that hold standing access to documents, search indexes and databases, and a platform side flaw like the CVSS 10.0 Azure AI Foundry bypass leaves nothing for the customer to patch. A CISSP walk through scoping those identities and monitoring them by shape and volume.

By Danny Mercer • Oct 8, 2026 •1 views
AI service identity monitoringnon-human identityLLMjackingmanaged SOCAzure AI Foundry security
Share:

There is a particular kind of advisory that leaves a security team with nothing to do and a lot to think about. On September 17, 2026, Microsoft published CVE-2026-85889, a missing authentication flaw in Azure AI Foundry with a CVSS score of 10.0, the ceiling, the number you only see when somebody forgot to put a lock on a door that leads to every other door. Microsoft mitigated it on its own infrastructure, reported no exploitation, and there was nothing for customers to patch. Two weeks later GitLab shipped a fix for CVE-2026-90970, a 9.9 template injection bug in its self-hosted AI Gateway that let an authenticated Duo Agent Platform user escape the prompt template sandbox and run commands on the gateway host. One of those you can patch. One of those you cannot. Both point at the same uncomfortable truth, which is that most small and mid-sized businesses have quietly created a new class of privileged identity in the last eighteen months, and almost none of them are watching it. AI service identity monitoring is the discipline that closes that gap, and it is a lot less exotic than the marketing around AI security would have you believe.

Here's the thing. The model is not the asset. The model is a very expensive autocomplete with a nice personality. The asset is everything the model has been trusted to reach, and in a typical deployment that list includes a storage account full of source documents, a search index built from those documents, a database connection for the "ask questions about our customers" feature, a Key Vault reference or two, and an inference key that anyone holding it can spend against your cloud bill. All of that hangs off an identity that is not a person, does not use MFA, never takes a vacation, and in most environments does not appear anywhere in the asset inventory.

The managed AI endpoint is a non-human identity with a data connection

When a team builds a retrieval assistant on Azure AI Foundry, Amazon Bedrock, or Google Vertex AI, the platform wires the pieces together with a service identity. On Azure that is usually a managed identity or a service principal. On AWS it is an IAM role assumed by the service. The identity is granted read access to the document store so the retrieval step works, often granted write access to the search index so the ingestion job works, and sometimes granted far more because the proof of concept stalled on a permission error at 6 PM on a Friday and somebody clicked Contributor at the subscription scope to make the demo happen.

That identity is what CISSP coursework would call a subject, and it holds standing access to objects that used to sit behind separate controls. HR documents lived in one SharePoint library with its own permissions. Contracts lived somewhere else. Customer records lived in a database with a service account only the line of business application used. The AI project pulls all three into one index so the assistant can answer questions across them, and in doing so it collapses three separately protected boundaries into one identity with one set of keys. That is not a bug. That is the feature the business asked for. It is also why the OWASP Top 10 for LLM Applications 2025 ranks Sensitive Information Disclosure (LLM02) and Excessive Agency (LLM06) so high (https://genai.owasp.org/llm-top-10/). The model does not need to be jailbroken for data to leak if the identity behind it can read everything.

The control objective here is old and boring and correct. NIST SP 800-53 Rev. 5 CM-8 says you inventory your system components, and AC-2 says you manage your accounts, including the ones that are not people (https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final). In practice, the AI service principal rarely shows up in either. It was created by a developer inside a resource group that also gets created and deleted on a whim, it was never reviewed in a quarterly access recertification because recertification lists come from the HR system, and its permissions were assigned once and never revisited. If you want a quick reality check, pull a list of every service principal and managed identity in your tenant that holds a role on a storage account or database, then ask which of them belongs to an AI project. Most teams find at least one they did not know about, and the one they did not know about usually has the broadest role.

Why a platform side auth bypass leaves nothing for the customer to patch

CVE-2026-85889 belongs to a lineage worth knowing. In August 2021 Wiz disclosed ChaosDB, a flaw in the Jupyter Notebook feature of Azure Cosmos DB that exposed the primary read and write keys of thousands of customer databases (https://www.wiz.io/blog/chaosdb-how-we-hacked-thousands-of-azure-customers-databases). In 2022 Orca Security disclosed AutoWarp, a flaw in Azure Automation that let one tenant obtain tokens for other tenants' managed identities. In 2023 the Storm-0558 intrusion used a stolen Microsoft signing key to forge authentication tokens and read Exchange Online mail at roughly two dozen organizations, a story the Cyber Safety Review Board later described as the result of a cascade of avoidable errors (https://www.cisa.gov/resources-tools/resources/CSRB-review-summer-2023-microsoft-exchange-online-intrusion).

None of those had a customer side patch. In every case the vendor fixed the platform and the customer was left with two questions the vendor could not answer for them. Was my tenant touched during the exposure window, and if it was, what could the attacker reach? The first question is answered by logs you own. The second is answered by how tightly you scoped the identity. The patch, the thing most vulnerability management programs are organized around, simply does not exist.

That is the shift a lot of SMB security programs have not made yet. A vulnerability management process that measures success by time to patch has nothing to measure when the vulnerable code runs on Microsoft's servers. The CVE gets a ticket, the ticket gets closed with "vendor mitigated, no action," and everyone moves on. The more honest closure note would read "vendor mitigated, exposure window unknown, tenant activity during the window not reviewed." Nobody likes writing that note, which is exactly why it is worth writing.

The GitLab bug is the self-hosted twin of the same problem. The AI Gateway sits between your source code and your model providers, and it holds the credentials it needs to talk to both. Code execution on that host is a foothold on a machine that, by design, sees prompts containing proprietary code and holds keys that can spend money. You can patch that one, and you should, but the patch only closes the door. It does not tell you whether anyone walked through it before October 2.

What an attacker actually does with an AI service identity

Strip away the novelty and the attack path looks a lot like any other cloud identity compromise. The interesting part is how short it is.

Initial access usually comes from one of three places. The first is a leaked inference key, committed to a repository, pasted into a front end bundle, or left in a notebook that got shared to the wrong channel. That maps cleanly to MITRE ATT&CK T1552.001, credentials in files. The second is a stolen token or a compromised service principal secret, which is T1528 and T1078.004, valid cloud accounts. The third is a platform flaw like the ones above, where the attacker never needs your credentials at all because the authentication check that would have demanded them never ran.

From there the attacker has two businesses to choose from, and frankly some of them run both. The first is data theft. The AI identity already has read access to the document store and the search index, so there is no privilege escalation phase and no lateral movement phase. The attacker queries the index, reads the storage account, and leaves with the contents (T1530, data from cloud storage). If the project has a database connection for structured questions, that connection was provisioned with real credentials, and those credentials work outside the AI service just as well as inside it.

The second business is resource theft, which the industry has started calling LLMjacking. In May 2024 the Sysdig Threat Research Team documented attackers using stolen cloud credentials to access hosted model services and resell that access, and estimated that a victim could rack up more than $46,000 per day in consumption charges if the abuse ran unchecked (https://sysdig.com/blog/llmjacking-stolen-cloud-credentials-used-in-new-ai-attack/). ATT&CK files this under T1496, resource hijacking, and OWASP calls the same exposure Unbounded Consumption (LLM10). For a forty person company in McKinney with a modest Azure commitment, a weekend of somebody else's role play chatbot traffic on your inference key is not a security incident in the traditional sense. It is a finance incident that arrives as an invoice, and it is usually the first sign anyone notices.

Here's what makes both paths nasty from a detection standpoint. Every action the attacker takes is something the AI service identity is supposed to do. It is supposed to read the storage account. It is supposed to query the index. It is supposed to call the model. Endpoint detection sees nothing because there is no endpoint. A rule that fires on "service principal reads storage" would fire every few seconds. The attack is indistinguishable from the feature at the level of individual events, and only becomes visible at the level of shape and volume.

Separating the AI identity from everything else it touches

Before monitoring can work, the identity has to be scoped tightly enough that anomalies are visible. A service principal with Contributor across the subscription generates so much legitimate noise that no baseline survives contact with it.

The first move is AC-6, least privilege, applied with some imagination. The identity that runs inference should not be the identity that runs ingestion. Inference needs read access to the search index and nothing else. Ingestion needs read access to the source documents and write access to the index, and it runs on a schedule, which means its activity outside that schedule is an alert all by itself. Neither identity should share credentials with the analytics team's service account or the line of business application's database login, because once they are shared you can no longer tell whose traffic is whose. Splitting them does two things at once. It shrinks the blast radius of any single compromise, and it turns cross purpose activity into a signal. An inference identity writing to the index is wrong by definition, and wrong by definition is the best kind of detection rule.

The second move is getting rid of static keys wherever the platform allows it. Azure AI services and Azure OpenAI resources support disabling local key based authentication entirely so that every call has to present an Entra ID token (https://learn.microsoft.com/en-us/azure/ai-services/disable-local-auth). AWS and Google have equivalent paths through IAM roles and workload identity federation. A key that does not exist cannot be committed to GitHub, and a token that expires in an hour is a much smaller prize than a key that has been valid since the pilot. The NSA-led joint guidance Deploying AI Systems Securely, published with CISA and FBI in April 2024, makes the same point in more formal language about protecting the credentials that sit around the model rather than just the model itself (https://www.cisa.gov/resources-tools/resources/deploying-ai-systems-securely).

The third move is SC-7 boundary protection. The AI project's data sources should be reachable over private endpoints rather than public storage URLs, and the AI resource itself should not accept inference traffic from the entire internet unless the product actually requires that. When the platform side auth check fails, network reachability is the control that still holds. It is the same argument we have made about edge appliances for years. If the bug runs before authentication, the only lever left is who can reach the front door.

Building AI service identity monitoring that watches shape, not signatures

You cannot write a signature for CVE-2026-85889 that runs in your tenant. You can write detections for what an attacker would do with the access it granted, and those detections work against the next platform bug too. That is the core idea behind AI service identity monitoring, and it rests on three telemetry sources most organizations already pay for and few actually look at.

The first source is the identity provider. Entra ID records service principal and managed identity sign-ins in their own log categories, separate from the interactive user sign-ins that most SOC dashboards are built around. Those logs show which identity authenticated, from what IP address, to what resource, and how often. An AI service principal has an extremely regular profile. It authenticates from the platform's own address ranges, to the same handful of resources, at a rate that tracks business hours. A sign-in from a residential ISP, a new resource audience, or a sudden spike at 2 AM on a Sunday is exactly the kind of deviation a human analyst can triage in minutes. Storm-0558 is the cautionary tale here. The organization that first spotted the forged token activity did so because it had the right mail access auditing turned on and someone was reading it.

The second source is the data side. Storage accounts, search services, and databases all produce access logs, and those logs should be keyed to the AI identities specifically. The question is not "did the AI identity read a blob," because it reads thousands a day. The question is whether the volume, the breadth, or the pattern changed. An inference identity that normally touches forty documents an hour and suddenly enumerates every container in the account is behaving like a person running a script, not like a retrieval pipeline answering questions. Same with a database connection that normally runs a handful of parameterized queries and suddenly runs a SELECT across a table it has never touched.

The third source is consumption. Token counts and request volume are a detection signal, not just a billing metric. Most platforms expose per deployment usage metrics, and setting an alert threshold at two or three times the normal daily peak catches LLMjacking within hours instead of at the end of the month. It also catches the more mundane failure where a developer's test harness gets stuck in a loop, which is not an attack but still costs real money and still deserves a phone call.

Two more pieces turn this from a dashboard into a control. AU-9, protection of audit information, says the logs have to live somewhere the monitored identity cannot write. If the AI project's diagnostic settings can be changed by the same identity that runs inference, an attacker can turn off the logging before doing anything interesting. Forward the logs to a workspace or storage account under separate ownership. And SI-4, system monitoring, has to include absence of telemetry as a condition. If the AI identity's sign-in logs or the storage access logs simply stop arriving, that silence is an event. It is either a broken pipeline or somebody covering their tracks, and either one is worth knowing about before Monday.

The NIST Generative AI Profile, AI 600-1, frames much of this under information security and value chain risk, and it is a useful document to hand to a leadership team that wants to know why the AI rollout needs a monitoring budget (https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf). The short version for the CFO is that the AI project concentrated access, and concentrated access needs concentrated watching.

What to do this quarter

None of this requires a new product category. It requires treating the AI service identity like what it is, which is a privileged account with a data connection that happens to answer questions in complete sentences.

Start with inventory. List every AI project, model deployment, and gateway in your environment, including the self-hosted ones and the ones a developer stood up in a personal sandbox subscription that is somehow billed to the company card. For each one, write down which identity it runs as, what roles that identity holds, and on which resources. That list becomes your CM-8 entry and your AC-2 review scope in one pass.

Then scope down. Split inference from ingestion, remove subscription level roles, disable local key authentication where the platform supports it, rotate any keys that predate the change, and move data source access onto private endpoints. If a vendor advisory like CVE-2026-85889 lands during the exposure window, pull the sign-in and data access logs for the AI identities for the affected period and actually read them, then write a closure note that says what you checked rather than "vendor mitigated."

Then watch. Baseline sign-in source, resource audience, data access volume, and token consumption for each AI identity, alert on deviation, forward the logs off platform, and alert on silence. Patch the self-hosted pieces like GitLab's AI Gateway on a written clock, because the one bug you can patch is the one auditors and insurers will ask about first.

The good news, and there is some, is that AI service identities are among the most predictable subjects in any environment. They do the same thing all day, every day, from the same places. That predictability is what makes AI service identity monitoring so effective once someone bothers to set it up. The attacker who compromises one has to act differently than the identity normally acts, and different is exactly what a well tuned baseline is built to see.

This is the kind of watching that CyberSphere's Managed SOC module was built to do without adding a sixth console to somebody's morning. It pulls identity provider sign-ins, cloud data access logs, and consumption signals into one place, baselines non-human identities like the service principal behind an AI project alongside the human accounts, and ties every detection to a remediation ticket in the same platform, with 24/7/365 coverage and dark-web breach monitoring that catches a leaked inference key before somebody else's chatbot finds it. Partners who want to deliver it under their own brand can white-label the client portal. The platform lives at https://cybersphere.thecyberone.com/, and the next platform side CVSS 10.0 is coming whether anyone has a patch for it or not, so the time to know what normal looks like is before the advisory lands.

Frequently Asked Questions

How quickly does Innovation Network Design respond to a security incident?

Our SOC triages and notifies within 15 minutes with confirmed details and containment steps, rather than handing you a queued alert. Incident response retainer clients carry a 2-hour guaranteed response.

Do you only work with businesses in the Dallas-Fort Worth area?

We are based in McKinney, Texas and work on-site across Plano, Allen, Frisco and the wider DFW metroplex. Our monitoring and response operate remotely, so we also support organizations nationwide.

How do we get started?

Start with a free assessment. Contact us or call 512-518-4408 and we will review your current setup and give you clear, prioritized recommendations before you commit to anything.

Need Help With This?

Innovation Network Design helps businesses across McKinney, Dallas, and nationwide with expert cybersecurity services.

D

Danny Mercer

Innovation Network Design

With nearly a decade in cybersecurity and IT infrastructure, our team delivers expert insights to help businesses in McKinney, Dallas, and across DFW make informed security decisions. Have a question? Get in touch.

Ready to Secure Your Business?

Get a free security assessment and find out where your organization stands.