The pattern is the same one we lived through with web apps in 2008 and with mobile in 2014. A new capability arrives, teams ship as fast as they can to capture upside, and the security model lags by several quarters. This time, the lag is worse, because "the model" itself is a moving target, and the orchestration code that wraps it changes every week.
This is a tour of the actual incidents we have seen in production AI deployments over the last twelve months, organized by attack surface. The point is not to scare you off shipping AI features. It is to make the surface visible enough that you can defend it.
1. The orchestration layer is the actual attack surface
Most "AI applications" are 5% model and 95% glue: API calls to a hosted model, a prompt template, a retrieval step against a vector database, some output post-processing, and a streaming response back to the user. Every line of that glue is regular Python or TypeScript, and every classical vulnerability still applies.
We see the same five problems over and over:
- Hardcoded inference keys.
OPENAI_API_KEYcommitted to a public repo, or pasted into a Slack channel that integrates with a third-party automation. Once leaked, an attacker can run inference on your dime and (worse) read the contents of any pinned context. - Server-side request forgery from tool-use loops. If your model can call URLs ("fetch the latest article from this link"), and you do not allow-list those URLs, the model will happily fetch
http://169.254.169.254/latest/meta-data/when prompted to. - Vector database exposed to the public internet. Qdrant, Pinecone, Weaviate, all wonderful, all routinely deployed without auth because "it's behind the API". We have found unauthenticated indexes containing entire customer support archives.
- Streaming response logs containing PII. When the model echoes user input, and your observability stack logs the response verbatim, you are now storing PII in Datadog. Without a DPA. Probably without consent.
- Outdated
transformersorlangchainversions. Both have shipped CVEs in 2025 that allowed remote code execution via crafted model files. Both are routinely pinned to "whatever main was when we started".
2. Prompt injection is the new SQL injection
The framing is exactly the same: untrusted input is concatenated into a sensitive operation, and the underlying system has no way to distinguish "instructions from the developer" from "instructions from the user". The fix is the same too: separation of channels, parameterization, and never trust the model to enforce its own policy.
"We assumed our system prompt would survive contact with the user. It survived for about four minutes after launch." Engineering post-mortem, fintech customer, March 2026
Concrete fixes that work:
- Use the model provider's role separation properly. System prompts go in the
systemfield, never concatenated intouser. Customer-supplied input is always quoted, marked as untrusted, and bounded. - Constrain tool use with explicit allow-lists. If the model can call functions, declare the set, validate every argument, and never let the model decide which endpoint to hit.
- Add a guardrail pass. A small classifier between the model's output and the user that filters obvious leakage (PII, credentials, system-prompt fragments). It is imperfect; it catches most cases.
- Rate-limit on prompt length and tool-call depth. Most jailbreaks involve long, recursive prompts. A simple cap defeats 80% of them.
3. Supply chain: where the model itself becomes the vulnerability
Pre-trained model weights distributed as .bin or .safetensors files are code. The deserialization step in PyTorch can execute arbitrary Python. Hugging Face flags this, but mid-2025 saw at least four public incidents where a popular community model contained a tampered weight file that downloaded a cryptominer on load.
The remediation is identical to what we do for npm packages:
- Pin versions. Never "latest".
- Verify checksums against a trusted source.
- Sandbox the loader. If you load a model in production, the process loading it should not have credentials to your data store.
- Subscribe to advisory feeds for the model hubs you use.
The same scanner, pointed at the AI-shaped half of your repo.
Our scanner already inventories every dependency in your codebase, including transformers, langchain, llama-index, openai, anthropic, vector-DB clients, and the orchestration libraries between them. When a CVE drops affecting any of them, you get the same plain-English alert and the same auto-PR you would for any other package.
For the AI-specific surface, we add three checks on top:
- Secret scanning targeted at inference-key patterns (
sk-*,nvapi-*,r8_*, etc.) across both repos and orchestration configs. - Detection of unauth'd vector-DB endpoints reachable from your cloud account.
- SBOM-style inventory of model weights pinned in your repo or pulled at deploy time, cross-checked against known-bad hashes.
4. The compliance angle nobody talks about
GDPR Article 22 covers automated decision-making. The EU AI Act covers high-risk systems. NIS2 covers operational technology. Most teams shipping AI features have not mapped which apply to them.
A short heuristic: if your AI output materially affects a customer's access to a service, a price they pay, or a decision about them, you are probably in scope for at least one of these. The documentation cost of "we use AI here" is no longer zero.
5. What to do this week
- Grep your repos for hardcoded inference keys. There is at least one.
- Audit which model libraries you actually use, and pin them.
requirements.txtwith==, not>=. - Put any vector database behind your VPC. If it has a public IP, it should not.
- Write down which AI features fall under which regulation. Even a one-page table is more than most teams have.
- Treat your prompt as code. Version it, test it, and assume it will be leaked.
AI is not magically dangerous, but it is conventionally dangerous in ways teams have not internalized yet. The defenses are the same boring ones that worked for the previous wave, inventory, patching, separation of privilege, secret hygiene. If you do those four things consistently, you are in the top quartile.