← Blog / AI security
AI security

Everyone is building with AI. Almost no one is protecting it.

Prompt injection, model theft, leaked API keys in inference layers, training-data poisoning, the attack surface around AI shipped to production is larger than the surface AI itself protects.

Prompt injection

External input rewrites the model's instructions. The 2025 OWASP LLM Top-10 puts this at #1, ahead of every classic web vuln.

Leaked credentials

API keys, vector-DB tokens, and inference endpoints hardcoded into orchestration code. The most common AI-related leak we see.

Supply-chain poisoning

Malicious model weights and tampered datasets distributed through public hubs. torch-load deserialization is the new pickle.

Data exfiltration via RAG

Retrieval-augmented generation pipelines leak training data and customer documents through clever follow-up prompts.

The pattern is the same one we lived through with web apps in 2008 and with mobile in 2014. A new capability arrives, teams ship as fast as they can to capture upside, and the security model lags by several quarters. This time, the lag is worse, because "the model" itself is a moving target, and the orchestration code that wraps it changes every week.

This is a tour of the actual incidents we have seen in production AI deployments over the last twelve months, organized by attack surface. The point is not to scare you off shipping AI features. It is to make the surface visible enough that you can defend it.

1. The orchestration layer is the actual attack surface

Most "AI applications" are 5% model and 95% glue: API calls to a hosted model, a prompt template, a retrieval step against a vector database, some output post-processing, and a streaming response back to the user. Every line of that glue is regular Python or TypeScript, and every classical vulnerability still applies.

We see the same five problems over and over:

2. Prompt injection is the new SQL injection

The framing is exactly the same: untrusted input is concatenated into a sensitive operation, and the underlying system has no way to distinguish "instructions from the developer" from "instructions from the user". The fix is the same too: separation of channels, parameterization, and never trust the model to enforce its own policy.

"We assumed our system prompt would survive contact with the user. It survived for about four minutes after launch." Engineering post-mortem, fintech customer, March 2026

Concrete fixes that work:

  1. Use the model provider's role separation properly. System prompts go in the system field, never concatenated into user. Customer-supplied input is always quoted, marked as untrusted, and bounded.
  2. Constrain tool use with explicit allow-lists. If the model can call functions, declare the set, validate every argument, and never let the model decide which endpoint to hit.
  3. Add a guardrail pass. A small classifier between the model's output and the user that filters obvious leakage (PII, credentials, system-prompt fragments). It is imperfect; it catches most cases.
  4. Rate-limit on prompt length and tool-call depth. Most jailbreaks involve long, recursive prompts. A simple cap defeats 80% of them.

3. Supply chain: where the model itself becomes the vulnerability

Pre-trained model weights distributed as .bin or .safetensors files are code. The deserialization step in PyTorch can execute arbitrary Python. Hugging Face flags this, but mid-2025 saw at least four public incidents where a popular community model contained a tampered weight file that downloaded a cryptominer on load.

The remediation is identical to what we do for npm packages:

What CyberDebunk does for AI workloads

The same scanner, pointed at the AI-shaped half of your repo.

Our scanner already inventories every dependency in your codebase, including transformers, langchain, llama-index, openai, anthropic, vector-DB clients, and the orchestration libraries between them. When a CVE drops affecting any of them, you get the same plain-English alert and the same auto-PR you would for any other package.

For the AI-specific surface, we add three checks on top:

  • Secret scanning targeted at inference-key patterns (sk-*, nvapi-*, r8_*, etc.) across both repos and orchestration configs.
  • Detection of unauth'd vector-DB endpoints reachable from your cloud account.
  • SBOM-style inventory of model weights pinned in your repo or pulled at deploy time, cross-checked against known-bad hashes.

4. The compliance angle nobody talks about

GDPR Article 22 covers automated decision-making. The EU AI Act covers high-risk systems. NIS2 covers operational technology. Most teams shipping AI features have not mapped which apply to them.

A short heuristic: if your AI output materially affects a customer's access to a service, a price they pay, or a decision about them, you are probably in scope for at least one of these. The documentation cost of "we use AI here" is no longer zero.

5. What to do this week

  1. Grep your repos for hardcoded inference keys. There is at least one.
  2. Audit which model libraries you actually use, and pin them. requirements.txt with ==, not >=.
  3. Put any vector database behind your VPC. If it has a public IP, it should not.
  4. Write down which AI features fall under which regulation. Even a one-page table is more than most teams have.
  5. Treat your prompt as code. Version it, test it, and assume it will be leaked.

AI is not magically dangerous, but it is conventionally dangerous in ways teams have not internalized yet. The defenses are the same boring ones that worked for the previous wave, inventory, patching, separation of privilege, secret hygiene. If you do those four things consistently, you are in the top quartile.

Scan your AI stack with the same workflow.

Your AI orchestration code is just code. CyberDebunk inventories it, watches for new CVEs across model libraries, secret-scans your repos, and drafts fix PRs, in the same dashboard as the rest of your stack.