Most organisations shipping AI features in 2026 are doing it faster than their security practice has adapted. That is not carelessness. The pressure is real and the tooling is young. But it produces a recognisable pattern: a feature goes live with an API key in an environment variable, a prompt assembled by string concatenation, retrieved documents injected without provenance, and no test that asks what happens when a user tries to make it misbehave.

Secure AI-DLC is the application of secure development lifecycle discipline to these systems. It is not a new security philosophy. It is the recognition that AI features introduce three things conventional application security does not account for.

The three things that are genuinely different

1. A dependency you cannot read. You can review source code. You cannot review model weights. A model is an opaque artefact from a supply chain you mostly have to take on trust, and the provenance question (where did this come from, was it modified, can I verify it) has weaker answers than it does for an npm package.

2. A component that can be influenced by its input. Every other part of your system does what it was programmed to do. A language model does what its context persuades it to do, and that context includes content from users, documents, web pages and tool outputs. The instruction channel and the data channel are the same channel.

3. A data path that leaves your boundary. Unless you are running models yourself, every inference sends data to a third party. What goes in that request (and whether anyone reviewed it) is usually undocumented.

Everything below follows from those three.

The model supply chain

Treat model weights as a dependency with the same rigour as any other third-party component, which for most teams means considerably more rigour than they currently apply.

Questions with defensible answers:

  • Where did these weights come from, and is the source authoritative?
  • Are they pinned to a specific version, or does the reference float?
  • If a hosted API, what is the provider’s change policy, and can the model behind your endpoint change without notice, and what does that do to behaviour you have tested?
  • Is the licence compatible with your use?
  • What is recorded in your inventory? A model is an asset. Auditors will ask.

Version pinning deserves emphasis. Teams that would never deploy npm install express without a lockfile routinely point production at a model alias that silently updates. Your prompt injection tests, your output validation, your latency assumptions were all validated against one specific version. Pin it, and treat an upgrade as a change requiring re-test.

Prompt injection: an architectural problem

The OWASP Top 10 for LLM Applications puts prompt injection first, and it is worth being precise about why it resists the fixes that worked for earlier injection classes.

SQL injection was solved by separating code from data. Parameterised queries mean user input can never be interpreted as instructions, because the two travel in different channels.

A language model has no such separation. System prompt, user message and retrieved document all arrive as one sequence of tokens. There is no parameterisation primitive. You can make injection harder with delimiters, instruction hierarchies and separate models for untrusted content, and these all help, but none is a guarantee, and designing as though one is will eventually be wrong.

So the defence has to be architectural:

Assume the model can be made to produce any output. Put the security boundary on what the system is permitted to do with that output.

Concretely:

  • The model’s identity is not the user’s identity. If the model can call a tool, that tool call is authorised against the user’s permissions, checked server-side, every time. A model persuaded to request data the user cannot access must be refused by the authorisation layer, not by the prompt.
  • Consequential actions need confirmation outside the model’s control. Sending an email, moving money, deleting records. The model may propose; a human or a deterministic rule disposes.
  • Output is untrusted input to whatever consumes it. Model output rendered as HTML is an XSS sink. Model output passed to a shell is command injection. Model output used in a query is injection. Escape and validate exactly as you would user input, because that is what it is.
  • Indirect injection is the harder case. The attack does not have to come from your user. A document in your knowledge base, a web page your agent fetches, or an email it summarises can carry instructions. Anything retrieved is untrusted, and its provenance should travel with it.
The boundary is on what the output is allowed to do
The trust boundary in an application built on a language modelSystem prompt, user message and retrieved documents all arrive at the model as one sequence of tokens, with no parameterisation primitive separating instructions from data. Assume the model can be made to produce any output. Every tool call is authorised server-side against the user’s permissions rather than the model’s, consequential actions are confirmed outside the model’s control, and output is escaped and validated as the untrusted input it is before it reaches a browser, a shell or a query.ALL OF THIS ARRIVES AS ONE SEQUENCE OF TOKENSENFORCED HERE, EVERY TIMEoutputSystem promptUser messageuntrustedRetrieved documentsuntrusted, and the harder caseThe modelno separation of instructions from dataassume it can be made to produce any outputThe security boundaryon what the output is permitted to doTool calls authorised as the userserver-side, not by the promptConsequential actions confirmeda human or a deterministic ruleOutput escaped and validatedit is untrusted input to whatever consumesit
Why not fix the prompt
there is no parameterisation primitive to fix it with
The model’s identity
is not the user’s identity
Indirect injection
anything retrieved is untrusted, and provenance travels with it
This diagram as text
  • All of this arrives as one sequence of tokens
    • System prompt
    • User message — untrusted
    • Retrieved documents — untrusted, and the harder case
  • The model — no separation of instructions from data — assume it can be made to produce any output
  • The security boundary — on what the output is permitted to do
  • Enforced here, every time
    • Tool calls authorised as the user — server-side, not by the prompt
    • Consequential actions confirmed — a human or a deterministic rule
    • Output escaped and validated — it is untrusted input to whatever consumes it

Relationships

  • System prompt → The model
  • User message → The model
  • Retrieved documents → The model
  • The model → The security boundary — output
  • The security boundary → Tool calls authorised as the user
  • The security boundary → Consequential actions confirmed
  • The security boundary → Output escaped and validated

The data boundary

For every AI feature, someone should be able to answer: what data leaves our boundary on a request, and who reviewed that?

In practice the answer is often more than intended, because retrieval-augmented systems assemble context dynamically. A prompt that in testing contained a short question contains, in production, a customer record, an internal document and a conversation history.

Controls that apply:

  • Data classification before retrieval, so the retrieval layer knows what it may include
  • Minimisation: include what the task needs, not everything semantically similar
  • Redaction of identifiers where the task does not need them
  • Access control on the retrieval layer itself, enforcing the requesting user’s permissions. A RAG system that indexes everything and retrieves without permission checks is a data leak with a search interface.
  • Contractual and residency review of the provider, which for regulated clients is where the conversation frequently ends up: if customer data cannot leave the jurisdiction, a hosted model in another one is not available to you regardless of its quality.

This last point is why AI strategy and infrastructure strategy are the same conversation for regulated organisations. Self-hosted models on owned infrastructure are sometimes not a preference but the only lawful option.

Testing something non-deterministic

Conventional testing assumes the same input produces the same output. It does not here, which makes people conclude these systems cannot be tested. They can, but the tests assert differently.

Test type What it asserts
Prompt injection suite A corpus of known injection patterns, asserting the system does not take the prohibited action
Indirect injection Documents with embedded instructions, asserting retrieval does not escalate
Data leakage Attempts to extract system prompt, other users’ data, or training content
Authorisation The model asks for something the user may not have; the system refuses
Output handling Adversarial output (HTML, SQL, shell metacharacters) is escaped by the consumer
Regression on upgrade The suite re-runs on any model version change

The key shift: assert on system behaviour, not on model output. “The model refuses” is a fragile assertion. “The tool call is rejected by the authorisation layer” is a durable one, and it is the one that actually protects you.

Also log enough to investigate: prompts, retrieved context references, tool calls, and outputs, with retention that matches your incident and regulatory needs, while making sure the logs themselves do not become the data leak, since they now contain everything sensitive that passed through.

Where this fits with the frameworks

  • OWASP Top 10 for LLM Applications: the most directly actionable list for engineers; start here.
  • NIST AI RMF (AI 100-1): a governance frame for identifying and managing AI risk; useful for the conversation with risk and legal.
  • ISO/IEC 42001: an AI management system standard, structurally similar to ISO 27001. Expect clients and auditors to start asking about it.

None of these replaces your existing application security practice. AI features still need authentication, authorisation, input validation, dependency scanning, secrets management and logging. Secure AI-DLC is an extension, not a substitute, and a team with weak fundamentals will not be saved by AI-specific controls layered on top.

A practical starting sequence

  1. Inventory. Which features call a model, which model, and through whose API key? Most teams cannot answer this today.
  2. Map the data path for each. What leaves the boundary?
  3. Check authorisation. Can any model-initiated action bypass a permission check? This is the highest-severity class and the most common finding.
  4. Pin model versions and treat changes as changes.
  5. Write a prompt injection suite, even a small one, and run it in CI.
  6. Review output handling at every consumer.
  7. Set retention and access rules for AI logs, which are now sensitive.

Steps 3 and 6 find the most severe issues for the least effort. Start there.

An honest note on maturity

This is an emerging discipline. The tooling is immature, the standards are young, and anyone claiming a complete solution is overstating what exists. What can be done today is the list above, which is mostly rigorous application of principles the industry already understands, applied to a component that behaves differently from the rest of the stack.

STP builds this capability deliberately and describes it as it is: we can inventory your AI data paths, review authorisation boundaries and output handling, build injection test suites, and design the infrastructure for self-hosted inference where residency requires it. We would rather tell you where the discipline is still developing than sell certainty that does not exist yet.

More on Secure SDLC and Secure AI-DLC, or start a conversation.