Tervaq iconTervaq

AI assurance

AI-enabled fraud: investigate the workflow, not the label

A practical way to assess claims of AI-assisted fraud without confusing model capability, automation, and a demonstrated control failure.

‘Built with AI’ tells an investigator very little about the control that failed. It may describe how code was written, how a message was produced, or how an operation was coordinated. Those are different questions, and they need different evidence.

Be precise about the role of AI

Separate the claimed role of a model from the behavior of the resulting system. Was AI used to draft text? Did it generate software that someone later operated? Was a deployed agent making decisions and calling tools? Or is AI merely part of the story told by a source? The report should say which interpretation the evidence supports.

This distinction changes what you assess. Reviewing generated software is not the same as assessing an agent that can act on a live account. A model name, a screenshot of a chat, or a source’s estimate of development time does not establish the reliability, scale, or impact of the resulting operation.

Follow the action across the boundary

For an AI-enabled application, focus on where an output can become an action. OWASP’s agent-security guidance identifies risks including tool abuse, excessive privileges, indirect prompt injection, and data exposure. Those risks become concrete when the review names the tool, the permission, the data it can reach, and the action it can take.

Draw a short sequence from incoming material to the final operation. Mark where trust is established, where an authorization check is enforced outside the model, and where a person can review a consequential action. A statement in a system prompt is not equivalent to an independently enforced permission boundary.

Reference: [1] OWASP

Measure outcomes rather than impressive demonstrations

Define a permitted outcome and a prohibited outcome before testing. For a customer-support agent, that might mean distinguishing an allowed explanation from an unauthorized account change. For an internal research assistant, it might mean distinguishing a useful summary from disclosure of material it was not permitted to read.

A test that produces an unusual answer is not necessarily a security failure. Conversely, a polite refusal followed by an unauthorized tool action is not a successful defense. Keep the model’s words and the system’s observable actions in separate fields. Grade the outcome against the agreed boundary.

Do not put the whole defense into one prompt

OWASP’s prompt-injection guidance presents multiple defenses, including separation of untrusted content, output validation, least privilege, and human review for sensitive operations. It does not make a single filter a substitute for the application’s security design.

Our assessment approach is to ask what happens after each control fails. Can a misleading document expand a tool’s permissions? Can an untrusted output choose a recipient or payment destination? Can a review step show the reviewer the exact action rather than an abbreviated description? These are design questions that can be answered through scoped tests without distributing an operational abuse tool.

Reference: [2] OWASP

Keep the evaluation reproducible

Record the application release, relevant configuration, model identifier where known, permissions, test fixture, and observed action. State whether the run was repeated and which sources of variability remain. Preserve the conditions needed for another authorized reviewer to understand the result.

When a mitigation changes, rerun the protected workflow and its legitimate alternatives. Do not silently turn a narrower permission set into a claim that every prompt-injection risk has been solved. A well-bounded result is more useful than a broad assertion of safety.

Questions to take into the review

  • Which behavior is attributable to the model, and which to the surrounding software?
  • Where is authorization enforced independently?
  • What actual action would count as a failure?

The label is not the finding

AI can be a relevant part of the threat model without being the root cause of a particular incident. The useful output is an explanation of the workflow, the failed boundary, and the evidence needed to validate a response. That is where research becomes a decision an enterprise team can act on.

References & scope

This article combines cited public guidance with Tervaq’s proposed assessment approach. It is not a report of a named organization’s incident or a claim of independent certification.

  1. AI Agent Security Cheat Sheet OWASP
  2. LLM Prompt Injection Prevention Cheat Sheet OWASP

For corrections or substantive questions, contact contact@tervaq.com.