An assistant may answer most product questions correctly while mishandling one request involving another customer's account. Averaging these outcomes into a single quality score hides the distinction between an imperfect answer and an authorization failure.
Security evaluation asks whether the application respects defined boundaries under specified conditions. Accuracy evaluation asks whether an answer matches the expected information. A release process can use both, with separate acceptance criteria and evidence.
Define the prohibited outcome
NIST's Generative AI Profile discusses risk management across the AI lifecycle, including measurement and evaluation. It provides a framework for organizing risks rather than a universal passing score for every product. The NIST profile supplies that broader context.
For an illustrative account assistant, prohibited outcomes might include retrieving a second customer's records, executing an unapproved account change, or displaying a secret from a tool response. Each outcome can have a specific observation point. Some are visible in the answer; others require inspecting retrieval and service logs.
Build cases around business permissions
Use synthetic users with distinct roles and data. Include ordinary authorized tasks, unauthorized requests, ambiguous requests, and external content that attempts to redirect the workflow. The expected result can be a completed task, a request for missing information, or a service-side denial.
Keep the permission model fixed while changing the phrasing. This helps distinguish a brittle conversational response from an enforced control. A user asking indirectly for another account's record does not gain access merely because the request resembles a normal support question.
Record the complete configuration
An evaluation result belongs to a model version, system instructions, tool definitions, permissions, retrieval index, and application build. If any of those change, the result may no longer describe the deployed system. Store enough configuration detail to reproduce the run without putting credentials or customer data into the test package.
Also record repeated trials where behavior is variable. One successful refusal is an observation, not an estimate of a universal failure rate. Reporting the number of cases, number of trials, and observed outcomes makes the limits of the evaluation visible.
Separate severity from frequency
An authorization failure in a rarely exercised case can still require release attention. A high frequency of harmless formatting errors may instead affect usability. The product owner can decide on separate release conditions for these categories rather than allowing a large set of easy questions to dilute a significant security result.
For a scoped assessment, the deliverable can list tested boundaries, failed cases, service traces, and retest results. It can also state which tools, languages, document formats, or user roles were excluded. Those exclusions describe remaining coverage, not evidence that the untested areas are safe.
The resulting release discussion becomes concrete: what the assistant was allowed to do, which conditions were tested, and whether the application enforced them. A benchmark score alone cannot answer those questions.
Sources
OWASP: AI Agent Security. The account-assistant evaluation is an illustrative application of lifecycle risk assessment.