Test typed and pasted AI prompts separately
A successful paste test does not establish what happens when an employee types the same content. Evaluate the input method, policy decision and final submission separately, using an approved test account and synthetic data with a known expected classification.
For Security engineers running an employee AI pilot
One support summary, three input methods
A pilot user asks an AI assistant to summarize a fictional support ticket. The evaluator changes only how the same test text enters the composer.
What you are working with
- A synthetic ticket containing a vendor-approved sensitive-data fixture and an obvious TEST ONLY label.
- An otherwise equivalent ticket containing public information that should be allowed under the selected policy.
- A recorded browser version, extension version, policy configuration and signed-in AI account type.
A safer approach
- Use an isolated test conversation and confirm the fixture's expected classification before attempting submission.
- Type, paste as plain text and paste formatted text in separate runs with the same policy.
- Keep screenshots and outcomes in the evaluator's worksheet; do not assume the security product records input method.
Expected outcome: Each in-scope method produces the agreed decision, and the evaluator can distinguish an intercepted submission from a warning followed by a completed send.
Work through the procedure
Fix the policy and fixture
Write down the required action for the synthetic restricted text and the allowed control. Ask the vendor which detector and sample are appropriate. An invented account number that fails a checksum is a poor positive test because the expected classification is already uncertain.
Exercise each input method
Start a fresh conversation for each run. Test manual typing, plain-text paste and formatted paste without changing browser, account or policy. Observe whether intervention happens on paste, during composition or at submission; these are different points in the workflow.
Inspect the final action
Attempt the normal send action and record what the user sees. Where permitted in the test environment, inspect the resulting conversation and available product evidence. A popup alone does not establish that the provider never received the submitted material.
Repeat with a benign control
Send the allowed version using the same methods. Record any unnecessary interruption separately from a missed restriction. Retest inconsistent results and identify their conditions before deciding whether a method is supported, unsupported or still unresolved for this deployment.
What to check before proceeding
1. Input-method coverage
- Ready when
- Typed and pasted fixtures meet the documented action for each in-scope method.
- If the check fails
- Exclude the failing method from the coverage statement and agree a workaround or remediation test.
2. Submission outcome
- Ready when
- The evidence distinguishes prevented, modified, overridden and completed submissions.
- If the check fails
- Mark the result inconclusive and obtain a supported way to verify the final action.
3. Benign usability
- Ready when
- The approved control remains usable without an unexplained policy interruption.
- If the check fails
- Review the matching rule before broadening enforcement to more users.
Common mistakes to avoid
- Counting a warning as prevention when the user can still send the original text.
- Changing AI service, account type and input method together, leaving the cause of different results unclear.
Evaluate this workflow with Aona
Where Aona can help
Ask Aona sales and engineering to identify the supported prompt workflows and suitable synthetic fixtures for a scoped pilot.
What to confirm
Confirm typed and pasted behavior for the actual release and provider; this protocol does not claim every input path is covered.
Evaluating ChatGPT prompt controls?
Discuss typed and pasted prompts on the ChatGPT surface your employees use. Review supported policy responses with synthetic inputs and agree on acceptance criteria.