There are two broad ways to put AI into test automation. In the first, a model helps people write tests: it drafts steps from a document, suggests wording, proposes a fix when the app changes. The output is a test that a person reviews and saves. In the second, a model reads each step while the test is running and decides, on the spot, what to click and what to type.
The second approach makes an impressive demo. It also gives up a property that testing depends on. We design for the first, and this post explains the architecture that makes it work.
What determinism means here
A test run is deterministic when the same test, against the same application state and the same data, takes the same actions and reaches the same verdict every time. The application may be slow one day and fast the next. The test's behavior should not change with it.
That property is what lets a failure be reproduced, a run be compared with last quarter's run, and an auditor be shown exactly what was checked. When a model interprets steps at run time, two runs of one test can do different things, and a failure may reflect the model's reading that day rather than the application.
Authoring time versus run time
| Concern | AI at authoring time | AI at run time |
|---|---|---|
| Same actions every run | Yes, from a compiled plan | Not guaranteed |
| Reproducing a failure | Re-run the same version | May not recur |
| Review before use | Every change reviewed | Decisions made live |
| Cost per run | No model calls | Model calls on every step |
| Suite with AI switched off | Keeps running | Stops |
The pipeline: author, compile, run, propose
A deterministic design has four stages. AI participates in the first and the last. The middle two contain no model at all.
Author
People write tests as plain-English steps in a governed grammar: a fixed set of verbs, quoted targets that name what a person sees on the screen, explicit values and optional qualifiers. AI can draft these steps from a process document, a recording or a requirement, but what it produces is a draft. A person reviews it and saves it.
Open the 'Invoice' ''
Check that 'Hold reason' shows 'Price variance'
Click the 'Release hold' button
Check that 'Status' shows 'Validated' within 30 seconds
Compile
Each saved step is parsed into a structure: verb, target, value, qualifier. Targets are bound to objects in a model of the application, so 'Hold reason' refers to one known field on one known page. A step that does not parse, or names a target the model does not know, fails here, at compile time, where it is cheap to fix. It never reaches a run as an ambiguous instruction.
The compiled result is versioned. Any run can be traced back to the exact version it executed, and any two versions can be compared line by line.
Run
The engine executes the compiled plan. Waits are explicit: a check retries until its condition holds or its timeout passes, which handles a slow application without changing what the test does. Values come from named variables and data bindings, not from inference. Every step records what it expected, what it saw and the evidence behind both.
Propose
When a step fails because the application changed, AI comes back in. It can read the evidence and propose a fix: the 'Release hold' button is now labeled 'Release'. The proposal goes back to authoring, where a person who is not the author approves it, and the test is recompiled as a new version. The run itself never changed mid-flight.
Handling a changing application without runtime AI
The common objection is that applications change and only a model can keep up live. In practice, most of what breaks tests can be handled deterministically:
- Slow pages and background processing, through retry-to-timeout on each check
- Moved controls, through targets bound to a model object rather than a position or selector
- Changing data, through variables and data bindings resolved before the step runs
- Real changes in the application, through proposed heals that are reviewed, then recompiled
The last case is the one where a person should be involved anyway. A changed application may be an intended change or a defect, and that is a judgment, not a lookup.
What this design buys you
- Failures that reproduce, because the same version runs the same way
- A version history that shows exactly when and why a test changed
- Run costs that do not grow with the number of steps or runs
- A suite that keeps running if AI is switched off for a project or a provider is unavailable
- Evidence an auditor can follow from requirement to result
How Qventis One helps
Qventis One is built on this design. Tests are written in Studio as governed plain English, with AI drafting and proposing heals for people to approve. The Qventis Engine compiles each test once against the App Model and runs it the same way every time across eight engines, with zero AI tokens at run time and evidence for every step. Book a demo to see a test go from draft to compiled plan to run.