Writing tests in plain English is an old idea that keeps getting better tools. The appeal is obvious. A business analyst can read the test. A product owner can check that it matches the requirement. A new SDET can understand the suite on day one.
Look inside a natural-language suite a year after it started, though, and you often find something like this:
Each line is readable. None of them is exact. Which orders page? Which order is "the new one"? What does "worked" mean, and what counts as "right"? A person can guess. A machine has to guess too, and that guess is where flaky tests and runtime AI costs come from.
Free-form steps drift for ordinary reasons. Every author has a habit. One writes "click", another "press", a third "tap" and a fourth "select". Some describe the element by its label, some by its position, some by what it does. Over time a suite collects five ways to say the same thing and no way to tell whether two steps mean the same thing.
The tool then has to interpret all of that variety. Some tools use pattern matching, which breaks on the sixth phrasing. Others send each step to a language model at run time, which handles variety well but brings cost, latency and results that can change between runs.
The people side suffers too. Reviews turn into debates about what a step means. Maintenance gets harder, because nobody can search for every step that touches the "Save" button when it is described ten different ways. In the World Quality Report 2025-26 from Capgemini, 50% of organizations named maintenance and flaky scripts as a challenge. Ambiguous steps feed both.
The fix is to give plain English a grammar. Every step should still read as a normal sentence. It should also be structured data underneath, with a small set of parts:
When a step is saved, it is parsed. If it fits the grammar, it becomes a step. If it does not, the author gets an error right away, the same way a developer gets a compiler error. Ambiguity is caught when the step is written instead of when the test fails at two in the morning.
Open the 'Orders' page
Click the 'New order' button
Choose 'Express' in the 'Delivery' list
Remember 'Order number' as ''
Check that the 'Orders' table shows '' within 5 seconds
Compare the last step with "make sure the new one is there". It names the table, the exact value and how long to wait. A reviewer can approve it without asking a question.
A grammar needs a vocabulary, and a small one works better than a large one. Twelve universal verbs cover most of what a business flow does:
Keeping the list short has three benefits. People learn it in an afternoon. Reviews get faster, because every step starts with a word everyone knows. And each verb can mean exactly the same thing on every channel, so "Click the 'Save' button" behaves the same way on a web page, a desktop client and a mobile screen.
Some channels need more. An API test sends requests. A data test compares tables. A mobile test swipes. The answer is to let channel packs add verbs only where the universal set cannot express the action:
Every extension follows the same pattern of verb, quoted target and value, so the suite still reads as one language.
The most common source of fragile tests is the way elements are identified. Many frameworks put CSS selectors or XPath straight into the test. Those strings describe the page structure, which changes often, and they mean nothing to a business reader.
A governed sentence identifies a target by the label a person sees, in quotes, followed by a role noun: "the 'Save' button", "the 'Email' field", "the 'Delivery' list". The quotes make the label exact. The role noun removes ambiguity when the same label appears twice, such as a "Save" button and a "Save" tab.
Behind the sentence, an object repository maps each target to however the element is actually found on that channel. When the UI changes, you update one entry and every test that uses it is fixed. The sentence never changes, because the business step never changed.
Values belong in quotes too, and data that flows between steps should be named. "Remember 'Order number' as ''" says exactly what is captured and what it will be called. A later step that checks "" is clearly linked to it. Free prose hides this flow in words like "it" and "the new one".
You can apply this thinking to any suite you have today. Export your natural-language steps and list the verbs in use. Collapse synonyms into one word each. Rewrite targets as quoted labels with role nouns. Replace "it" and "there" with named values. Then add a check that rejects new steps that do not follow the pattern. Within a few weeks, your suite will read more consistently and fail for clearer reasons.
Qventis One is built on governed plain English. Every step is a sentence backed by a schema, written in Studio, compiled once and run with zero AI tokens at run time on web, desktop and every other channel. Book a demo to see your own flows written as governed sentences.
Sources