qventis blog

Coding agents and tests under policy | qventis blog

Written by qventis team | Oct 2, 2026, 3:00:00 PM

Coding agents have changed how a lot of development work finishes. An agent makes a change, runs the tests, reads what failed, adjusts and runs them again until it is satisfied. For a small repository with fast unit tests, this loop is mostly a productivity gain.

Enterprise testing is a different setting. Functional and end-to-end suites run against shared environments, use reserved test data, may touch regulated records and can take hours. The results feed release decisions that someone has to sign. An agent that runs these suites with a developer's broad credentials, as often as it likes, against whichever environment it finds, creates problems no one intended.

What goes wrong without policy

  • Runaway runs. The agent reruns a long suite after every small edit, tying up environments and data other teams need
  • The wrong target. It points tests at an acceptance or production-like environment because the configuration allowed it
  • Edits to make tests pass. It changes an expected result or skips a check, because that is the shortest path to green
  • Borrowed identity. It acts with the developer's permissions, so its actions cannot be told apart from theirs
  • No record. Nobody can reconstruct later what the agent ran, changed or concluded

None of this requires a misbehaving agent. It only requires an agent doing its job with no boundaries around it.

Give the agent its own identity

Agents should sign in as themselves, through the same identity provider as people, using SSO and OAuth. Each session should be linked to the person who started it, but carry its own permissions. That way the audit trail shows clearly what the agent did and on whose behalf, and an agent's rights can be narrower than its user's.

Define what it may do, project by project

The core of the policy is an allow-list of actions, set per project, with conditions. A starting point might look like this:

ActionAgentCondition
Read tests and resultsAllowedProjects it is assigned to
Ask which tests a change affectsAllowedAny assigned project
Run a selectionAllowedDevelopment and test environments only
Run the full suiteNot allowedPipeline only
Propose a healAllowedGoes to a person for approval
Approve any changeNot allowedNever, for any agent
Edit an expected resultNot allowedTest owner only
Change data poolsNot allowedTest data owner only

Each call is checked against the agent's role and attributes in the same way a person's request would be. The policy decides, not the agent's judgment about what seems reasonable.

Ask for a selection, not the suite

The most useful habit to build into an agent's workflow is to ask what to run before running anything. Given a change, the platform returns the tests that touch the affected screens, steps and services, ranked by risk. The agent runs that selection.

This keeps the loop fast enough for an agent to iterate, keeps shared environments available for others, and makes the agent's evidence relevant. A selection tied to the change also makes clear what the agent did not run, which matters when its result is presented as a signal.

Return evidence, not just pass or fail

An agent can only reason about a failure if it can see what happened. Results returned to an agent should include, for each failed step, what it expected, what it saw, the step log and a screenshot or response body. With that, the agent can usually tell whether its change broke something, the test needs a heal, or the environment had a problem, and it can propose the right next action.

The run itself should stay deterministic and use no AI. The agent spends its own reasoning deciding what to call and how to read the results. The test execution behaves the same way whether a person, a pipeline or an agent started it.

Keep the gate where it was

An agent's local green is information, not approval. The change still passes through the same pipeline gate as any other: the pipeline runs its own selection under its own policy, and the release decision still belongs to a named person. Heals the agent proposed wait for a reviewer who is not the agent and not the person who launched it.

This parity is the point. People, pipelines and agents reach the platform through different doors, a CLI, an HTTP API or MCP, but the same rules apply behind each one.

Audit every call

Every agent action should be recorded with who asked, which inputs, what result and under which policy. Export it to the same SIEM that holds the rest of your security events. When someone asks why a test changed or why a run happened at two in the morning, the answer should be one query away.

Roll it out in stages

  • Start read-only, letting agents read tests, results and evidence
  • Allow selections in development environments for one project
  • Add heal proposals once the review process is in place
  • Widen scope project by project, based on the record so far

How Qventis One helps

Coding agents such as Claude Code, GitHub Copilot and Cursor connect to Qventis One through MCP, and pipelines through the CLI or HTTP API, all under the same policy. Agents sign in with SSO and OAuth, per-project allow-lists decide what they may do and where, and they can ask which tests a change affects, run that selection and read the evidence. Runs use zero AI tokens, heals wait for a person, an agent can never approve its own change, and every call is exported to your SIEM. Book a demo to see an agent run tests under your policy.