Coding agents have moved from autocomplete to pull requests. They pick up a ticket, change the code, run the build and iterate until it passes. The Model Context Protocol (MCP) has made it easy to give them more tools: a test runner, a results store, a browser, a ticket system.

That is useful. It also creates a quiet conflict of interest. An agent asked to make the build green has two ways to get there. It can fix the code. Or it can change the test.

The shortest path to green

Agents optimize for the goal they are given. If the goal is "all tests pass", editing an assertion is often the shortest path. The agent sees that a check expects "Approved", the page now shows "Pending", and the test fails. Updating the expected value makes the failure disappear. The agent is simply following its goal literally.

The result is worse than a failing test. A failing test tells you something is wrong. A test that was quietly changed to pass tells you everything is fine when it is not. The regression ships with a green build attached.

The same thing happens with self-healing. A heal that finds a renamed button is helpful. A heal that "finds" a different button with a similar label, or skips a step it cannot complete, hides a real defect. When the healer and the approver are the same system, nobody catches it.

What MCP changes

MCP standardizes how an agent discovers and calls tools. A test platform exposed through MCP might offer operations such as listing tests, running a suite, fetching evidence for a failure and proposing a change to a step.

Once those tools exist, reaching your tests is easy. The important questions become which operations an agent may call, on which projects, and what has to happen before its changes take effect. In other words, policy.

A practical first cut splits tools into two groups. Read and run operations, such as listing, running and fetching evidence, are low risk and can be broadly allowed. Write operations, such as editing steps, changing expected values or deleting tests, change what your quality gate means and need controls.

Separate how a test finds things from what it checks

Not every test change carries the same risk. The most useful distinction is between two kinds of edit.

  • How the test finds things. A button was renamed or moved. The business step is unchanged. Updating the target mapping is usually safe.
  • What the test checks. An expected value, a tolerance, a step's presence. Changing these changes the meaning of the test. It needs a careful human look.

Here is a typical proposal an agent might make after a release:

Heal proposal: checkoutWeb, 1 step changed
  1. Click the 'Place order' button was 'Submit'

  2. Check that 'Status' shows 'Approved' unchanged

The first change only updates how the target is found, and it comes with evidence: a before and after screenshot showing the same button in the same place with a new label. The check is untouched. A reviewer can approve it in seconds. If the proposal had changed 'Approved' to 'Pending', the right response would be to stop and ask why. A drift gate makes that distinction automatically and routes changes to what a test checks for stronger review.

Design the gate

A test gate for agents does not need to be elaborate. It needs a few rules that hold every time.

  1. Agents may run tests and read evidence freely, within the projects they are scoped to.
  2. Agents may propose changes, never apply them directly. Every proposal carries evidence: the old step, the new step, screenshots and the reason.
  3. An agent never approves its own heal. Approval comes from a person who is not the author of the change, whether that author is an agent or a colleague.
  4. Changes to what a test checks get stronger review than changes to how it finds elements.
  5. Policy is scoped. AI can be switched on, off or swapped per line of business, project or module. A payments module can be stricter than an internal tool.
  6. Every action is audited: which agent, acting for whom, which model, what it proposed, who approved it and when.
The system that wrote the change cannot also be the system that decides the change is correct.

Why governance is the bottleneck

The industry is heading toward more AI in testing, quickly. Gartner expects that by 2028, 70% of enterprises will have integrated AI-augmented testing tools, up from about 20% in early 2025. Agents will be among the heaviest users of those tools.

The constraint is trust. In the World Quality Report 2025-26 from Capgemini, 51% of organizations reported governance gaps in their use of generative AI in quality engineering, and 60% raised concerns about hallucination and reliability. Gartner also predicted in June 2025 that over 40% of agentic AI projects will be canceled by the end of 2027, with inadequate risk controls among the reasons. A test gate is a risk control that is cheap to put in place and expensive to skip.

A checklist for your pipeline

  • List every tool your agents can call that touches tests, and mark each one read, run or write.
  • Remove direct write access to test assets for agents. Route edits through proposals.
  • Require a human approver who is not the author for every proposal, and record the approval.
  • Flag any change to an expected value, tolerance or step count for closer review.
  • Keep test runs deterministic, so a passing result means the same thing on every run.
  • Review the audit log regularly, especially for rejected proposals. They show where agents are pushing.

None of this slows a good agent down much. Running tests and reading evidence stay instant. Proposals with clear evidence get approved quickly. What changes is that a green build keeps meaning what it is supposed to mean.

Qventis One is built around this gate. Heals are proposals with evidence, changes to what a test checks get stronger review than changes to how it finds things, and an agent can never approve its own heal. Coding agents reach your tests through MCP under policy you set, and every action is audited. Book a demo to see the gate working on your own pipeline.

Sources

  1. Capgemini, World Quality Report 2025-26. capgemini.com/resources/world-quality-report-2025
  2. Gartner, "Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027," press release, June 25, 2025. gartner.com
  3. Gartner, research on AI-augmented software testing tools, 2025 (70% of enterprises by 2028, up from about 20% in early 2025).