Is Your Agentic Test Framework Worth the Complexity?

The useful question isn’t whether your framework is agentic. It’s whether the autonomy is paying for itself.

In the frameworks we are asked to review, we increasingly see AI roles, configuration files and orchestration layers wrapped around work that was already predictable. The architecture works. The harder question is what all that machinery is buying.

Some teams built an agentic QA framework once coding assistants became normal. Others inherited one from a vendor or another team and are now being asked to extend it.

The first argument is usually about whether the thing is really agentic. Skip it. There isn’t a settled definition, and the answer doesn’t change what your team does on Monday. What matters is whether the architecture is worth what it costs to keep.

Autonomy Is a Cost Until It Produces a Benefit

We aren’t arguing against agentic testing. AI is a large part of how we deliver.

The moment you add agents, you add complexity. Runs become less predictable. Failures can be harder to reproduce. There is more to maintain and more to figure out when something goes wrong. That can absolutely be worth it. If the agent finds paths your scripts would miss, adapts when the application changes, or gets past failures that would normally stop a run, it is doing useful work.

But if you end up with the same coverage you could have had with a straightforward automated test, you haven’t made testing smarter. You’ve made it more complicated.

Five Things We Check and What They Cost

These are the five capabilities we look for. They aren’t an industry standard; there isn’t a settled one. We look at what the framework demonstrates against your application today and at what it will cost to maintain as the product moves. Design intent matters less than either.

You can run these yourself. Each check has a cost test underneath it, because a capability that works but costs more than it saves is still a problem.

Can It Reason?

We give the framework a goal and an application that has changed. Can it work out what to do, or is it running a pre-written sequence?

Cost Check

Build the same test the ordinary way and compare engineer hours. If the framework took longer and covered no more, the reasoning isn’t paying its way.

Can It Adapt?

We change the interface, the workflow or the environment. If an engineer has to edit selectors, prompts or configuration before the test works again, the framework isn’t adapting much.

Cost Check

Take a recent product change and count the files your team had to touch. If every change sends engineers into role definitions and configuration, upkeep grows with the product.

Can It Recover?

We cause a failure mid-run. Can the system diagnose what happened and try another route, or can it only report the failure? At what point does an engineer have to step in?

Cost Check

Time how long it takes to get from a red run to a known cause. If your engineers spend that time guessing, or can’t see which route the AI took, trust in the suite fades and people start skipping tests.

Can It Generate?

We look at what the system leaves behind. Is it finding new scenarios or producing reusable components that didn’t exist before, or is it mainly executing what the team already gave it?

Cost Check

Count how many of the scenarios it produced a person later kept. If most get thrown out, the review time is the cost.

Can You See Inside It?

Can you see why the system made a decision, and trace the prompts, tool calls, model versions and outcomes behind it? Can a human review or approve the actions that matter? And when the model changes, can you tell whether its behavior changed too? Where test evidence has to be repeatable and auditable, this becomes a governance issue and, in regulated environments, potentially a compliance one.

Cost Check

Sit through a code review of a change the framework made. If an engineer can’t see what changed, why, and who approved it, you are carrying a governance cost nobody has priced.

Two Checks for the Whole Framework

Count the maintained roles, tools and configuration files behind a single workflow, and ask what each one contributes. Then run the framework against a newer model and see whether the results change. If nobody can say, that uncertainty is part of the cost.

If the framework shows little or none of the five, stop judging it as autonomous testing and compare it, on cost, with orchestration. There is nothing wrong with orchestration. Plenty of teams need it. But then the comparison is with a scheduler and a test runner, and those are much cheaper to own.

Where Autonomy Earns Its Keep

Every testing problem can be scripted. Some of them only at a price that keeps climbing. That’s where an agentic framework starts to justify the added complexity. For example:

  • The Steps Depend on What the Application Shows: Take a loan application that branches on the applicant’s answers, with different screens and follow-up questions each time. You could script every branch in advance, but the number of paths can make that expensive to build and maintain. An agent can decide what to do next based on what appears on the screen.
  • The Interface Changes Often: A workflow may move from one screen to three, or present different controls depending on what happened earlier. A fixed test has to be updated to reflect the new path. An agent that understands the goal can work through the changed flow without having that exact sequence written down first. That’s worth paying for when maintaining those paths costs more than the autonomy does.
  • A Failure Is Worth Diagnosing: Say a test fails halfway because a promo code has expired. A script stops and reports red. An AI can work out the cause, try another route, finish the run and tell you what it found. That matters when a stopped pipeline costs your team a day.
  • The Application Holds Scenarios Nobody Wrote Down: While moving through a real application, the AI may reach a combination nobody thought to test, such as a discount applied to a gift card, and test it. That’s coverage you didn’t have to think of in advance.
  • Several Routes Are All Correct: If there are five legitimate ways for a user to complete checkout, a test that knows only one route answers a narrower question than whether checkout works.

In these cases the AI is doing what a fixed script can’t do at a reasonable price: choosing its own steps, surviving change and finding new scenarios. That’s what the extra architecture buys.

If none of these describes the work, ask whether an agentic layer is buying anything that deterministic automation or AI assistance couldn’t deliver more cheaply.

Prove It With a Scorecard

Treat an agentic layer like any other engineering investment. If it is producing real value, these numbers will show it.

See where your framework stands and share the results with your team.

GET YOUR COMPLETE SCORECARD
Measure What To Compare Warning Sign
Test creation effort Agentic vs. conventional, same use case Agentic takes longer with no extra coverage
Maintenance effort Engineering hours per product change Hours rise with every release
Execution time Comparable workflows and coverage Slower runs for the same coverage
Coverage Scenarios executed or discovered Nothing was found that the team didn’t write
Recovery rate Failures diagnosed and recovered automatically Every failure still needs an engineer
Reliability False failures and flaky runs Failures stop being taken seriously
Model and API cost Cost per workflow or run Spend grows faster than coverage
Release velocity Product change to green pipeline Testing becomes the bottleneck
Architecture overhead Components the team must maintain Nobody can say what a component contributes

Most teams can fill this in within a few weeks. If the agentic column wins on effort, coverage or recovery, you have your answer and should keep going. If it loses on most rows, you are paying agentic prices for orchestration. And if you cannot assemble these numbers at all, that is a finding in itself.

The Three-Lane Rule

Our default is straightforward. Every piece of testing work belongs in one of three lanes:

Lane 1

Deterministic Automation

Wherever you already know what should happen. Most of a regression suite lives here.

Lane 2

AI Assistance

Wherever it reduces human effort—writing tests, migrating code, creating data, triaging failures.

Lane 3

Autonomy

Only where reasoning, exploration, adaptation or recovery buys something you can point to.

Most quality systems need all three. The expensive mistake is running everything in lane 3.

That standard applies to us. QASIP, our own platform, sits in lane 2. It generates test design, automation and migrations, and a QASource engineer verifies what it produces before anything reaches your repository. It is not an agent, and we do not sell it as one.

The frameworks we’re asked to look at usually aren’t in trouble because somebody used AI. They’re in trouble because autonomy was applied to lane-1 work, where it added cost without removing work.

When the Agentic Layer Isn’t Worth It, Simplify It

What do we do when an assessment comes back that way?

  1. Rebuild on a Design, Not a Pile of Roles:

    For web automation that means Playwright, built in layers your team can maintain: tests written as code, reusable page and component pieces, clear reporting, CI/CD integration and clear ownership. We build from how the system should work now, not by copying the shape of the old framework into a new tool.

  2. Generate the Missing Coverage:

    Manual cases that were never automated become scripts, and a QASource engineer reviews each one before it goes into your repository.

  3. Migrate What Works:

    Working tests get migrated, not casually rewritten.

  4. Take the Machinery Out of Test Data:

    Creating data through the user interface is usually slower and more fragile than creating it through APIs. If the endpoints don’t exist, asking the product team for them can cost less than maintaining years of setup logic.

  5. Keep What’s Useful:

    A functional engineer should be able to describe what they need in plain language and get useful work back. But every agent, role definition and layer your team maintains has to earn its place. The ones that don’t are inventory.

What you end up with is a suite your engineers can read, a maintenance load that grows with your coverage rather than with your product, and a release decision somebody is willing to put their name to. That is the point of the exercise. The simplification is only how you get there.

We Ran It on Ourselves First

We ran Selenium for 20 years. On our own estate, Playwright turned out to be materially better: quicker to execute, cheaper to maintain, and with fewer false positives.

That left us with a problem. A large part of our client estate had been built over those two decades. Knowing the newer framework was better didn’t make migrating all that existing work affordable.

We used QASIP, our in-house AI platform for test generation and migration, to do that work. We started on the codebase we could most afford to break while finding out whether the approach held: our own.

Execution time
50–60% less
Maintenance effort
40–50% less
Migration effort
70–80% less

Reduction against the Selenium baseline

Execution and maintenance are Playwright against Selenium on comparable suites. Migration effort is QASIP against a hand migration.

If your agentic framework was built on Playwright from the start, the migration is not your problem and this section is not your argument. The checks above still are.

It was the least forgiving migration available to us. That experience is also why we’re careful to talk about what a framework transition costs, not only what it gains.

Put Your Own Framework Through the Five Checks

A scoped assessment puts our engineers in your codebase and gives you a written finding:

  • Which of the five capabilities your framework actually demonstrates
  • How it scores against the scorecard above
  • What it costs you to own
  • What we would keep, simplify, replace or build

Sometimes the answer is to keep it. If the answer is rebuild, we can do that too—this is an engineering assessment of an architecture, not an independent assurance opinion, and we are clear about which one you are buying.

The finding is yours either way and commits you to nothing.

We judge an AI layer the way we judge any other architecture. Does it remove work? Does it give you coverage you trust? Does it make change cheaper? Can you measure what came back? If it does, keep it. If it has become one more system your engineers maintain and nobody can point to what came back, simplify it.

Autonomy Is a Cost. Spend It Where It Pays.

Tell us what your framework costs you to own. We’ll come back with a proposal within two business days of the first conversation.