Field Notes

Code Velocity Went Up. Code Confidence Went Down

One senior engineer’s experience with AI-assisted development and what it reveals about the growing challenge of validating software.

We recently sat down with a senior software engineer at a financial technology company. His team uses AI coding assistants every day. He is a developer, not a tester, and he asked that we not identify him or his employer.

We wanted to understand how AI had changed the way his team develops software. What we heard was a candid account of what happens when writing code becomes much faster, but testing it doesn’t.

His central observation was simple:

“I’m just not able to test as fast as I can write.”

The problem wasn’t that AI couldn’t write tests. It could. The problem was that his team still had to prepare environments, understand dependencies, review generated tests, and determine whether changes affected the rest of the product.

AI had accelerated one part of engineering without accelerating all the work required to release software confidently.

This is one engineer’s experience, not a study of the industry. We’ve kept his observations separate from our own conclusions.

What He Sees in Code Review

Tests that no one reads. AI now writes many of his team’s unit tests. He told us that he may still review tests on his own pull requests. But when reviewing someone else’s work, he often doesn’t.

“When I’m reviewing somebody else’s PR, I don’t look at their tests anymore. I assume that they have tested it.”

Test creation has become easier. Verifying whether those tests check the right things hasn’t.

A test suite can grow while fewer people examine what the tests actually assert. He said his review now focuses more on behavior and context: does this change make sense in the application?

Bigger changes. AI makes adjacent changes easy. A developer asks for a feature, notices something else that could be improved, and asks the assistant to fix that too. A cleanup here. Another improvement there.

“Every PR is bigger,” he said.

The reviewer now has more code to understand but no more time to review it.

Code that looks right but isn’t. This was one of his biggest concerns.

“It looks right, it makes sense in your head, but is actually a wrong business case. Then you ship it.”

In our view, this is where generated tests offer the least protection. When AI writes both the implementation and its tests from the same incomplete understanding, the same mistake can appear in both. The tests may pass while the business requirement is still wrong.

That is a different problem from catching syntax errors or broken functions. It requires someone who understands what the software is supposed to accomplish, not just how the code works.

New engineers trusting the model. He described how he reviews changes by understanding the complete flow, not just the lines that changed. An engineer who has been working in the codebase for a month may not yet have that context. Instead, they rely on the assistant.

“They’re just trusting [the model] because, well, my company trusts [it],” he said.

That creates additional work for the engineers who know the system well enough to recognize when something doesn’t make sense.

The Environment Tax

When we asked what slows his testing down, he didn’t start with writing tests. He talked about getting somewhere to run them.

His team has two main preproduction options, and neither is particularly convenient.

Developer environments take approximately 30 minutes to provision. Engineers then have to seed data, configure services, and prepare the specific conditions needed for their feature. These environments also rely on mocks, so some behavior differs from production.

Staging is closer to production infrastructure, but the data and services are frequently out of sync. Testing there takes time, and not every engineer on his team does it consistently.

He gave us an example.

An engineer develops a feature that requires a new event consumer. They create the consumer in their developer environment. Everything works. But they forget to include the script that creates the queue the consumer reads from in production.

The application code is deployed. The queue doesn’t exist, so the consumer has nothing to read. The feature fails.

Staging might have caught the missing dependency, but testing there is sometimes skipped because of the time involved.

The code worked in the environment where it was tested. The problem was that the environment didn’t match what would actually be deployed.

This is something engineering leaders should pay attention to. In our experience, if setting up a realistic environment takes too much time, developers will find ways around it. And when that happens, environment differences become a source of defects that ordinary unit tests may never catch.

When Production Becomes the Testing Environment

His team also validates features in production using feature flags.

There is nothing inherently wrong with that. Progressive rollouts, feature flags, and production monitoring are legitimate engineering practices. They can reveal problems that preproduction testing cannot.

But his concern was different. When development and staging environments aren’t reliable, production becomes the first place a feature is meaningfully validated.

He described what happens when he cannot get sufficient confidence before deployment:

“I just pray and say I’m gonna put this feature flag on, and I’m gonna enable it for just myself.”

A feature flag can limit who sees a change and allow a team to disable a feature quickly. But it doesn’t guarantee that newly deployed code cannot affect existing behavior.

He said that, on his team, he sometimes doesn’t feel confident that a feature is working correctly until it has been stable in production for approximately two weeks.

That is a long time to wait for confidence in something that has already shipped.

The issue isn’t testing in production. It’s relying on production because the earlier validation process hasn’t provided enough confidence.

Who Owns the Testing Nobody Gets Around To?

His company doesn’t have a dedicated QA function responsible for testing individual features. Developers test their own work.

There are specialized teams supporting infrastructure, capacity, load testing, and testing tools. But feature-level validation remains largely with the developers.

We put the strongest argument for this model to him. If developers are responsible for quality, they cannot simply throw code over the wall to a QA team. They have to care about whether their work is correct.

He agreed with the principle. His concern was whether it works consistently under delivery pressure.

“I’m always under pressure to deliver a feature. I’m never gonna invest in making testing better. Everything is a fire. Everything is important, and that’s how you make mistakes.”

He acknowledged that some engineers are very good at testing their own work. But improving testing infrastructure, building end-to-end coverage, and maintaining reusable scenarios compete with the immediate pressure to deliver features.

The feature usually wins.

His team has unit tests and some integration tests, but he described those integration tests as limited. They have no end-to-end tests for the workflows his team owns.

He also described a blind spot that comes from testing your own changes.

“I’ve seen that my specific scenario works, but I haven’t tested: did other things break?”

That question gets at something important. A developer naturally concentrates on the change being implemented. Someone responsible for broader validation needs to understand how that change interacts with the rest of the application, including features delivered weeks or months earlier.

He called building and verifying your own work a “conflict of interest.”

His preferred solution wasn’t necessarily a separate department. He suggested engineers could rotate between development and testing roles, spending time on both sides. But he was clear that testing is a specialization.

The organizational structure can vary. What matters is that the work has an owner, sufficient expertise, and time to be done properly.

What Falls Between Teams

His organization has specialized engineering capabilities. There are teams responsible for capacity planning, synthetic load testing, testing infrastructure, and tools that identify inefficient database queries.

Yet he described several questions that aren’t consistently addressed when his team introduces a new endpoint.

  • Does it need rate limiting?
  • Should it use caching?
  • What response time should trigger an alert?
  • What error rate is acceptable?
  • How should it behave under load?

These aren’t obscure testing questions. They are basic considerations that affect performance, reliability, security, and operational cost.

The problem is that the specialized teams operate at the platform level, while developers concentrate on individual features. Some questions fall between them.

In our view, capacity is a good example. Increasing it ahead of a major traffic event and reducing it afterward can be a reasonable operational decision. But without sufficient performance data, an organization may not know whether the additional capacity is necessary or whether the application could handle the workload more efficiently.

AI can generate a new endpoint quickly. It doesn’t automatically establish the endpoint’s operational requirements or ensure someone has validated them.

The faster software changes, the more important it becomes to know who owns these questions.

Where AI-Assisted Testing Helps, and Where It Doesn’t

We asked whether AI makes dedicated testing expertise unnecessary.

His answer was that testing is “way more important” now.

He wasn’t arguing against AI. He uses it throughout his development process. His concern was that AI can generate code faster than developers can establish confidence in what it produces.

He also had a specific idea about how AI should be used in testing. Let AI explore an application and work out a user flow. Have it generate a test script. Then have an engineer verify the script and incorporate it into the automated test suite. From that point forward, run the verified script rather than asking the model to rediscover the same flow every time.

As he put it:

“I don’t want this LLM to run and generate this flow off its head every single time.”

We agree with the principle, particularly for known regression scenarios.

There is a legitimate place for adaptive, agent-driven testing. It can be valuable for exploration, unfamiliar workflows, and situations where the path changes frequently. But when the expected behavior is known, repeatable automation is often the better choice.

A generated test should establish three things:

  1. It checks the intended business behavior, not simply the current implementation.
  2. It would detect the relevant failure if the code were wrong.
  3. It can run reliably and be reviewed and maintained by engineers.

This is also how we approach AI-assisted testing through QASIP, our internal quality engineering platform. QASIP helps generate test cases and automation using application context and existing engineering assets. Engineers verify the generated work before it becomes part of the client’s testing process. More on how we use AI for QA →

The goal isn’t to introduce AI into every execution step. It’s to make useful test coverage faster to create and easier to maintain without giving up control over what gets tested.

Redesigning Validation, Not Adding Headcount

The lesson from this conversation isn’t that every company needs to hire more testers. It’s that validation needs to keep pace with the way software is now being developed.

We would start with five changes. The first and third came directly from him; the rest are our conclusions.

  • Make environments easier to use. Developers need production-like environments that can be provisioned quickly, with realistic dependencies and appropriate test data. If testing requires too much preparation, some of it will be skipped.
  • Verify AI-generated tests. Generating more tests isn’t enough. Someone must establish whether those tests check the right behavior and would catch meaningful defects.
  • Use repeatable regression automation. AI can help discover and create test scenarios. Once the expected behavior is understood, verified automated tests can provide consistent regression coverage.
  • Give validation clear ownership. Whether that responsibility sits with an embedded quality engineer, a dedicated team, or a rotating engineering role, someone must look beyond the immediate feature.
  • Use production validation deliberately. Feature flags, progressive rollouts, and monitoring should add confidence, not compensate for an inadequate preproduction process.

He also described the working relationship he would want between developers and testers.

When a feature is ready, the person responsible for validation gives the developer a short set of checks to complete first. That prevents obvious problems from consuming validation time. Then the tester examines the broader risks: regressions, unexpected user behavior, dependencies, and scenarios the developer may not have considered.

The developer remains responsible for the release decision. The validation function makes sure that decision is informed by what has been tested, what hasn’t, and what risks remain.

That is a more useful model than treating QA as a gate at the end of development.

Five Questions Engineering Leaders Should Ask

If your development team is using AI coding assistants, these are five questions worth answering.

  1. How much time does your team spend preparing environments before testing can begin? Include provisioning, seeding data, configuring dependencies, and resolving differences between environments.
  2. Who verifies that AI-generated tests check the correct business behavior? A passing test isn’t useful if it confirms the wrong requirement.
  3. Are deployment dependencies tested as part of the release process? Consider consumers, queues, configuration, database migrations, and other infrastructure changes that may exist in development but not in production.
  4. Who owns non-functional quality for individual features? Rate limits, caching, performance objectives, acceptable error rates, load behavior, and security controls all need clear ownership.
  5. How long after release does your team actually trust a feature? If confidence comes only after days or weeks in production, what was missing from the validation process before release?

These questions won’t produce a complete quality strategy. But the answers will show where development speed may have moved ahead of validation.

Release Confidence Is a Business Question

The engineer was candid about the limits of his experience.

“That’s my opinion. I don’t have any metrics.”

His observations describe one team. Other teams in the same company may work differently. And we came into the conversation with our own perspective. QASource is a quality engineering company. We believe independent testing expertise has value.

But what made the conversation useful wasn’t a general argument for QA. It was the specificity of the problems he described.

  • A reviewer who no longer reads all the tests accompanying other engineers’ changes.
  • A feature that works in development but fails because a production dependency was never created.
  • A staging environment that is difficult enough to use that testing sometimes gets skipped.
  • New engineers trusting AI-generated code without fully understanding the system around it.

These are engineering problems with potential business consequences. A change that implements the wrong business rule can affect customers and revenue. An endpoint without appropriate operational controls can become a reliability problem. A feature that isn’t trusted until it has been in production for two weeks creates uncertainty about what the business can safely depend on.

Toward the end of our conversation, the engineer said something that stayed with us:

“I am more pro-human now than ever.”

He wasn’t rejecting AI. He was making the opposite point. AI has made engineers faster, but that makes experience and judgment more important, not less.

The point isn’t for AI to slow down. It’s for the rest of engineering to catch up.

Writing code has become faster. Understanding what it does, what it might break, and whether customers can depend on it still takes work.

The next improvement in software delivery may come less from writing code faster and more from making it easier to trust what we ship.

← All insights

Shipping Faster Than You Can Verify?

Tell us where release confidence is slipping. We’ll look at how your validation process is designed, where the risks are concentrated, and what needs to change.