Hero Image Acutest Banners

Should IT delivery teams take advantage of AI for requirements and testing?

22 July 2026 Time to read:  minutes

AI can help testers work through scattered information faster. But testing remains a social, judgement-based activity, and tools should support the enquiry and not define it.

Start with the requirement, not the tool

Programme leaders are under increasing pressure to adopt generative AI in IT delivery. Requirements and test design are often presented as obvious places to start.

Our experience is that AI tools have a place, but they bring limits and risks, and human skill and judgement remain essential.

Organisations should explore the technology for themselves rather than rely on vendor claims or industry hype to decide where it belongs.

When considering introducing AI tools in your workflows, your first thought might be to look for a new tool. In many ways, that’s exactly what happened when ChatGPT appeared. People gained access to the technology first, then spent time experimenting to see what the possibilities were.

You may be looking at AI because you want to address familiar delivery pressures:

  • Increase speed of testing to keep up with code generation?
  • Reduce variation between how tests are documented?
  • Improve test coverage and cut the number of bugs going undetected before production?

These are reasonable aims, but the way they are sometimes framed hide assumptions about testing. Speed is taken as the measure of value, even when sometimes the risks need careful investigation. Consistent formats can be mistaken for consistent quality of thought. Test coverage can sound objective, when it always depends on what someone has chosen to count. A test case can look like the work when it is often only a partial record of what people do when they test.

Knowledge, first-hand experience and ownership are all areas where teams can compromise without noticing. Using any automation tool creates distance between the person using the tool and the thing it is working on.

Using AI, generated outputs can look complete, professional and consistent yet be superficial. If you put yourself at arm’s length from the people and the systems, you can miss the contradictions and inconsistencies where the bugs that matter hide. Development is a social activity. People disagree, forget things, make trade-offs, and carry knowledge that never makes it into the document.

The danger is not the tools. The danger is taking something socially complex and trying to turn it into an automated routine.

The place to start is by asking your teams if they are already using AI for their work. A 2026 Google and Public First study suggests UK workplace AI adoption has risen sharply, from 34% in 2025 to 73%, but only 15% of workers are advanced users. There is a good chance people are already experimenting, but unevenly, informally and with very different levels of skill. Speaking to testers and business analysts who are already using these tools can show where existing enterprise AI helps, and where it falls short.

So, the first step is not procurement. It is finding out what people are already doing.

Test cases are not testing

The hard problem is not whether AI can generate requirements and test artefacts. It can. Most general enterprise AI tools can do this now. More importantly, these artefacts were never the destination. A test case is only a description of a possible test. Writing it will not find a bug. Understanding the business, challenging assumptions and learning the software might.

The harder problem is scattered knowledge

That work starts earlier: gathering and building knowledge of the business or product, finding information, and getting first-hand experience of the systems. Decisions, assumptions, conversations and knowledge are scattered across meetings, emails, documents and people’s heads. This information is spread unevenly across teams, gets lost in translation, causes delays, and creates bottlenecks around the few people who know where information is or what it means.

This is where we see the possibilities and value. Not to write test cases or requirements, but to navigate scattered information, find connections, expose inconsistencies and explore ideas.

One recent experience came from an enterprise client working through a large testing programme, changes to legacy systems and spotty documentation. Our approach was not revolutionary, but it worked: interview people across the programme, including system experts and users, record the meetings, keep the transcripts, and collate existing documentation, designs and specifications. We used the material to identify gaps and inconsistencies, build a shared understanding and focus testing on the areas of greatest risk.

Generative AI helped us work quickly with unstructured documents and sources: exploring them, reorganising them, analysing them, and checking our understanding. But the most important parts were still the interviews, reviews and conversations about risk and who might be affected, how badly, and what the programme was willing to accept. This is social grounding.

The transcripts and documents did not contain the understanding. They gave us material to question. We came to understand through the interviews, reviews and disagreements around them.

This is where the distinction matters. Testers can use generative AI for the parts of testing work that can be written down, repeated or reorganised: summarising documents, finding inconsistencies, drafting scenarios, producing checks and working through large volumes of information. But that is not the same as testing. Testing depends on judgement, context, risk, interpretation and shared understanding. Those are social activities. AI can imitate the language of that work, but it does not participate in it. It has no stake in the outcome, no standing in the group, and no responsibility when the judgement is wrong.

There are also practical failure modes testers need to be ready to spot. Beyond a certain volume of information, the model only works on part of what is available at a time. Generative AI can be confidently wrong, over-agreeable, biased, or sensitive to what was included, excluded or emphasised in the prompt and source material.

Tools should support the enquiry, not define it

This leads to tooling choices. There is a distinction between AI models and the tools they are embedded in. Tools and harnesses control your interaction with the model: how information is stored, retrieved, provided to the model, reduced, dropped from active memory, and shaped by background instructions. These choices are often invisible to testers. They reflect the tool developer’s view of the task, and shape what the tool treats as good, complete or useful work.

Every AI-enabled testing tool embeds a model of testing. Can testers see it, challenge it, and change it when it does not fit? The more opinionated the tool, the more of somebody else’s view of testing you inherit. That is a problem because the inherited view may be wrong, too rigid, or just not yours. It can quietly narrow judgement, reduce room for disagreement, and make it harder for testers to experiment, make their own mistakes and learn from them.

It becomes a problem when the tool controls the context, the coverage model, or the judgement about what is good enough. Tester-led use starts with testers framing the work: deciding what sources matter, what questions to ask, what risks to pursue, what coverage means, and what evidence is good enough.

Start with the testers, not the shopping list. What do testers need? Where does tooling remove drag? Where does it start taking over?

  • Existing enterprise AI can help testers pull together scattered information, compare sources, draft working artefacts and learn from the material.
  • Generalist tools can be configured with prompts, templates and review patterns so the team can repeat what works and avoid repeating what fails.
  • Specialist tools are easier to make the case for when the problem is workflow, traceability, audit, evidence capture or other procedural drag.
  • Bespoke capability may make sense where testers need tighter control over the work itself: procedural steps, standards, risk models, coverage measures, prompts, source handling or evidence.

These are not binary choices. Testers may use different tools for different parts of the work. The danger is when the tool stops supporting the enquiry and starts defining it. Using tools is fine. That implies intent. The problem starts when a tool is sold as a way to automate complex testing activity and delegate judgement to software.

Use tools to support the enquiry, not define it.

Tester Led Judgement

Figure: tool choices around tester-led judgement. The useful question is not which option is most advanced, but where testers keep control of context, coverage and judgement.

Read the framework above as a way of protecting testing judgement, not as a maturity model or procurement route. Use it to ask where speed is useful, where transparency is needed, and where packaged automation starts to hide the reasoning testers need to challenge.

Capability lives in the team

Capability is not a prompt library. Testers build capability by using the tools together, comparing results, sharing prompts, challenging weak outputs, spotting failure patterns and agreeing where the tool should stop. That means building shared practice: coverage models testers understand, review habits they actually use, examples of what works and fails, and clear choices about which models and tools are acceptable for which tasks.

This is not a one-off setup job. The skill is not only technical. Capability lives in the team as much as in the tooling.

Control the dependency

Programme leaders still have decisions to make, but they do not own the testing judgement. Their job is to foster the conditions in which testers can exercise that judgement well: enough time, access to the right material, safe tools, review discipline, clear accountability, and authority to challenge the tool’s output. Otherwise “human review” becomes another rubber stamp.

The right balance depends on context: time pressure, regulatory expectations, delivery scale and the consequences of getting the judgement wrong.

Because these factors vary significantly across programmes, there is no single “correct” approach. Anyone promoting a universal answer is selling snake oil.

All approaches introduce dependency, and all fail in similar ways if requirements are weak, judgement is abdicated and review discipline is absent.

Dependency is not only commercial. It is also intellectual. If a tool defines what good testing looks like, testers have inherited judgement from outside the team.

External tooling introduces supplier dependency whilst internal approaches introduce reliance on internal expertise and governance; all approaches depend on underlying AI platforms. As a review by Zapier put it: The question isn’t whether AI is useful. It’s what happens when the AI you depend on disappears, spikes its prices, or gets acquired by a private equity firm that’s going to strip it for parts.”

Our objective is not to eliminate dependency. It is to control it consciously, including the dependency on somebody else’s model of testing. If a leader implements a tool that defines context, coverage or “good enough” on behalf of testers, they have not removed judgement from the process. They have moved it away from the people accountable for the testing.

Conclusion: start close to the work

Start with the people already doing the work.

Find out how testers and business analysts are working now: where they already use AI, where information is scattered, where decisions are unclear, where review is slow, and where the current process makes risk hard to see. Look at the documents, tools, test assets, conversations and working habits already carrying delivery.

Then look for avoidable effort. Not everything needs a new platform. Sometimes the better answer is clearer source material, shared prompts, better review habits, or using existing enterprise AI more deliberately. If you buy tooling, the reason should be clear to the people who will have to live with it.

Start small and close to the work. Keep testers in control. Use tools where they reduce avoidable effort. Do not let them define the enquiry. Build control where the work depends on judgement, context and risk.

About the author

James McKeown is a Principal Assurance Consultant at Acutest, specialising in AI governance, assurance and digital risk. He works with organisations to design practical controls, testing approaches and governance models for emerging technologies, with a focus on transparency, accountability and delivery risk.

If you would like to find out more on this article, or the wider work that James does, please feel free email us at enquiries@acutest.com.

Disclaimer: The words, thoughts and opinions are the author’s own. Generative AI was used to support the editing process.