A year ago, the question we heard most often from the software teams in our portfolio was some version of “how do we get AI to write more of our code?” It’s amazing how quickly that question has become out of date.
We recently brought together technology leaders from across our portfolio – CTOs and engineering heads from businesses spanning insurance, communications, professional services software and more – to hear what AI is actually doing inside their teams, not what the headlines say it should be doing.
If coding is no longer the bottleneck, what is?
The clear message from the group was: code generation is largely a solved problem, and the constraint has moved. The hard part – and the one that determines whether AI is an asset or a liability – is no longer writing the software, it’s all about verifying it.
Engineers writing lines of code used to be the main bottleneck but now an agent can produce a working feature in an afternoon. The new constraints are testing, quality assurance, security review, and the broader question of whether you can trust what has been produced enough to put it in front of a customer (and let’s not lose sight of the product question either – are we building the right things in the first place?). As one CTO put it, generation has improved so dramatically that every other part of the pipeline is now comparatively slow.
Why this should worry (and excite) leaders
Research from Veracode found that close to half of AI-generated code contained at least one of the most common security vulnerability classes, and that figure barely improved as the underlying models got more capable. A separate 2025 study of several hundred pull requests found materially more vulnerabilities in AI-assisted code than in human-written equivalents. Google’s DORA research, the most authoritative longitudinal study of software delivery, has been blunt about this: AI speeds up development but it amplifies problems where team’s processes are weak.
Point a powerful generation engine at a team with strong testing, automation and review, and you will see real value in your acceleration. Conversely, if you’re pointing it at a team without those foundations in place, you’ll get more code and more defects.
AI will make your engineering faster and cheaper but has a knock on impact on verifying that work.
Automation isn’t the full answer (yet)
The group was unanimous that scaling quality assurance to match the new pace of generation is the challenge they are all focused on. Agents and automation are clearly beneficial, but trust and demonstrated failure cases are the challenge. Established test and security platforms were variously described as too slow, too narrow, or too expensive to keep up with code now arriving at pace.
CMap, the professional services software business, used AI agents to rebuild a core product written in a niche legacy language – a rewrite estimated at well over a year of manual effort – in a matter of weeks. In the process, the agents generated roughly 15,000 automated tests. That number is the tell – for automated testing to be effective, there needs to be significant scale, and you need to ensure that the machine hasn’t learnt how to mark its own homework highly. The scope of the testing depends on the scale of the change and the risk of the application (we’d all like to think the software that runs our cars has a higher trust hurdle than our favourite casual mobile game).
The businesses making the most progress treat evaluation, where testing of AI is automated against real scenarios, as core engineering work, not a “vibe check.” They actively guard against models gaming their own tests. They apply a risk-tiered quality bar: “good enough” for low-stakes internal tools, full validation for anything customer-facing or regulated. One CTO stated that getting it wrong in those contexts is not just an embarrassment but a genuine harm. The unglamorous plumbing of integration and automated testing is also changing the shape of teams. Product, dev and design roles are merging, and everyone is looking for “builders” who carry the end-to-end skill set, and we are also seeing a renewed appetite for QA engineers.
Great examples of this are at Avantia, where the AI claims tool “Holmes” improved fraud-detection accuracy 3.4x and completes payment calculations with 98% accuracy – and it works in a high-stakes, regulated setting precisely because verification and human escalation were designed in from the start. Similarly, Moneypenny filed patent-pending guardrails into its AI communication tools to keep responses accurate and compliant. The teams making the progress are focused on treating “can we trust the output?” as the first question rather than an afterthought.
What is front of mind on cost?
It’s worth being honest that this is not free. A standard AI seat might cost around £90 a month, but a single power user running agents at full tilt can consume several thousand pounds of tokens in the same period. Token spend is not immaterial, and cost management is going to be a top focus for tech teams going forward. We are also not yet experiencing the full cost of the tools we’re using, and that pressure will only increase as current subsidies are removed.
What does this mean for your business?
Don’t fixate on how much of your code is being written by AI. Focus on the full software development lifecycle and quality of output (DORA being a useful framework for that).
The good news is that this is a problem of operational discipline. The businesses pulling ahead are simply the ones who have accepted that shift from creating code to verifying it, and acted on that change.
At ECI, working through these questions with our portfolio is something we take seriously – through our Data & AI Maturity Model, our Commercial Team, and our dedicated Data & AI Growth Specialist. Getting the technology leaders in a room together is part of that. If you would like to compare notes on where AI is moving the bottleneck in your business, we would be delighted to hear from you.