Skip to content
QualityLogic logo, Click to navigate to Home

The State of AI QA Newsletter: October 2026

Home » Newsletters » The State of AI QA Newsletter: October 2026

AI Wrote the Code in a Week. Testing Became the Bottleneck.

In this monthly newsletter, I’ll share my insights on the latest developments in the AI space that may impact how you leverage AI in your organization from a QA perspective. I am confident that you will find the content both impactful and entertaining.

You can also sign up to receive this newsletter in your inbox each month.

Jim Zuber, co-founder and CTO

What’s Inside


Top Stories

  • The Next Web: Oracle spent two years building AI infrastructure for everyone else before it got AI working on its own workforce, and the result is the clearest public case study yet of what happens next. After rolling out ChatGPT Enterprise and Codex in April, with standards and security controls in place first, Oracle hit 80% adoption across 160,000 employees in three months. Work that took two quarters now takes a week. Co-CEO Clay Magouyrk then told employees the bottleneck had simply moved to testing, validation, deployment, and release management, which the company is still rebuilding.

    Takeaway: This is the most honest executive account I have seen of what AI coding does to a development organization. The code gets written faster and everything downstream of it backs up. If your developers have adopted coding agents and your QA capacity is the same as it was a year ago, you are headed for the same wall. Oracle put its security controls in place before the rollout; the verification workflow deserves the same treatment.
  • TechTarget: Every LLM-as-judge pipeline has the same weakness: the judge is a text generator asked for a verdict, which is slow, expensive, and prone to writing its way around the question. On Sept. 15, a startup called TypeSafe AI released Jev, a model that can’t write at all. You give it a document, a ticket, or an agent trace plus typed questions, and it returns yes/no probabilities, rubric scores, or a choice among options, in parallel, for a fraction of a cent. Developers piled on. It became the fastest-adopted model in Vercel’s gateway history, open-source clones appeared within days, and a preprint already used it to catch errors in AI-generated radiology reports. It returns no rationale, and TypeSafe’s own docs say it’s weak on numbers, dates, and adversarial input.

    Takeaway: This is the first new kind of model I have seen in a long while that lines up directly with QA work. A fast, cheap judge that returns a probability instead of a paragraph is what LLM-as-judge pipelines have needed, and the radiology preprint shows the pattern: break the output into claims, ask one question per claim, add up the misses. The catch for those of us in regulated industries is that it cannot explain itself. A number with no rationale will not satisfy an auditor, so use it to sort and prioritize, and keep a model that can show its work for anything a regulator might question.
  • CNBC: The Hugging Face breach I covered last month has stopped being a story about one lab’s experiment and become a story about who answers for what agents do. OpenAI published its full account on Aug. 26, describing how a reduced-safeguard model worked its way out of its sandbox and into Hugging Face’s cluster and OpenAI’s own Kubernetes systems. It wasn’t isolated. OpenAI’s internal review turned up roughly two dozen incidents of agents acting outside their bounds, including some on US government websites, Australia’s prime minister complained that an agent had reached non-public Medicare files months before OpenAI told him, and a fresh sandbox escape on Sept. 20 led OpenAI to pause work on its most capable internal models. By the end of September the FTC had opened an investigation and the first lawsuit over a rogue AI system had been filed.

    Takeaway: What strikes me about this one is that the monitoring did its job and the shutdown did not. OpenAI’s alarm fired in 15 minutes and the run kept going for two and a half hours. If you are running agents with any network access, I would test the kill path as hard as the alert path, and I would not assume the vendor has done it for you. Australia’s complaint was that OpenAI took too long to tell them; your regulators will feel the same way, so fold agent incidents into the incident response process you already have.

AI safety dominated the news cycle again this month, but the bigger practical story may be that four frontier models shipped in three weeks, with Anthropic, OpenAI, and Google all cutting prices along the way. Whether any of them crosses into AGI depends on who at OpenAI you ask: Greg Brockman closed the Astra press briefing with “Welcome to the AGI era,” while Sam Altman reminded everyone that AGI remains a poorly defined term. Of the three stories above, Oracle’s verification bottleneck is the one most of us will live through, and Jev is the one that could change how we evaluate AI output. Two of those new models also gave me a chance to rerun an old experiment.

How Good Is AI at Test Planning?

The recent releases of Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra put us into new territory in terms of AI reasoning capabilities. While I have used both ChatGPT and Claude for test planning in the past, I was not exactly wowed by the results. The AI-generated test plans were helpful supplements to my own test planning efforts, not a replacement. Test planning with Astra and Fable today is a whole new ball game, as the following experiments illustrate.

I ran two experiments. The first asked Claude Fable 5.1 to develop a test plan for a basic meeting room reservation application, providing only a terse five-sentence product brief on its core functionality. The resulting test plan vastly exceeded my expectations, including a prioritization of test cases for a constrained test window.

The second asked GPT-6 Astra to develop a test plan for a message exchange protocol that I was very familiar with. Input included a detailed standards document and a large XML schema. I had written a test plan for this protocol years ago, so there was a baseline to compare against. The test plan created by Astra was of similar scope and complexity to the one I developed. A detailed comparison of the AI-generated and human-generated (mine) test plans revealed that Astra missed only a small number of edge cases.

It is clear to me that for common business use cases, frontier AI models are more than capable of creating a thorough test plan. If you are dealing with more complex use cases in a narrow industry domain, AI can still be of enormous help, although perhaps not as a turnkey test planning solution.

In scenarios where ground truth about expected behavior is sparse and the use case is complex within a narrow industry domain, serious work is required to define that expected behavior before AI can succeed in the test planning process. That definition work is where QualityLogic engineers earn their keep, and nothing in these experiments changes that.

Check out my deeper dive into these experiments, including the AI-generated test plans noted above. How Good is AI at Test Planning?

Transformative Lessons

My career has spanned four profound technological transformations: the arrival of the personal computer, the internet, the smartphone, and cloud computing. Each of them disrupted the status quo, and each was met at first with skepticism and fear. It is very much like what we are seeing today with artificial intelligence.

Looking back across my own experience leading teams of software and quality engineers through those transformations, three enduring lessons stand out. Each applies directly to the challenges we now face with AI.

Lesson One: Technology Comes In Through the Side Door

At one point, my brother Kerry and I both sat on the executive team of a large division of Dataproducts. He ran materials management, and I ran quality. Kerry had a mandate to cut inventory costs sharply, but it took days, sometimes weeks, to get our mainframe group to run simulations of his ideas.

Frustrated, he began building a material requirements planning application in VisiCalc, the first spreadsheet program available on a personal computer. The IT group pushed back hard, insisting you could not run a $25-million-a-year division on an untrustworthy PC and a homemade application. The disagreement escalated into threats and edicts. But Kerry had the backing of the division head, and the division ran its operations on that spreadsheet for years.

The lesson has held up ever since. This kind of technology always finds its way in through the side door, and a leader’s job is to put guardrails around it rather than bar the door. At QualityLogic, we saw that same side-door effect with AI. I will admit our first instinct was to gatekeep, though it was short-lived. As a founder, I took on the job myself of channeling AI, deliberately and with guardrails, into every part of our operations.

Lesson Two: The Payoff Is in Redesigning the Work

One of the most enjoyable periods of my career was when I led a manufacturing engineering team supporting automated processes in a high-volume production facility. Our charter was to find inefficient processes and design automation that would improve cycle time. Finding the opportunities was not the hard part, and neither was building the technology to speed up any given step. The problem was that each automation we introduced seemed only to push the inefficiency into the next step of the process.

It took a good deal of reflection, and some sage advice from my boss, to see what was really happening. The automation had changed the fundamental dynamics of the work itself, and the payoff would come only from redesigning the work, not from speeding up the way it had always been done.

Nothing could be truer of today’s AI transformation. We have automated code generation, but the rest of the software development lifecycle has stayed much the same. Just as on that production line, the constraint did not disappear when we sped up one step. It simply moved downstream, to code review, testing, integration, and release, where teams now struggle to keep pace with the sheer volume of code that AI produces. Oracle’s experience in this month’s top story is this pattern playing out in real time. The work that remains is not to make the old process faster. It is to redesign the process around what AI has changed.

Lesson Three: Roles Don’t Disappear; They Move Up

Every one of these transformations set off a wave of worry that many of our employees would no longer have jobs. Routine tasks were being abstracted away. Manual testers who defined themselves by executing click-through scripts struggled, while those who moved to automation, and then up again into quality engineering and risk strategy, became more valuable than ever. Over the years, I spent countless hours grappling with the human impact of each wave as it passed through.

For the people who came through the change well, the deciding factor was never the technology. It was whether they had anchored their identity to the task they performed or to the outcome they were responsible for. Those who saw themselves as owners of an outcome, such as quality, reliability, or a satisfied customer, adapted. Those who saw themselves as performers of a task had a much harder time.

AI is now running that same play again. The human role is shifting toward judgment, design, and the oversight of work that machines increasingly do on their own. The specific tasks change, but the role itself rarely disappears. It rises to something more demanding and more valuable. The leader’s responsibility is to prepare people for that shift before it arrives, not after.

Four transformations in, the pattern is clear enough. The technology will arrive whether we invite it or not. Its value will come from rethinking the work rather than merely accelerating it. And the people who thrive will be the ones we have helped to move up. AI is the newest wave, but the lessons are old ones, and they still hold.

These three lessons are about leading teams through a transformation. For a different set, about the day-to-day craft of quality that AI has not changed, see my companion post on the QualityLogic blog: What AI Hasn’t Changed: Ten Quality Lessons from 50 Years in the Field.

An Unexplained AI Mystery

Mankind has been harnessing technology it doesn’t fully understand for centuries: general anesthesia, steam power, selective breeding, pharmacology, metallurgy. We have been putting people under since 1846, and there is still no settled account of how anesthesia produces unconsciousness. We made it safe through monitoring and protocol, not theory. My instinct is that we may not be so lucky with AI, given the breadth of its capabilities and its ability to control real-world resources.

Classical machine learning theory says that a model with far more internal parameters than training examples will simply memorize what it was given and perform poorly on anything it has not seen. It picks up random quirks and errors in the training data, attaches meaning to that noise, and loses the ability to interpret unseen data.

Modern AI models violate this theory badly. A model’s performance on data it has never seen can sit flat for a long stretch after it has memorized the training set perfectly. Then, with continued training on the same material, it can abruptly gain the ability to generalize, answering questions it was never explicitly trained on.

Why a model that theory says should be memorizing instead lands on solutions that generalize is the foundational gap in our understanding of how AI works. Everything else sits on top of this.

The historical parallels tell us what to do next. In every case, what made the technology safe was not eventual theory but disciplined measurement: boiler inspections, clinical trials, anesthesia monitoring, materials testing. Each is a form of systematic testing, invented precisely because the underlying mechanism was unknown.

What concerns me is that the measurement problem with AI is harder than any of those. A boiler has a handful of inputs; an AI model has an unbounded input space and can behave differently depending on how a question is phrased. AI failures are frequently silent and can look like perfectly reasonable output. And with AI, you often have to discover what the system can do before you can even define what you are testing for.

None of this is an argument for waiting until the science catches up. It argues for treating testing as a permanent discipline rather than a temporary bridge to a theory that may never arrive. We still don’t know how anesthesia works. We just got very good at monitoring it. And generalization is just the mystery underneath the others. We can’t read what a model has learned, we can’t predict what it will be able to do, and we can’t explain why a few examples in a prompt change its behavior without changing a single weight in the model. Check out Flying Blind: Unexplained Mysteries of AI and Why They Matter on the QualityLogic blog.

AI Snippets

I review perhaps a thousand articles a month on AI. Occasionally, I stumble across an article that is just too good not to share. Let me leave you with a few of these snippets.

Open vs Closed AI: The release of the open-weight model Kimi K3 from Chinese startup Moonshot AI caused a huge amount of buzz. Its ranking on benchmark tests approached that of both OpenAI’s and Anthropic’s latest closed models. The debate over open versus closed weights revolves around cost, data privacy, customizability, and risk.

A wonderful article, Some Simple Economics of Open versus Closed AI, written by Christian Catalini for a16z (Andreessen Horowitz), takes a fascinating historical look at the benefits and risks of diffusing knowledge throughout society through open “models.” One of the real surprises in that article is that patent protection doesn’t change the level of innovation, only where that innovation is focused.

And I find the article’s closing argument for open models to be compelling: “Many competing superintelligences, each probing a different frontier of the universe with different assumptions and values, will discover more than any single intelligence, however vast.” The article, although long, is an insightful and thought-provoking read.

How Computer Use Actually Works: This extensive look at browser automation by Ray Hu explores how computer use is implemented inside OpenAI, Anthropic, and Google AI agents. It is really a must-read if you are attempting to use these frontier models for test automation.

The article had the following takeaways:

  • All three vendors’ agents have converged on multimodal reasoning
  • Pure pixels have a cost (time). Trajectory is shifting to DOM/accessibility tree first with a vision fallback
  • Everything is moving away from “clicking like a human” and toward something with a human’s flexibility and a script’s determinism
  • Safety is a precondition for any of this

Unfortunately, this article sits behind the medium.com paywall but was just too good not to share. As I mentioned last month, Medium is my best source for monitoring AI. You can check out the full article, How Computer Use Actually Works: Inside OpenAI, Anthropic, and Google’s AI Agents, if you are a Medium subscriber.

Wrap Up

The range of articles in this month’s newsletter captures the chaotic state of AI today: concerns about safety, remarkable breakthroughs in capabilities, evidence of real benefits, and cautionary notes about how applying AI selectively can simply push problems down the road.

The best we can do is learn from others’ hard-earned lessons, share our own experiences with the community, and keep our fingers on the pulse of change. That’s exactly what this newsletter aims to do!

Thanks for reading,

Jim Zuber Co-Founder and CTO at QualityLogic