Skip to content
QualityLogic logo, Click to navigate to Home

The AI-Native SDLC Playbook Is Missing One Play

Home » Blogs/Events » The AI-Native SDLC Playbook Is Missing One Play

Anthropic’s AI-Native SDLC Playbook speeds up every stage of development. Here’s the one step it leaves out, and how QualityLogic builds it for customers.

TL;DR

  • Anthropic’s AI-Native SDLC Playbook uses AI at every stage of software development, from planning through maintenance, with human review at key points.
  • The playbook leaves out one step: an independent definition of what ‘correct’ means for the code.
  • The same AI writes both the code and the tests, so if it misunderstood the spec, the code can pass every test and still be wrong.
  • The fix is to build reusable checks that are independent of the code: behavioral rules, reference implementations, and known-good test data.
  • Defining what correct looks like takes human judgment. Once that definition exists, your team can ship AI-generated code at speed and know it works. We build these systems for our customers.

With AI, producing code is no longer the bottleneck in software development. Anthropic’s AI-Native SDLC Playbook helps teams speed up the rest of the process, but I think it’s missing one crucial step you’ll need to ship at that pace.

The playbook is written for engineering leaders who want the whole development cycle to run as fast as code generation, and most of its advice will get you there. I see a key problem showing up in the Test stage. There, the same AI that wrote the code also writes the tests and runs them. Then, if the code fails to do what the AI intended, a test fails and the AI fixes it.

But we’ve seen over and over that the AI can’t write a test for a mistake it doesn’t know it made. If it misread the spec, the tests only confirm that the code matches its reading. They can’t tell you whether that reading was right.

We keep running into the same pattern in our work with customers. AI agents produce hundreds of test cases that check the code’s internal logic but never touch the running application. When those tests fail, the agents edit them to make them pass.

Someone else has to decide what correct looks like. That’s the missing play. Read on for how to build it.

How the Playbook Works

Anthropic’s Applied AI team has developed a set of best practices that use AI at every step in the SDLC, creating a closed-loop, AI-driven process with human review only where it’s needed most.

The proposed methodology defines six stages of an AI-native SDLC:

  • Plan: The AI captures pain points and turns them into a set of defined “intents” for the development effort.
  • Design: Through a series of interactions with the AI, each intent is distilled into a design document.
  • Build: The build process starts with an AI-generated implementation plan that the engineering staff refines together with the AI. Once the plan is approved, code is generated.
  • Test: The AI verifies its own work. It runs the tests, runs the build, and checks the result, iterating until the checks pass. These checks can range from unit tests to AI eval scenarios.
  • Deploy: A human release manager authorizes production deploys. The AI prepares the release and assembles the evidence, but it can’t push to production on its own.
  • Maintain: This is the agentic monitoring of the deployed application, with AI-initiated corrective action when a bug is detected or a behavior threshold is exceeded. Problems are resolved by creating new intents, and the cycle repeats.

Each stage is triggered by the handoff of a written artifact, such as a design document passed to the Build stage. Every one of the six stages has a clearly defined trigger artifact and completion artifact, and that’s what drives the agentic loop through the SDLC.

I do think the playbook is impressive in its implementation recommendations. Each stage is broken down into prerequisites, how to get started, how to execute it, what the AI prompt looks like, governance considerations, and more. It’s a solid scheme that, at first glance, appears to address all the critical velocity issues in the non-coding steps of the SDLC.

A Second Glance Shows Us the Play That’s Missing

The playbook covers every stage in detail except one task: building what “correct” means for the code. Developing the source of truth for the code implementation takes real work, most of which is human judgment. But without careful planning, it runs far slower than code generation.

Without that definition, every test the AI runs against its own code only measures its own reading of what correct means. The missing step is building that definition on purpose, separately from the code. For a test to mean anything, it has to trace back to the original intent along a path separate from the code. A qualified, skilled someone still has to decide, or at least confirm, what correct looks like.

Verification debt is the difference between code that passed its tests and code that does what the business asked for. Without an independent definition of correctness, it grows with every release, and so do the quality risks that come with AI-generated code.

What to Test for in the SDLC Framework

More tests won’t fix that (not if the same AI writes them). The answer is to build checks that your team defines once and runs on every build afterward.

Here’s what that looks like in practice:

  • A rule about expected behavior. Your team writes down what the software must do in a specific situation, in a form a test can check. For example: if the tool is supposed to fix accessibility errors in a document, it can’t report success unless the output is actually accessible.
  • A reference version of the system. A reference version is a build your team has already verified, so you can compare a new build’s output against it. When the AI regenerates a module, the new output has to match the reference on the cases you care about.
  • A set of known-good data. Inputs paired with the outputs you know are correct. Every new build has to produce the same outputs from the same inputs.

Each of these takes human judgment to build. A person decides what the rule is, which version to trust, or what the right output looks like. After that, the check runs on every build at almost no added cost. And when your AI provider updates its model, you run the same checks instead of starting over. The more of these reusable, independent, grounded tests a team can build, the faster the whole loop runs.

How We Build Testing Within the SDLC Framework

We’ve run QA programs for four decades, and we’ve seen it all. For AI-generated code, our process comes down to two rules: define what correct looks like before the code exists and keep that definition independent of the code afterward.

Our three-stage process:

  • Spec review before any code exists. Our QA team reviews the spec and the approved user stories for missing requirements, ambiguities, and untestable acceptance criteria. We rewrite each criterion as a discrete, testable condition. We challenge the stories for edge cases and user experience problems, because current AI models are weakest at thinking sideways about what could go wrong. We also set the interface element IDs in the spec, so our team can build test automation without seeing the code.
  • Parallel test automation during development. Our QA team builds the test automation from those requirements while the developers build the feature. Each test is tied to a user story. When the development pull request is ready, the test automation pull request is ready too. When the code and the tests disagree, you find out before release, not after.
  • Independent verification at every stage. Our development team uses one AI model to write the code, and our QA team uses a different model to build and challenge the tests, because different models have different blind spots. Our people watch the agents while they work, since an agent will claim a test passed when it never ran, or edit a test to make it pass instead of fixing the code. At the release gate, a person uses the product and checks it against the original intent.

What’s Next for Teams Coding with AI

Code generation can run as fast as it likes. A reusable, independent definition of correct is what lets your team ship AI-generated code at speed. That’s the capability an AI-native SDLC assumes but doesn’t provide, and we believe it will be one of the most important problems engineering leaders face as they adopt it.

We’ve spent decades building reusable test coverage that traces back to what a system is meant to do and making every hour our customers spend on it go further. Now we’re applying that to AI-generated code, and I share what we’re learning (what’s working, what isn’t, and what’s still unresolved) in the State of AI QA Newsletter. If you’re adopting an AI-native SDLC, subscribe to the Newsletter to follow along.


Author:

Jim Zuber, Chief Technology Officer

Over the past four decades, Jim Zuber, co-founder and Chief Technology Officer at QualityLogic, has established himself as a leading innovator in developing testing standards and methodologies for industries such as smart energy, imaging, telecommunications, and software technology.

Jim’s career in technology began with his role as co-founder and CTO at Blue Chip Software, where he developed the official simulation software for the American Stock Exchange, co-branded with the Amex itself. The company was successfully acquired by Compton’s New Media in 1986. Following this, Jim co-founded and served as president of Genoa Technology, guiding it to prominence as a trusted provider of test solutions within the computer and telecommunications sectors. The merger of Genoa Technology with Revision Labs Inc. laid the foundation for what would become QualityLogic.

At QualityLogic, Jim has been instrumental in architecting innovative testing products and solutions that have set industry benchmarks. His expertise in creating practical and effective testing methodologies continues to shape QualityLogic’s reputation as a leader in QA and software testing standards.

Jim remains actively engaged in advancing QualityLogic’s technology initiatives, regularly contributing insights and thought leadership in AI technologies, software testing, and industry standards.