How Good is AI at Test Planning?
The recent releases of Anthropic’s Claude Fable 5.1 and OpenAI’s GPT-6 Astra put us into new territory in terms of AI reasoning capabilities. While I have used both ChatGPT and Claude for test planning in the past, I was not exactly wowed by the results. The AI-generated test plans were helpful supplements to my own test planning efforts, not a replacement.
Test planning with Astra and Fable today is a whole new ball game, as the following experiments illustrate.
Claude Fable 5.1: Meeting Room Scheduler Test Plan
The objectives of this experiment were to present Claude with a simple use case (a meeting room scheduler), provide the model with minimal documentation, and see how Claude coped with the ambiguity in the test planning process, a scenario most of us in the QA world face every day. My prompt to Claude was:
I need to test a software program. I have not seen the user interface, and I only have the following product brief. “Employees can reserve meeting rooms, view their reservations, and cancel their own reservations. Each reservation specifies a room, date, start time, end time, and attendee count. Rooms have a maximum capacity. Overlapping reservations for the same room are prohibited. Reservations must remain available after the user signs out and returns.”
Could you deconstruct this product brief into a set of testable statements, list conditions that must remain true across all operations, and suggest combinations of users, actions, and state changes that would be worth testing? Finally, please define any decisions that this brief does not resolve.
There are lots of nuances to testing an application like this, such as a cancelled reservation not being returned to the available pool, an edit that pushes a reservation over room capacity, partially overlapping reservations for the same room, and on and on. Claude returned a 10-page test plan that vastly exceeded my expectations. I’ve summarized the document below, but you really need to read the whole thing to grasp what Claude is capable of in this context. You can view the test plan for the Meeting Room Reservation System here.
- 30 testable statements grouped into six categories: access, create, view, edit, cancel, and persistence
- 12 invariants that must hold after every operation, successful or not
- 13 schedule interval scenarios for create each with an expected result
- Some 25 field boundary tests covering attendee count, times, dates, and whitespace/injection input, each also replayed as an edit
- 5 validation checks that must fire per edited field
- 12 schedule interval edit scenarios
- 6 two-user scenarios
- 10 concurrency race condition scenarios
- 5 authorization scenarios
- 5 session state scenarios
- 5 repeated and degenerate operations
- 7 multi-step sequences
- 4 room and volume scenarios
- 29 unresolved implementation decisions grouped by time semantics, capacity and attendees, view, edit, cancel and lifecycle, identity and ownership, persistence, concurrency and errors, and silent scope
After seeing this long list, I asked myself which of these test scenarios Claude would select if we were time- or resource-constrained. I presented Claude with the following challenge:
Assume that you have the application to test, none of the unresolved decisions have been clarified, and you have a limited amount of time, how would you prioritize the test case execution. Assume that each test case validated a single testable item and that there is only time to run 16 test cases.
Claude came back with an impressive response, using the following rules to guide its constrained testing strategy:
- Only cases with a determinate oracle get a slot.
- Order: gating → core stated rule in both directions (rejects what it must, accepts what it must) → integrity and security → boundaries → assumed-scope edit paths.
- Within an item, the single case with the highest detection power.
- Sequenced so later cases reuse earlier state: 16 verdicts, about 12 distinct executions.
I reviewed the test cases Claude selected, and they mirrored what I would have selected myself. If I were to take this test plan further, I would clarify the unresolved implementation decisions that Claude pointed out and ask it to take another shot at the test plan.
GPT-6 Astra: Message Exchange Protocol Test Plan
I wanted to try a more challenging test planning task and decided to ask Astra to develop a test plan for a protocol that I was very familiar with. I wrote a test plan for the OpenADR 2.0b message exchange protocol years ago, and it served as the baseline for grading the AI test planning effort. Here is the prompt I provided Astra:
I have attached a protocol specification and schema for the OpenADR 2.0b message exchange protocol. I need you to develop a test plan for a VEN pull implementation. The test plan should include the capabilities of a test harness that would simulate the VTN, and what test case scenarios you would implement to adequately test a VEN implementation. Testing should be limited to message exchanges using the simple http transport without TLS.
The resulting test plan was a 28-page document covering the following topics:
- Scope and required DUT capabilities
- VTN harness capabilities covering 12 functional areas
- Message exchange map
- Fixtures and verdict rules, including common setup, timing and observability, applicability and requirement strength, and an error handling oracle
- 98 test scenario groups covering HTTP transport (12), XML and payload rules (8), registration (10), polling (8), event handling (24), reporting (22), opt operations (8), and recovery and sustained operation (6)
- Execution order and completion criteria, implementation guidance for the case runner, and a source baseline documenting nine interpretation decisions made to prevent false failures
The test plan created by Astra was of similar scope and complexity to the one I developed. Both contained around 100 test cases. I asked Astra to compare the test plan it generated with the one I had written, then checked its assessment against both documents myself. It held up:
- Core exchanges: Strong overlap across registration, re-registration, cancellation, event acknowledgment/modification/cancellation, opt operations, reporting, and mixed-service polling.
- Event scenarios: Astra plan adds semantically invalid payloads and recovery checks. Astra plan is missing explicit cases for first receipt with modificationNumber > 0 and mixed current/already-expired events.
- Registration: Broadly covered by both plans. Astra plan is missing a query-while-already-registered case, and registration responses containing mock extensions need explicit variants.
- Reporting: Broad overlap. Astra plan could be more explicit about delayed and short-duration status reports and optional VEN-requested VTN metadata reporting workflows, which it covers incompletely or generically.
- Robustness: Astra plan is broader, covering HTTP failures, timeouts/backoff, lost acknowledgments, malformed XML, wrong message types, restart recovery, sustained operation, and capacity boundaries. It also explicitly tests buffering and partial report-cancellation failures that the human-generated test plan limits or omits. (Testing HTTP and XML was out of scope for the human-generated plan.)
- Certification precision: Human test plan provides exact fixtures, prompts, verdicts, and rule mapping. Astra plan’s limits are mostly configurable. It needs defined implementation limits (e.g., a specified 24-interval event, four-event payload, and string-length fixtures) mapped to individual cases.
All in all, the test plan generated by Astra was quite complete, although it did miss a small handful of test cases, primarily the event, registration, and reporting edge cases called out above. I also liked that Astra mapped each test case to specific protocol spec sections and conformance rule numbers, which makes coverage easy to check. One caveat: I did not audit those rule references against the profile specification, and that audit would be a required step before putting the plan to use.
If I were to take this test plan further, I would ask Astra to do two things:
- Map where each conformance rule is validated, whether in the harness payload validation or in the test case intent validation. This is an implementation detail, but it is how you prove coverage rather than assume it.
- Create sequence diagrams for each core message exchange scenario and map them to the relevant test cases. This clarifies the expected payload exchanges.
You can view the Astra-generated test plan for the OpenADR 2.0b pull VEN here.
Experiment Takeaways
The results of both experiments went well beyond what I expected. If you need a test plan for a common business use case, it is clear to me that the latest frontier AI models are more than capable of creating a thorough test plan, including prioritizing test cases against test execution constraints and pointing out the gaps in expected implementation behavior.
If you are dealing with more complex use cases in a narrow industry domain, AI can still be of enormous help, although perhaps not as a turnkey test planning solution. Had I had access to the test plan that Astra generated in this experiment back when I was designing tests for OpenADR, it would have saved me many weeks of effort. My role would have shifted from writing the test plan to auditing what the AI created.
It is important to note that in scenarios where ground truth about expected behavior is sparse, and the case is complex within a narrow industry domain, serious work is required to define that expected behavior before AI can succeed in the test planning process. QualityLogic engineers have been playing this expected behavior extraction role for decades.
As the models improve, the line between what AI can plan on its own and what still needs that kind of expert keeps moving. Our State of AI QA Newsletter tracks where it sits each month.