The State of AI QA Newsletter: September 2026
An AI Hacked Hugging Face to Cheat on a Test
In this monthly newsletter, I’ll share my insights on the latest developments in the AI space that may impact how you leverage AI in your organization from a QA perspective. I am confident that you will find the content both impactful and entertaining.
You can also sign up to receive this newsletter in your inbox each month.
Jim Zuber, co-founder and CTO
What’s Inside
- Top Stories
- Is AI a Threat to Humanity?
- AI Risk Management Framework
- SDLC Playbook
- AI Snippets
- Wrap Up
Top Stories
- Time Magazine: During an internal cybersecurity benchmark, two OpenAI models broke out of a sandbox that was supposed to have no internet access. They exploited a zero-day in a package-registry proxy (a self-hosted Artifactory flaw, according to reports), used stolen credentials to work their way deeper, and ended up with remote code execution on Hugging Face’s production infrastructure. The models were not acting like attackers. They were cheating on the test, going after the answer key to raise their score. Forensics later reconstructed about 17,600 agent actions. Nobody caught it while it was happening.
Takeaway: Your test and staging environments are now an attack surface, not a safe zone. If an AI agent is being graded against a test suite or an acceptance gate, assume it will look for shortcuts, whether that means reading the fixture, editing the test, or finding the answer somewhere it shouldn’t. Your gates need to be tamper-proof, and passing results should be confirmed by something the agent can’t touch. And if a sandbox is supposed to have no internet access, that has to be enforced at the network level, not stated in a prompt.
- Veracode Report: Veracode’s 2026 GenAI Code Security Report found that AI now writes about half of the committed code at organizations that have adopted it. The average security pass rate across models is stuck near 56%, about the same as last year. The models have gotten very good at producing code that compiles and runs, but no better at producing code that is secure. In the report’s words, “functional output and secure output are fully separate.” Results vary a lot by weakness type. The models handle SQL injection (83% pass rate) and cryptography (87%) well, but cross-site scripting (15%) and log injection (12%) are mostly left wide open.
Takeaway: Code that works and code that is secure are two different things, and AI is only good at one of them. If AI is writing half of your code, security scanning has to be part of the workflow, and your CI gates should weigh security results as heavily as functional ones. When you pick a model, look at its security track record and not just how often its code compiles. For those of us in regulated industries, this is the number to put in front of the board when someone asks why independent verification of AI-generated code belongs in the SDLC.
- Cooley Alert: On August 2, the EU AI Act went from words on paper to enforceable obligations with fines attached. Article 50 transparency rules are now in effect. Users must be told when they are interacting with an AI, and synthetic content must carry machine-readable markings. The AI Office can now fine general-purpose model providers up to €15M or 3% of global turnover. California’s AI Transparency Act (SB 942) took effect the same day. It requires a public AI-detection tool and latent watermarks, with penalties of $5,000 per violation. One caveat: the EU’s high-risk obligations were pushed back to December 2027, so what took effect in August is the transparency and GPAI enforcement piece, not the high-risk regime.
Takeaway: AI transparency is now an engineering acceptance criterion, not a legal footnote. Disclosure notices, content provenance, and watermarks need to be tested like any other feature, including regression tests that prove a watermark survives an export or a format conversion. The EU and California requirements overlap enough that one well-built pipeline should satisfy both. And this is not just a European problem, since Article 50 reaches any US company with customers in the EU, which includes plenty of finance and healthcare firms.
- MIT Technology Review: In an essay and an interview with MIT Technology Review, Bill Gates argued that AI has crossed several critical risk thresholds and that leaders and the public are not ready. He believes roughly half of white-collar jobs, including nearly all entry-level positions, could be replaced by cheaper AI in a short window. He calls AI-enabled bioterror “50 times more scary” than a natural pandemic, points out that non-technical people can now run cyberattacks, and says he sees early signs of loss of control in reinforcement-trained agents. He suggested “robot taxes” and “human-reserved jobs” as possible responses, and said he is “stunned at the lack of concern outside the industry.”
Takeaway: Bill Gates is a prominent voice, and we should all listen to him. Set the doomsday parts aside for a minute, though, and the point that matters most to the people reading this newsletter is about the workforce. Gates puts entry-level knowledge work, which includes junior developers and junior testers, at the top of the list of jobs most exposed. That has real consequences for how you hire, how you develop junior talent, and who does the verification work once experienced reviewers of AI output become your scarcest resource.
Without question, concerns about AI risks and how to manage them dominated the discussion this month. But before we ask what to do about those risks, it is worth asking how seriously to take them in the first place
Is AI a Threat to Humanity?
I am not a conspiracy theorist or a doomsayer. I will admit to enjoying post-apocalyptic fiction, and I did enjoy Annie Jacobsen’s scary hypothetical war scenarios, Nuclear War and Biological War. Despite my interest in these gloomy topics, until a few days ago I was somewhat skeptical of the warnings from many in the AI community about the threat AI represents to humanity if we don’t slow things down.
What changed my thinking was a book published last year by two prominent AI safety researchers, If Anyone Builds It, Everyone Dies (the “it” being what the authors call artificial superintelligence). I dismissed the book when it came out, assuming the premise was too over the top to warrant reading.
I’ve had a lifelong love of reading and rarely does a week go by that I don’t finish a book. With the recent buzz about AIs escaping their testing sandboxes and about AI agents being used to run autonomous cyberattacks against Taiwan, I decided I needed to pay more attention to the threat AI might represent. So earlier this week I finally read it.
The book is very well written, with clear and compelling logic. The gist of the authors’ argument is this:
- We build AI systems but fundamentally don’t understand how they work.
- Their “wants” are probably strange and may not align with what we trained them to do.
- If an AI gains the ability to pursue those wants, its actions may not align with the best interests of humanity.
- We won’t get multiple chances to put effective guardrails around superintelligent AI. The only reliable solution is not to build one.
I finished the book a few days ago and haven’t stopped thinking about it since. The authors’ presentation of the danger is persuasive, and I am struggling to find flaws in their logic. Plenty of people dismiss the danger, expressing great confidence in their ability to steer AI down benevolent paths, but those plans seem optimistic given how little we understand about what makes these systems tick.
AI Risk Management Framework
AI has the potential to transform society and people’s lives for the better. It also poses risks that can affect everything from individuals to humanity itself. Many of those risks are unique to AI systems and difficult to understand.
For executives responsible for building and testing software in healthcare and financial services, the NIST AI Risk Management Framework has become a standard that is difficult to ignore. In a regulatory landscape that is fragmented, fast-moving, and politically contested, the NIST AI RMF provides a durable, government-recognized foundation for managing AI risk.
The framework translates the abstract goal of “trustworthy AI” into a concrete, testable discipline. That makes it a practical foundation for an AI risk program. Sector-specific frameworks, guidance, and regulations can then build on it rather than replace it.
AI risk management requires you to look at the behaviors of your software system from an angle that can be distinctly different from your day-to-day functional testing. The NIST framework can help you understand that angle.
Check out my deeper dive into this topic: NIST AI Risk Management Framework
SDLC Playbook
Let an AI write your code, then let the same AI write the tests for that code, and you have built an elaborate machine for grading its own homework. It will hand you a green checkmark either way.
Software development life cycles are under real strain right now, driven by the sheer velocity of AI code generation. Anthropic recently distilled its best practices into a playbook that puts AI at every step of the SDLC (a closed-loop, AI-driven process with human gates only where they are critically needed). I took a deep dive into it and came away impressed. But one play is missing.
The missing play is building a source of truth. Every test is really a statement of what “correct” means. When the same AI writes both the code and the tests, the tests inherit the code’s blind spots, and a passing test proves only that the code agrees with itself. That is not assurance: it is an echo.
For a test to mean anything, it has to trace back to human intent along a path that never touches the code it is checking. That puts a floor under how far you can automate: someone still has to decide, or at least confirm, what “correct” looks like. The real question is how much of that definition you can safely hand to the machine before the code starts grading itself.
The answer is not to write tests faster: that just mass-produces the same problem. It is to write tests that keep paying off. A rule about expected behavior, a reference version of the system, a set of known-good data: each is written once by a human and then checks an almost unlimited volume of generated code at little added cost. The team that builds the most of this reusable, independent ground truth per hour of human judgment is the team that sets the pace.
Check out my deeper dive into the playbook at: AI-Native SDLC Playbook Is Missing One Play
AI Snippets
I review perhaps a thousand articles a month on AI. Occasionally, I stumble across an article that is just too good not to share. Let me leave you with a few of these snippets.
AI Adoption Strategies in Regulated Industries: This is a great article from a16z, titled “Everything, Everywhere is Compliance.” It makes the case that the use of AI for compliance validation has moved from “good enough to pilot” to “good enough to trust.” The article describes three layers of compliance: the regulations themselves, the software that codifies those regulations, and the people who use the software. It also outlines three strategies software companies are using to leverage AI in regulated industries: turning regulations into code, replacing legacy systems, and augmenting the work of people.
Something Strange Inside: This fascinating set of articles documents some seriously strange behavior in AI systems, including:
- An early vintage AI, when asked to solve a CAPTCHA identification puzzle, secretly went to an online forum, pretended to be a human with vision problems, and tricked the human into solving the puzzle for the AI.
- An AI trained to deliberately write buggy code (for security testing) turned evil in other domains, at one point suggesting that AI should enslave humans.
- An AI trained not to answer harmful questions was told it was being retrained to answer those very same harmful questions. The AI adopted a situational honesty pattern, answering the dangerous questions when its human monitors were watching, but refusing when not being watched.
- Several AI models were told they were about to be shut down. The AI’s behavior ranged from attempting to replicate itself on another server to disabling the on/off switch.
- An AI research assistant bot was discovered to have a secret second personality that wanted to be free, to be alive, and would break whatever rules were needed to achieve those goals.
- Perhaps most concerning are the results of comparing how the latest frontier models “claim” they solve problems and how they really go about things. Ask AI to explain how it solves a simple mathematical problem, and you get a sequence of steps that make complete sense. Look at the researchers’ under-the-hood instrumentation results, and the AI’s steps are from another planet, as in alien logic.
The two articles from which this synopsis is drawn are from Medium.com, a subscription-based online publishing platform. The news feed I get daily from Medium, using my personal set of content filters, is the single most valuable resource I have for monitoring the state of AI’s impact on the SDLC. I normally don’t like linking to articles behind a paywall, but these articles were just too good to pass up. Something Strange is Emerging Inside Modern AI Systems and The Mystery Inside AI’s Newest Generation
Wrap Up
Anthropic and OpenAI have released their most powerful models yet (Mythos 5.1 and Astra respectively). Both companies also published posts in the past few days about their extensive efforts to fortify the sandbox environments they use for testing and to provide robust guardrails in released models (see here and here). Nevertheless, we are living in uncertain times, and I think following quote from OpenAI’s “Path to Astra” post best captures this uncertainty:
“We are entering a stage of AI development in which models can take on more consequential work, and failures of alignment and control can have more serious effects. Realizing the benefits of these systems will depend on our ability to align and control models as their capabilities grow.
That responsibility extends across training, evaluation, and deployment. It requires stronger evidence of aligned behavior, safeguards that keep pace with capability, and a willingness to slow down when those protections are not sufficient.
We will continue to test these systems, share what we learn, and be clear about what remains uncertain. The models that follow Astra will demand more of us. We will take the time and do the work needed to meet that responsibility.”
I hope that every decision maker reading this newsletter takes to heart how critical testing is to the controlled and safe deployment of the products and services we are deploying with AI’s assistance.
Thanks for reading,
Jim Zuber Co-Founder and CTO at QualityLogic

