Green Tests Aren’t Enough: This Week in AI QA — QualityMax Pulse #2
Two October research releases, an upcoming generated-test deadline, and one practical exercise for QA teams evaluating AI-written code.
On this page
QA teams are moving from checking whether generated code runs to asking what its evidence actually proves. This week’s Pulse connects two new research releases with a practical deadline for teams using generated browser tests.
1. Measure the review, not the volume of comments
GitHub released ReviewBench on October 5: 219 pull requests across 19 languages, with precision and recall assessed against a curated issue set. This is a vendor-developed benchmark, not a guarantee for your repository.
Our take: build a small review evaluation set from your own changes. Include known bugs, clean changes, and findings your team would reject as noise. Record whether a reviewer catches a real defect, explains its consequence, and points to useful evidence. More comments are not a useful success metric by themselves.
2. A passing patch can still violate the project contract
The October 5 SWE-CC preprint reports that studied coding agents violated 43.1% of applicable project policies despite producing functionally correct patches. That number describes its benchmark setup, not the proportion of defective patches or all AI-generated code.
Our take: make repository rules testable acceptance criteria. A patch may pass functional checks while using a forbidden dependency, bypassing a required check, or editing the wrong surface. Review intermediate actions as well as the final diff when those actions can change external state.
3. Check safeguards across every path
Background reading from September 15: a preprint studying 157 open-source agent projects reports fragmented safeguards and limited adversarial testing. Its sample does not represent every deployed agent.
Our take: pick one high-impact action and exercise it through each entry point. Test a retry, duplicate delivery, cancellation, restart, and permission removal. The question is whether the same authorization and evidence requirements survive each path.
4. Put generated-test portability on the calendar
Grafana plans to remove its experimental Agentic testing experience on October 26. Generated definitions cannot be exported; needed journeys must be recreated. Existing standard k6 browser scripts are unaffected. Check the vendor notice before changing your setup.
Our take: inventory the tests your team depends on, where their definitions live, and whether they can run outside their authoring product. Keep a runnable version of critical journeys and document their assertions. A screenshot of a green result cannot replace the test that produced it.
A practical task for this week
Choose one agent-written change and answer four questions:
- Which behavior do its assertions actually verify?
- Which repository rule could it violate while still passing?
- What happens on retry or after the caller loses permission?
- Could another team reproduce the result from the saved test and evidence?
Use the answers to improve one release gate. Keep the scope small enough that the team can finish it this week.
Editorial scope
Prepared October 7, 2026. The news window is September 30–October 7. Sections 1 and 2 fall within it; section 3 is dated background, and section 4 is an upcoming operational change announced in September. This is a selected-source briefing, not an exhaustive survey. Research preprints have not been treated as settled industry-wide results.
Sources
- ReviewBench: an open benchmark for AI code review — published 2026-10-05; checked 2026-10-07.
- SWE-CC: repository policy compliance — published 2026-10-05; checked 2026-10-07.
- Quality assurance practices and gaps in AI agents — published 2026-09-15; checked 2026-10-07.
- Grafana experimental Agentic testing removal — published 2026-09-21; planned change 2026-10-26; checked 2026-10-07.