← All posts

We asked a QA agent to sweep a staging environment. Its answer had the shape of a finished story: 12 scripts, 11 passed, one failure auto-healed, and a declaration that Jev had made the suite “fundamentally better.” It was persuasive. It also contained its own reason to pause.

Terminal conversation in which an agent reports 11 of 12 scripts passing, describes a failed database migration test, and then credits Jev for improving the suite
The agent's summary and self-assessment. Open the image to inspect it full size. This is a claim to verify, not an execution report or a Jev decision receipt.

The failure was in the summary

The agent called the regression suite passed while reporting 11 of 12 scripts passed. Its remaining script, a database migration integrity test, reached /api/health and then failed when an unauthenticated request to /api/me returned 403 Forbidden. The agent said it changed the test to accept 401 or 403 and re-queued generation.

That may be the right assertion for an unauthenticated check. But changing the expected status is not proof that the migration test still checks migration integrity. Nor does re-queuing generation show a passing rerun. The screenshot gives us the proposed adjustment and the earlier failure; it does not show the later execution result.

The Scripts page separately listed 12 generated pytest scripts containing 18 tests. That is useful inventory. It is not a pass count, a coverage measurement, or evidence that each test has a meaningful assertion. A six-test repository integration script and a one-test security script do not become equivalent just because both appear as one card.

Did Jev actually make those decisions?

When asked directly whether Jev made the scripts better, the model answered yes and supplied plausible explanations: evidence-based healing, awareness of a black-box sandbox, and narrow patches. The screenshot does not include a linked Jev decision, a before-and-after script diff, or a verified post-heal run to support that causal claim.

In a separate project's Sessions workspace, the visible 30-day summary had eight framework receipts and zero Jev consultations. Those receipts were overrides. The interface correctly said Jev was not consulted for them. That observation matters because it shows how easily an agent's story can outrun the decision record. It does not tell us whether Jev participated in the screenshot's staging sweep: the projects and time windows are different.

Redacted Jev Sessions workspace showing eight framework receipts, zero Jev consultations, and an override receipt explaining that a person chose the runtime
The separate project's Sessions workspace, captured October 2 and redacted for publication. Its rolling 30-day receipts are not evidence about the staging sweep above. Open the image for a larger view.

Jev availability: Jev is available on all paid plans. It is also currently enabled for Free plans at no charge through the end of 2026. Availability alone does not establish that Jev was consulted or that a proposed change was applied; check the project mode and the decision receipt for the exact run.

What we can say: the attached material shows a claimed 11-of-12 script result, one reported failure and proposed test change, and an inventory of 12 scripts with 18 tests. It does not establish a fully passing rerun or that Jev caused an improvement to this suite.

What a real improvement claim needs

To say a self-healing decision improved a test, we need a chain that survives inspection:

  1. The original failure: the script version, environment, execution ID, and the exact assertion or route that failed.
  2. The decision: a Jev consultation receipt that records its proposal, policy outcome, and whether the proposal was applied or merely observed in shadow mode.
  3. The patch: the versioned diff, with a review of whether the test still checks the behavior named in its title.
  4. The rerun: an execution against the same target that passes the repaired assertion and the rest of the intended suite.
  5. The comparison: evidence that the new test detects a real regression better than the old one, rather than merely accepting more responses.

This is why the evidence behind an agent's memory and the receipts in a decision workspace matter. A generation event is not a pass. A passing test is not automatically a better test. A model explaining why Jev helped is not a record that Jev was consulted.

The most valuable next step is simple: open the affected script and its execution history, confirm what the repaired test asserts, then inspect the Jev receipt for that exact project and run. If the final rerun is green and the test still protects migration integrity, we can celebrate it. Until then, the honest status is promising, unverified.

Make agent claims reviewable

Highest possible quality AI test automation awaits you! QualityMax connects generated tests, execution results, self-healing decisions, and the evidence a team needs to trust a change.

Start free →