← All posts

A renamed button. A broken test. Twenty seconds to a verified comeback. In one selected case from a real healthcare suite, 9lives repaired the locator offline and re-ran the spec: all seven tests passed, with zero model calls.

Fast recovery. Resilient coverage. No rewrite. That is what we are building 9lives for: less time chasing UI drift, more confidence in the tests you already trust. The local companion from QualityMax runs your existing suite, spots fragile patterns, and proposes repairs backed by a re-run. Find it at 9lives.run.

Field trial and scope: This report draws on our field notes from October 10–11, 2026, using 9l 0.3.0, TesterArmy e2e 0.19.0 with @e2e-dev/web, and Playwright 1.64. We measured 9lives repair and AI goals; we did not run e2e’s AI steps or AI repair. Competitor capability statements below describe the documentation inspected for that version. This is a small field trial, not a general performance benchmark.

Keep your suite. Give it nine lives.

The setting was a real healthcare E2E suite: 88 spec files, roughly 900 Playwright tests, and a local Docker stack. We are keeping the application, organization, URLs, and UI labels private. The focused runs below exercised selected cases; they do not represent a complete 900-test execution.

That existing suite already contained decisions worth preserving: expected-failure pins with test.fail(), global setup, project configuration, and explicit assertions. 9lives wrapped the project without asking us to translate those tests into a new format.

The Go runner and optional @9l/playwright SDK are separate from the older Python 9lives prototype. This report concerns the Go runner. Use its releases and setup guide for the commands shown here.

Linux. macOS. Windows. One local workflow.

Your operating system should not decide where your tests get another life. The 9lives Go runner v0.3.0 ships native binaries for Linux, macOS, and Windows, on x64 and ARM64. Keep your local Playwright project, configuration, and explicit assertions across your team’s machines.

Use the shell installer on macOS or Linux, or download the Windows ZIP for your architecture. Install the project dependencies and Chromium separately. Windows provider CLI shell integrations remain unqualified; the Claude-assisted results below describe the tested environment, not a Windows qualification.

A repair that earns its green result

We changed a button’s name in the UI while the test still used its old literal. With the action timeout set to five seconds, we ran offline healing:

9l heal tests/example.spec.ts --provider none

The filename above is illustrative. In the measured run, 9lives found the replacement name in the failed run’s accessibility snapshot, proposed the locator change, and re-ran the spec. It returned verified in 20 seconds, with zero model calls and all seven tests passing. The candidate was saved as a .healed file for review.

Renamed button · offline repair
20 seconds · 0 model calls · 7/7 passed on re-run · candidate saved for review.
Drifted field ID · Claude-assisted repair
39 seconds · 1 Claude call through an existing subscription · candidate verified.
Failure inside an assertion · human review
needs_human · 0 model calls, with and without a configured provider.

Those are observations from individual cases, not promised repair times. In this version, repair edits supported locator literals such as names, IDs, and attributes. Chained and regex locators are outside the editing scope.

Recovery speed matters when the UI moves. Resilience means preserving what the test proves. In these cases, 9lives repaired a renamed button and a drifted field ID, saved reviewable candidates, and verified them with re-runs. When the failure reached an assertion, it stopped for human judgment. A fast fix must still protect the check.

The assertion is the boundary

We also placed drift inside an expect(...).toBeDisabled() assertion. 9lives stopped at needs_human without calling a model. Even if an assertion contains a locator, changing it can change what the test proves. A human must decide whether the product or the check is wrong.

A passing re-run establishes that those assertions passed. It does not establish that the assertions cover every requirement, or that the application is defect-free. That is why the diff, the original failure, and the verification result all matter.

Your Claude subscription, put to work

AI steps and AI repair ran on a plain Claude subscription through a signed-in Claude CLI. No suite rewrite. No API key setup. For the steps, we used:

9l run tests/example.spec.ts --sdk --goal-provider claude

Two login-page goals completed in 5.7 and 7.5 seconds, with two decisions each. Explicit assertions after the goals passed. A goal that asked for a nonexistent button returned unresolved after one decision and the test failed.

The distinction matters: completing an AI goal is not itself proof of the intended outcome. Keep deterministic assertions after the interaction. Offline repair makes no model call; when you choose model assistance, the configured provider receives the inputs needed for its work. Subscription access remains subject to the provider’s own limits.

Check the tests as well as the application

9l assess tests reviewed all 88 files in approximately three seconds and flagged patterns including fixed waits and conditional skips. Its static count was 775 tests, distinct from the suite’s approximate runtime inventory. Static assessment is advisory; it does not execute those tests or establish coverage.

In an earlier trial on a real fix in the same repository, 9l confirm correctly distinguished confirmed, fix-ineffective, and regressed outcomes. It compares a reproduction on unfixed and fixed revisions. A confirmed result shows that the spec distinguishes those revisions; reviewing the assertion still determines whether it checks the reported defect.

Faster recovery starts with keeping your Playwright investment

We put 9lives alongside YC-backed TesterArmy’s open-source e2e. For this trial, the difference that mattered was the workflow: 9lives worked with our existing Playwright suite, configuration, and assertions; e2e uses its own runner and test format.

For the 0.19.0 documentation we inspected, test.fail() and globalSetup were listed as unsupported in the Playwright migration guide. The provider documentation described Claude API-key access and subscription options for other providers, but no Claude subscription sign-in. These are dated documentation observations; check current support before choosing.

We did not configure a provider for e2e, run its AI steps or AI repair, or send application data to TesterArmy. We cannot claim that its healing failed or that 9lives is faster at AI interaction. The small deterministic comparison was effectively a tie: four tests took 1.5 seconds in e2e and 1.7 seconds in Playwright.

The installation footprint in this setup was a 9.4 MB checksum-pinned 9lives binary versus 46 npm packages totaling approximately 53 MB for e2e and its web engine. That binary comparison excludes the existing Node, Playwright, and browser installation required by 9lives. We observed no 9lives telemetry during the trial; e2e documented default-on usage telemetry, which we disabled with E2E_TELEMETRY_DISABLED=1. Neither observation is a universal privacy guarantee.

Your Playwright coverage is an investment. 9lives lets you keep it: repair supported locator drift locally, inspect the diff, and verify the result before accepting a change.

A local companion to the full QualityMax platform

Use 9lives when the work starts in an existing Playwright repository: run, assess, propose a repair, and inspect the evidence. Use the QualityMax platform when the team needs the broader workflow of discovery, generation, managed execution, and review. The self-healing guide explains the hosted repair workflow.

9lives brings that evidence-first approach to a small local tool. Start with one failing locator, keep the assertion intact, and review the candidate before applying it.

Built without external funding. Built with ambition.

QualityMax and 9lives were built without any external funding. Our ambition is to set the market standard for fast, resilient modern test automation: practical AI that cuts maintenance while protecting the coverage teams rely on. We intend to earn that position through results teams can inspect and reproduce.

9lives puts that ambition in your repository. An open-source runner, your existing Playwright suite, and a repair you can inspect before it becomes part of your code. This field trial is a concrete step toward that goal: one changed button, one saved candidate, and a green re-run.

Your existing tests deserve another life.

Give your next broken locator a comeback. Get the Go runner, follow its setup guide, and try it on one Playwright spec.

Get 9lives at 9lives.run →

Source and documentation · 9lives on the QualityMax homepage