A fix isn’t fixed until something tries to break it.
RiftX runs the retest unattended. It re-tests each reported finding against the live target, attacks the fix, and returns one sealed verdict with the evidence to check it yourself.
How often the verdict is right, and what it does when it cannot tell.
An unattended verdict is worth what its worst case is worth.
Verdicts that matched a known correct answer
Every scenario runs against one of 43 applications built both vulnerable and patched, so the correct verdict is fixed before the run and the whole set is re-scored on every change. We wrote the suite, so read it as our own bench test rather than an outside audit.
- 298
- matched
- 20
- said Not Fixed against a fix that had held
- 10
- returned Needs Review
- 328
- scenarios
“Fixed” is the answer it works hardest to disprove
It reads what the defense actually does, then attacks it with the bypasses that class of fix is known to miss.
A finding is never reopened without proof
Reopening one takes a captured execution artifact, not an inference about one. A run that cannot produce one comes back for review.
Every verdict is re-derived by a second, independent judge
It reads the same evidence, reaches its own conclusion, and can send a verdict back for review but never upgrade one.
RiftX isn’t a scanner. It’s a retester.
You report a finding and its steps-to-reproduce. RiftX retests exactly that and reaches its own verdict.
Not Fixed
The finding still reproduces, or a bypass got around the fix. RiftX shows the live capture of it happening.
Fixed
The fix held against every attempt RiftX planned against it, and the report names them.
Needs Review
RiftX could not prove it either way and says so. Fixed and Not Fixed clear the same bar; anything short of it lands here.
Reproducing isn’t the product. Beating the fix is.
Anyone can replay the reporter’s steps. RiftX treats the fix as an adversary: it characterizes the defense, plans bypasses against it, and only then judges. A confirmed bypass overturns “Fixed.”
Reproduce
Replays the reporter's steps in a real, isolated browser.
Observe the defense
Characterizes what the fix actually does before touching it.
Plan & execute bypasses
Maps the fix to techniques worth trying, then tries them.
Judge
A separate judge sees the captured evidence and nothing else, then calls it.
Verdict
A confirmed bypass flips “Fixed” to “Not Fixed.”
What stops it marking something Fixed that is still live?
- The risk is not invention, it is coverage: a bypass nobody thought to try. So the claim is never that a fix is unbreakable. It is that the fix survived a named set of attempts, and the report lists each one and whether it held. A bounded result, with the bounds printed.
A replay tells you the finding once existed. Beating the fix tells you whether it still does.
Every engagement, your best people retest the same findings by hand.
Navigate, inject, screenshot, write it up. The loop is identical every engagement, it needs no human judgment, and it runs on your senior bench.
Every client you win makes the queue longer. Practice does not make it shorter.
You cannot fix it by pushing it down the bench. Whoever signs off on a retest has to be senior enough to defend the call, and those are the people who resent the work most.
Retestable findings are a minority of what you report. Criticals and highs get retested for closure. Lows are usually accepted through other validation.
Retest consumes your most expensive capacity. It does not have to.
The only lever most firms have is another head. Nobody has sold you a second one because two years ago a browser agent could not reach a verdict on its own.
One engagement's retest round, six retestable findings. Testing and write-up only.
By hand
An hour testing each finding, half an hour writing it up.
With RiftX
Two minutes to submit each finding, eight to review the evidence and sign it.
Access, credentials and waiting stay with your team either way. Reconciled against a ten person retest team at a mid-size consultancy.
The version to forward
Same work, minus the salary line, with evidence you can hand a client.
Every verdict arrives with the evidence to overturn it
Your reviewer checks the capture, not our word for it.
Can I put a machine’s verdict in a client deliverable?
- You are not meant to take the verdict on faith, and the product is not built as though you would. The run is unattended; the verdict is not. Every verdict ships with the evidence it was derived from, and your retester reads that file and signs the call, which is the judgment they were always paid for, minus the hour of reproducing that used to come first. Your name still goes on the finding, which is exactly why the evidence travels with it.
A second, independent check rechecks every verdict, always on. By design it can only downgrade toward Needs Review, never inflate. The failure mode is caution, never a false all clear.
The same file, open to your questions. Ask why the call landed where it did, what the payload was, or which step settled it. The answer comes back out of this retest’s own artifacts, named so you can open them yourself.
Where your client’s data goes, and where it does not.
Anthropic’s API, under our account
Reasoning runs there and nowhere else, under published commercial terms that do not train on inputs or outputs.
Encrypted, on our own machine in Austria
An encrypted object store we operate, never a hyperscaler bucket. The machines that run retests keep nothing locally.
Kept for a year, then deleted
A retest's evidence is held for 365 days from the run and then deleted.
Tenant scoped
Keyed to your account from the API key, never the request. One client's evidence cannot be read from another's.
No state reuse
Each retest's browser is wiped and closed before the next. No cookies or sessions carry over.
Tamper evident
Every bundle is sealed over its files. Any later edit is detectable.
Am I even allowed to point this at my client’s environment?
- Retest authority is usually implied by the engagement that produced the finding, but implied is not written. We hand you a scope clause bounded to the findings you reported. If the client will not sign, point RiftX at staging.
Take on more retest volume without taking on another head
The queue grows with every client you win, and whoever signs off on a retest has to be senior enough to defend the call to the client. That is the line this moves.
Bring a finding you have already retested by hand. We point RiftX at it live, walk the evidence bundle it produces, and you check the verdict against the answer you already own.
