Skip to content

Autonomous retest · for pentest consultancies

Your seniors retest every “fixed” finding by hand. RiftX does it for them, and never guesses.

A fix can look solid and still fail under attack. RiftX reruns each reported finding in a real browser, tries to beat the fix, and returns one evidence backed verdict: Fixed or Not Fixed, sealed for the report. It never clears a vuln it can’t prove; when unsure, it returns Needs Review. 5× cheaper than a retest by hand, and none of it costs a senior hour.

Private beta. No credit card. Set up in one session.

One retest · one verdict
Fixed
as reported

The team shipped a patch and closed the ticket.

RiftX retested the fix
Not Fixed
#VT-2026-0847

Reflected XSS: attribute context bypass

Evidence · dialog intercepted● captured
payload landed in value="…"
alert(document.domain)fired
HMAC-SHA256 · sealedHAR · MP4 · PNG
5× cheaperthan a retest by hand
~6 hrs backper engagement, off your seniors
Never guessesunsure → returns Needs Review

Finding got automated. Retesting didn’t.

AI assisted researchers and autonomous agents now surface vulnerabilities faster than any team can work through them. A reported finding is only a claim until someone proves it again against the live target, and that step is still manual, still senior.

  • 560+

    valid vulnerabilities submitted by autonomous AI agents in a single year.

    HackerOne · 2025
  • #1

    XBOW, an autonomous AI, topped HackerOne's US leaderboard, ranked above every human researcher.

    2025
  • 20

    real vulnerabilities found by Google's Big Sleep with no human intervention.

    Google · 2025
  • 1 in 5

    of curl's 2025 security reports were AI slop, dropping its real bug rate to roughly 1 in 20. Real find or false alarm, the only way to know is to retest it.

    curl · 2025

The bottleneck moved. Finding is cheap now. Proving what’s real, and whether the fix holds, is the work that’s left, and it still lands on your seniors.

Sources: HackerOne 9th Hacker-Powered Security Report (2025); Google "Big Sleep" via TechCrunch (Aug 2025); XBOW HackerOne US leaderboard (2025); curl "Death by a thousand slops," D. Stenberg (Jul 2025).

Every engagement, your best people retest the same findings by hand.

Navigate, inject, screenshot, write it up. Then do it again next engagement. It is structured, repetitive work that requires no human judgment, and it still burns senior pentester hours.

That is not a pentester problem. It is a delivery capacity problem, and it scales with every client you win.

6 hrsper engagement, gone
~6findings, same loop each time
1 houreach, no judgment required

~6 findings/engagement is conservative against Cobalt State of Pentesting 2024 (~9.6/pentest). Retest time from hands on engagement experience.

RiftX isn’t a scanner. It’s an autonomous retester.

You report a finding and its steps-to-reproduce. RiftX retests exactly that, never hunting for new bugs, and reaches its own verdict.

Report X, retest X, nothing else.

Not Fixed

The finding still reproduces. RiftX proves it with a live capture: the fix did not hold.

Fixed

The fix holds under attack. RiftX could not reproduce it, and shows the evidence for why.

How it thinks

Reproducing isn’t the product. Beating the fix is.

Anyone can replay the reporter’s steps. RiftX treats the fix as an adversary: it characterizes the defense, plans bypasses against it, and only then judges. A confirmed bypass overturns “Fixed.”

  1. 01

    Reproduce

    Replays the reporter's steps in a real, isolated browser.

  2. 02

    Observe the defense

    Characterizes what the fix actually does before touching it.

  3. 03

    Plan & execute bypasses

    Maps the fix to techniques worth trying, then tries them.

  4. 04

    Judge

    A read only judge reasons over the raw evidence it captured.

  5. 05

    Verdict

    A confirmed bypass flips “Fixed” to “Not Fixed.”

Retest log● running
00:00reported finding received
00:24goal set: reproduce reflected_xss
01:13fix observed: output HTML encoded
02:41bypass planned: attribute context break
03:58payload executed: dialog intercepted
04:57verdict: Not Fixed · evidence sealed

A replay tells you the finding once existed. Beating the fix tells you whether it still does.

Evidence & Trust

Evidence is part of the verdict, not an afterthought

Every retest produces a sealed bundle. Your pentester reviews evidence, not assertions.

#VT-2026-0847
Not Fixed
Finding
Reflected XSS
Severity
High
Confidence
90%
Target URL
https://app.target.com/search
Parameter
q
Verdict
Not Fixed
Payload
<script>alert(document.domain)</script>
Evidence
HTTP/1.1 200 OK
Content-Type: text/html; charset=utf-8
X-Request-Id: 7f3a...
...
<div class="results">Search: <script>alert(document.domain)</script></div>
Steps to Reproduce
  1. Navigate to https://app.target.com/search
  2. Enter payload <script>alert(document.domain)</script> in the search field
  3. Observe JavaScript execution via browser dialog interception

HMAC-SHA256: a3f2...8e91. Evidence integrity sealed

Every verdict includes

Full HTTP traces

Complete request and response for every step

Screen recording

MP4 of the browser executing the reproduction steps

Reproduction steps

Exact steps followed, ready for your report

Integrity seal

HMAC-SHA256 hash proving evidence was not modified

The auditor

An always on, independent QA service rechecks every verdict. By design it can only downgrade toward Needs Review, never inflate. The failure mode is caution, never a false all clear.

The boundary

Hard safety limits bound every retest’s time, actions, and spend. When it cannot be sure, it returns Needs Review rather than risk clearing a vuln that is still live.

The agent sees real traffic. Here is how that data is held.

AES-256

Encrypted, off host

Evidence lives only in S3 under AES-256. Workers hold no cloud keys and keep nothing locally after upload.

API key scoped

Tenant scoped

Every artifact is keyed to the customer from API key auth, never the request body. Cross tenant reads are rejected.

Per job browser

No state reuse

Each job's browser is wiped and closed before the next. No cookies or sessions carry across jobs or tenants.

HMAC-SHA256

Tamper evident

Every bundle is HMAC-SHA256 sealed over a Merkle root of its files. Any later edit is detectable.

Scope

The consultancy owns and authorizes the target, and the agent is SSRF guarded: it cannot be turned against cloud metadata, loopback, or RiftX’s own infrastructure.

This already runs. Today.

Not a concept deck. A production system you can watch run a real retest end to end.

29

Vulnerability profiles, live

XSS, SQLi, access control, CSRF, security headers, open redirect, clickjacking, session management, TLS, open ports. The web findings consultancies actually retest.

Every verdict, independently rechecked

A separate, always on auditor rechecks each verdict and can only flag it for review, never inflate.

200+

Accuracy cases scored

Regression scenarios with a known correct verdict across the supported vuln classes. Every run is scored against that ground truth.

1

Working product

Submit a finding, watch it retest, read the sealed verdict. A production stack, not a prototype.

Live coverageXSSSQLiAccess controlCSRFSecurity headersOpen redirectClickjackingSession managementTLSOpen ports

5× cheaper, and it frees your best people.

At five engagements a month, that is roughly thirty senior hours handed back, close to four working days your best people spend finding bugs instead of proving old ones again.

$19.99

What RiftX charges to run that same retest. 5× cheaper, and it spends none of your senior hours.

$100

What one manual retest costs in senior pentester time, roughly an hour of repeatable, mechanical work.

605

Minutes of your time per finding. Submit and review; the agent runs the rest unattended.

Why now: agents can finally drive a real browser and reason over the evidence they capture. That was not practical until recently.

$19.99 is RiftX's list price per retest. $100 is a conservative loaded cost for ~1 hour of senior pentester time; senior rates run $120 to $190/hr loaded (Bright Defense, 2026).

Retest is the starting point. Autonomous AppSec is the direction.

Most tooling automates how teams find bugs. Very little automates how they retest the fix. That gap is where RiftX starts.

  1. Now · shipping

    Autonomous retest

    The repetitive retest your team does today, run on its own with evidence you can defend to a client.

  2. Next

    Retested finding workflows

    Inside the tools teams live in: tickets in, sealed verdicts back out.

  3. Later

    Retesting across the SDLC

    Continuous checking, grounded in evidence, wherever findings are produced.

Private Beta

Get your pentesters out of the retesting loop

Your pentesters should be finding vulnerabilities and writing reports, not manually retesting the same XSS for the third time this month.

Limited beta spots. No credit card. Set up in one session.