onionwright · proof

How we measure, and how to check it.

Every claim we make about detection is one number on one detector, measured the same way for stock Chrome and for onionwright. This page is that method written out — including the two checks we do not pass — because a benchmark a buyer can disprove in an afternoon is worth less than no benchmark at all.

The result

One detector, the same launch flags for both browsers, one run each, on a real host. Stock headless Chrome, then onionwright:

Checkstock Chromeonionwright
navigator.webdrivercaughtpasses
BotD verdictcaughtpasses
mainWorldExecutioncaughtcaught
CSP bypasscaughtcaught

2 of 4 clear on onionwright where stock Chrome is caught. The other 2 still catch it, and they are in the table on purpose — the mechanism for each is below.

What each row actually is

Four words in a table is not evidence. Here is the mechanism behind each, the two we pass and the two we fail in the same detail — so you can decide whether our "pass" means what you need it to mean.

  • navigator.webdriver we pass
    Removed by deleting the branch in the engine that sets it, not by overriding the property afterwards. An override leaves its own trace — a getter on the prototype where none should be — which is the second thing a detector checks once the flag itself reads false.
  • BotD verdict we pass
    FingerprintJS's open-source bot detector returns bot=True on stock headless Chrome and bot=False on onionwright, same machine, same run. It aggregates dozens of the signals below into one verdict, so flipping it is a coarse but honest summary of the static layer.
  • mainWorldExecution we fail
    Script injected through CDP runs in the main world and leaves an anonymous stack frame that page code can read. We do not clear this yet. A stealth product that claims a clean pass here is either not testing it or not telling you.
  • CSP bypass we fail
    CDP ignores a page's Content-Security-Policy, so an injected request that CSP should block still goes through — a discrepancy a page can probe. Also open, also tracked in public rather than papered over.

Why only four

There are hundreds of things a page can measure. We publish four because these four we have measured on the same detector, the same way, repeatably — not because they are the only signals that matter. A longer scoreboard would look more impressive and mean less: every row we cannot reproduce on demand is a row you cannot check, and a claim you cannot check is one you should discount.

This is the difference we are actually selling. CloakBrowser's README says it "passes every bot detection test" while its own issue tracker carries DataDome, PerimeterX, Cloudflare and reCAPTCHA failures. "Passes everything" is the claim that reads best and survives contact with a technical reader worst. We would rather show four numbers that hold than forty that do not.

Reproduce it

The point of measuring this way is that you do not have to take our word for it. Point the same open-source detector at stock headless Chrome and at an onionwright session, with matched launch flags, and read the two verdicts. The two rows we pass should flip; the two we fail should not. If your own run disagrees with this table, that is a bug in our claim and we want to hear about it.

An honest evaluation runs against the detectors that matter to your traffic, not ours — which is exactly what a scoped evaluation is for.

Scope an evaluation

Read the rest