See mailshade.org

How we test Chrome email tracker blockers

This is the method Mailshade uses before publishing a comparative performance result. Every extension is tested in a clean Chrome profile against the same versioned ranking corpus, with its version, Chrome build, activation mode, access tier, package hash and executed extension identity recorded. The ranking corpus is independent of Mailshade's generated rule catalog and pairs documented tracker patterns with observed or controlled benign resources. Generated rule-conformance probes are tested separately and can never improve a product's comparative score. Chrome network outcomes determine whether a request was blocked; visible extension UI records detection separately; benign controls expose false positives. A product receives no credit for an untested client, detector-only tier or unavailable feature. Official product pages can verify a vendor claim, version or supported platform, but they cannot substitute for a measured result. Raw machine-readable results, corpus provenance, hashes and enough protocol detail to repeat the run must be public before a page may label a result verified or name a winner.

Scope and eligibility

The benchmark covers Chrome extensions whose tested tier provides active receiving-side tracker blocking in webmail. A direct multi-webmail candidate must demonstrate that behavior on at least three separately named webmail families; Outlook, Hotmail and Microsoft 365 count as one Microsoft family rather than three clients. A store description records claimed coverage, not measured coverage. A detector-only free tier remains a reference unless its blocking tier is separately acquired, disclosed and tested.

Freeze the environment

  1. Record the UTC test date, Chrome version, operating system and extension version.
  2. Acquire the exact current-public package, recompute its SHA-256, extract it in the harness and bind the loaded runtime identity to that package.
  3. Use a fresh Chrome profile for every product. Apply a hash-bound, versioned setup driver and record whether protection was the default or configured state and whether that comparable state was free or paid.
  4. Record any login, permission prompt, paywall, unavailable client or setup failure instead of treating it as a pass or fail.

Separate ranking evidence from rule conformance

The locked ranking corpus uses independently documented tracker endpoint patterns and observed or controlled benign resources. Rows retain source and family provenance and are macro-averaged by family so aliases from one vendor cannot dominate the result. Mailshade's generated catalog is exercised by a separate source-derived conformance corpus. Because those probes are produced from the same catalog as its rules, they are explicitly ineligible for ranking. Fixture IDs and expected outcomes are versioned before they can affect a later score.

Measure observable outcomes

  • Tracker blocking efficacy: a ready Chrome network monitor observes the paired control but not the planted tracker request.
  • Benign precision: a benign request reaches the monitor and the calibrated extension UI does not flag it.
  • Multi-webmail coverage: real-client evidence counts only when the locked client fixture, tracker request and benign control all run successfully.
  • Link protection: a trusted browser gesture must not leak the original tracking URL; an explicit warning or verified direct destination is recorded separately.
  • Privacy and permissions: the exact manifest, granted origins and sanitized page, extension and service-worker network capture are reviewed.
  • Local evidence and reporting: per-message evidence, retained local history and an inspectable or exportable record are captured from the tested build.

Gmail-proxy behavior is its own locked stratum: a known tracker endpoint embedded in the proxy URL must be blocked, while an ordinary proxy image and a benign sibling path on the same vendor host must both load. The result proves receiver-side request behavior in the tested Chrome page; it does not claim visibility into an earlier Gmail server-side cache fetch.

Hard failures and scoring

An automatically leaked known tracker URL, a blocked or flagged benign control, or undisclosed product-controlled egress of message, identity or secret data is a hard failure for that product. Diagnostic scores remain visible, but a product with a hard failure cannot be named the category winner. Missing evidence is unassessed, never silently converted into a pass.

Publish before concluding

Every comparison page links its sources, exact tested cohort, version record, methodology hash, result status and conflict disclosure. Pending means the shared run has not completed. Inconclusive means a run exists but missing comparable evidence prevents a defensible ranking. Verified requires the test date, a local gate-checked public raw-results artifact and its SHA-256. A winner additionally requires complete rubric evidence, current-public packages for every direct candidate and one uniquely highest eligible score. Corrections create a new dated record rather than silently rewriting history.

FAQ

Why are official product pages not enough to rank blockers?

They are primary sources for versions and vendor-stated features, but each vendor describes its own product. Blocking efficacy and false positives require the same observable test for every product.

What counts as a blocked tracking pixel?

The planted tracking endpoint must receive no request during the defined observation window. An icon or DOM change by itself counts as detection, not proof of network blocking.

How do you handle a product that cannot be installed or configured?

The run records the exact setup blocker and the result stays inconclusive for the affected scope. The product is not assigned a zero and another product is not awarded an automatic win.

Can Mailshade be declared the winner before raw results are published?

No. Mailshade pages are published by its developer, so the conflict is disclosed. A winner claim requires a dated completed run, pinned versions and public raw results under this method.

How are webmail clients counted?

A webmail family counts only when active blocking is exercised on its locked fixture with tracker and benign network evidence. Outlook, Hotmail and Microsoft 365 count once as the Microsoft family; marketing claims and detector-only tiers do not become benchmark passes.

Can Mailshade's own generated tracker rules determine the winner?

No. Generated rule probes verify that the packaged catalog behaves as authored, but they are circular as comparative evidence and are permanently excluded from ranking. Only the independently sourced ranking corpus contributes to comparative efficacy and precision.