How to Check a Published Benchmark’s Numbers Yourself, Step by Step

ost people reading a technology benchmark stop at the summary table, trusting that the averages reported match the underlying data. A benchmark that actually publishes its raw data invites a more useful habit: recomputing the headline numbers directly, rather than taking the summary on faith. The process takes a few concrete steps, and it works the same way for any benchmark structured like this one.

None of these steps require specialized tools or technical training beyond basic spreadsheet skills, which is part of what makes this kind of verification worth doing rather than skipping in favor of simply trusting the published summary at face value.

Start With What Got Published

A verifiable benchmark publishes more than a conclusion. Phrasly’s September 2026 humanizer benchmark provides a separate CSV file for each of the four tools tested, each listing every input text, its Pangram detection scores, its hallucination count, its word count, and a SHA-256 fingerprint of the exact output text that was scored.

Before checking any specific number, it is worth confirming the file actually contains what the methodology section claims it should, the right number of rows for the stated sample size, with the specific fields described in the published methodology present for each one.

This first pass takes only a few minutes but catches a surprising range of problems, a mismatched row count, a missing column, a file that does not actually match what the summary table describes, all of which are worth knowing about before trusting any number calculated from the file.

Recomputing the Headline Averages

The average Pangram human score reported in a summary table is calculated by averaging a specific field, fraction_human multiplied by 100, across every completed output for a given tool. Recomputing this directly from the CSV means pulling that column, confirming the row count matches the number of completed outputs the methodology states, and averaging it independently.

The same logic applies to the other published measures. The share of fully human outputs is the count of rows where fraction_human equals exactly 1, divided by the total completed outputs. The hallucination averages follow the same pattern, calculated only across outputs with a recorded count, as the methodology explicitly notes for cases like WriteHuman’s three declined texts or the handful of outputs across other tools missing a hallucination count.

Doing this once, for a single tool’s single measure, is usually enough to build confidence in the rest of the summary table, since a benchmark careful enough to get one recomputed figure right is far more likely to have gotten the others right as well. Spot-checking a second, unrelated measure afterward, say a hallucination average instead of a detector score, adds further confidence without requiring a full recomputation of every number in the table.

A basic verification pass on a benchmark like this involves:

•          Downloading the per-tool data files rather than relying on the summary table alone

•          Confirming the row counts match the sample sizes and exclusions stated in the methodology

•          Recomputing at least one headline average directly from the raw figures

•          Checking that excluded or declined cases are actually marked and excluded consistently with how the text describes them

Verifying a Specific Output, Not Just the Averages

Where a benchmark withholds the actual output text and publishes only a cryptographic fingerprint instead, verification works a step further. Anyone who later obtains a specific withheld text, through an approved request process or otherwise, can compute its SHA-256 hash and compare that value against the one published in the CSV.

This check takes only a free, widely available hashing tool and a few seconds to run, and it answers a question no amount of trust in the publisher alone could settle, whether the specific text in hand is genuinely the one that produced the specific score listed next to its fingerprint. A mismatch at this stage would be a serious red flag, while a match confirms the chain from input to score to published figure held together end to end.

A matching hash confirms the text received is identical, character for character, to the one originally scored in the Phrasly Benchmark, which closes the verification loop even when the underlying text itself was never made publicly available from the start.

What This Process Actually Protects Against

None of this guards against every possible way a benchmark could mislead, a publisher still controls which texts were selected and how the test itself was designed. What it does protect against is a narrower but still meaningful risk, a published summary that does not actually match its own underlying data, whether through an honest calculation error or something more deliberate.

A benchmark that survives this kind of check, where an outside party can recompute the same averages independently from the published raw data, has cleared a real bar that most industry comparisons never attempt to meet in the first place, since most never publish enough underlying detail for this kind of check to even be possible.

Building this habit pays off well beyond any single benchmark. Once a reader has verified one published dataset firsthand, spotting the difference between a genuinely checkable claim and an unverifiable one becomes far easier the next time a new comparison shows up making a similarly confident claim, often within seconds of glancing at whether any underlying data was published at all.

For more on how AI detection benchmarks are designed and verified, further reading on the Phrasly blog covers the underlying research for anyone building this kind of verification habit into how they evaluate new tools.

Disclaimer: This article is for informational and educational purposes only. The steps and information provided are general guidance for independently checking published benchmark figures and may vary depending on the source, methodology, and type of benchmark. Readers should verify data using the original source and official documentation where available.