Proving it on a live pipeline.
Every benchmark on this site is industry research, attributed to whoever measured it. The validation program replaces borrowed proof with cadohq's own: the reply and meeting lift that comes from scoring a cadence before it ships, measured directly on a live sales pipeline.
One defensible sentence, stated with its sample size and window:
Scored cadences booked N percent more meetings than unscored cadences, across M accounts over a defined window.
Until that sentence is sound, the public proof stays on benchmarks. The moment it is, it leads and the borrowed figures move to support.
Cadences that clear the cadohq gate, scored at or above the workspace threshold and revised against the enumerated edits, outperform the same reps' as-drafted, unscored cadences on reply rate and meetings booked, without raising spam complaints.
The variable under test is the score and the gate, not personalization in general and not the research brief alone. Same reps, same period, same lists. The only difference is whether the cadence was scored before it shipped.
Accounts are randomized into two arms, within each rep and within each week. Randomizing inside a rep removes selling skill as a confound. Randomizing inside a week removes seasonality and list drift. This is the cleanest comparison: scored versus not-scored, same reps, same window.
If account volume is too thin for a clean split, compare a 6 to 8 week as-drafted baseline against a 6 to 8 week treatment period after cadohq is switched on, matched on account tier, persona seniority, channel mix, and rep. This is weaker: anything else that changed between the two periods confounds the result. It is the second choice, used only when the randomized design is not feasible.
- Meeting-booked rate. Meetings booked per account worked. The metric closest to revenue, and the one cadohq headlines.
- Reply rate. Replies per unique contact reached, deduplicated.
- Positive reply rate, classified consistently across both arms.
- Composite score at send. Confirms the treatment arm actually shipped higher-scoring cadences. If scores do not differ, the test is invalid.
- Touches to first reply. Tests whether scored cadences convert in fewer touches.
- Spam complaint and bounce rate. Scored cadences must not raise complaints. Tracked against the Gmail and Yahoo 0.1 percent danger zone and 0.3 percent ceiling.
Detecting a roughly 3 point lift in reply rate, for example 6 percent to 9 percent, needs on the order of 350 to 500 contacts per arm. Below that, the result is directional rather than significant, and is reported that way. Meeting-booked rate moves slower and needs a longer window, so reply rate reaches significance first. When volume cannot reach those counts in a reasonable window, the test runs as an ongoing internal benchmark with the current sample size always stated.
- Lock the metric definitions before the test starts. A reply and a meeting are counted the same way in both arms.
- Pre-register the primary metric and the minimum sample size before unblinding. No fishing for whichever number happened to move.
- Confirm the manipulation: the treatment arm must ship higher composite scores than control. Report the score gap next to the outcome gap.
- Report confidence with the sample size, never a bare percentage. A figure without an n is the kind of claim this product exists to catch.
Until the headline sentence is earned, these are the numbers we borrow, stated with their owners. Each one is somebody's measurement, not an industry fact, and we treat it exactly the way we ask reps to treat a proof point: cited or not used.
- Gmail enforces a 0.30% spam-complaint ceiling, with 0.10% recommended, and has escalated to permanent rejections since November 2025. Google's own sender guidelines, the one primary source on this list.
- Average cold email reply rate across 7.5 million sends in 2025 was 0.45%, falling 20% within the year. Belkins' client campaign data: one agency's book, measured against all sends, not a market census.
- The highest-impact AI application in prospecting was audience filtering, at 356% more closed deals. It was not copy generation. Sopro, State of Prospecting 2026, their campaigns, their analysis.
- Fewer than 10% of 300+ surveyed RevOps leaders report ROI from AI, while roughly 45% plan to expand usage. Default, The State of AI in RevOps. A vendor-run survey, quoted as such.
Two artifacts come out of this. A single headline sentence with its sample size and window, once significance is reached. And an internal methodology note that any prospect's RevOps team can inspect.
The credibility of the deliverability claim on this site comes from showing the method. The validation does the same.