cadohq
Start free trial
Validation

Proving it on a live pipeline.

Every benchmark on this site is industry research, attributed to whoever measured it. The validation program replaces borrowed proof with cadohq's own: the reply and meeting lift that comes from scoring a cadence before it ships, measured directly on a live sales pipeline.

The output

One defensible sentence, stated with its sample size and window:

Scored cadences booked N percent more meetings than unscored cadences, across M accounts over a defined window.

Until that sentence is sound, the public proof stays on benchmarks. The moment it is, it leads and the borrowed figures move to support.

Hypothesis

Cadences that clear the cadohq gate, scored at or above the workspace threshold and revised against the enumerated edits, outperform the same reps' as-drafted, unscored cadences on reply rate and meetings booked, without raising spam complaints.

The variable under test is the score and the gate, not personalization in general and not the research brief alone. Same reps, same period, same lists. The only difference is whether the cadence was scored before it shipped.

Primary design: randomized hold-out

Accounts are randomized into two arms, within each rep and within each week. Randomizing inside a rep removes selling skill as a confound. Randomizing inside a week removes seasonality and list drift. This is the cleanest comparison: scored versus not-scored, same reps, same window.

Control
Cadence is drafted as normal and shipped as-drafted. No score is shown for these accounts during the window.
Treatment
Cadence is drafted as normal, then scored. Below threshold, the rep applies the enumerated edits until it clears, then ships.
Fallback: matched pre/post

If account volume is too thin for a clean split, compare a 6 to 8 week as-drafted baseline against a 6 to 8 week treatment period after cadohq is switched on, matched on account tier, persona seniority, channel mix, and rep. This is weaker: anything else that changed between the two periods confounds the result. It is the second choice, used only when the randomized design is not feasible.

Metrics
Primary
  • Meeting-booked rate. Meetings booked per account worked. The metric closest to revenue, and the one cadohq headlines.
  • Reply rate. Replies per unique contact reached, deduplicated.
Secondary
  • Positive reply rate, classified consistently across both arms.
  • Composite score at send. Confirms the treatment arm actually shipped higher-scoring cadences. If scores do not differ, the test is invalid.
  • Touches to first reply. Tests whether scored cadences convert in fewer touches.
Guardrail
  • Spam complaint and bounce rate. Scored cadences must not raise complaints. Tracked against the Gmail and Yahoo 0.1 percent danger zone and 0.3 percent ceiling.
Volume and power

Detecting a roughly 3 point lift in reply rate, for example 6 percent to 9 percent, needs on the order of 350 to 500 contacts per arm. Below that, the result is directional rather than significant, and is reported that way. Meeting-booked rate moves slower and needs a longer window, so reply rate reaches significance first. When volume cannot reach those counts in a reasonable window, the test runs as an ongoing internal benchmark with the current sample size always stated.

Attribution discipline
The borrowed benchmarks, named

Until the headline sentence is earned, these are the numbers we borrow, stated with their owners. Each one is somebody's measurement, not an industry fact, and we treat it exactly the way we ask reps to treat a proof point: cited or not used.

What we publish

Two artifacts come out of this. A single headline sentence with its sample size and window, once significance is reached. And an internal methodology note that any prospect's RevOps team can inspect.

The credibility of the deliverability claim on this site comes from showing the method. The validation does the same.

Status: design. Measurement runs on the Nutmeg Labs pipeline.