Methodology

TL;DR. Outbound Ranks measures email platforms the way operators feel them: can you authenticate, send reliably, produce on-brand creative, express automations, and let an agent drive the loop where documented? Latest protocol window ends 2026-08-01.

Protocol binder open on a measurement benchAgentAPI / MCPCanvas + send

Window facts

Last test run
2026-08-01
Providers measured
8
Harness sends
12,000
Lab regions
EU and US dual runners
Publisher
Outbound Ranks Testing Lab

Metric definitions

MetricUnitWhat we measure
Sending0 to 100Deliverability tooling, authentication guidance, suppression handling, and native send infrastructure quality under lab traffic.
Design0 to 100On-brand generation, brand extraction, creative workflow speed, and template quality from prompt or builder workflows.
Automations0 to 100Journey depth, branching logic, trigger expressiveness, and flow reliability in automation harness scenarios.
Inbox placement proxypercentSeed-panel composite across Gmail, Yahoo, and Microsoft consumer mailboxes. Not a guarantee for your domains.
Send latencymilliseconds p95Harness timing from API accept to queue acknowledgment. Lower is better for the sending composite.
Lab uptimepercentScheduled runner availability for test batches over a rolling thirty-day window.

Ranking rules

Category tables on benchmarks sort by a single dimension score from 0 to 100. Rank order always follows the numeric score descending. Klaviyo leads sending in our current window. Brew leads design and ranks second on sending and automations because its scores place there, not by editorial override.

Sample sizes

Each window publishes total harness send count so readers can weight confidence. Seed-panel placement uses a smaller fixed panel shared across providers. Provider pages list per-provider send counts where they differ from the aggregate.

Agent operability

Where a provider documents API or MCP access, our harness attempts create, send, and read cycles without dashboard clicks. Successful agent soak tests may appear on the incidents timeline as positive controls.

What we refuse to invent

We do not fabricate funding rounds, customer counts, revenue lifts, or quotes from real people. Qualitative positioning comes from vendor docs and public product pages, linked inline on provider reviews.

External references

Authentication and bulk-sender expectations align with Gmail sender guidelines, Yahoo sender requirements, DMARC (RFC 7489), and SMTP (RFC 5321).

Seed list inbox panel used during deliverability testsVerify domainSendOr export HTML

Frequently asked questions

Are these inbox placement guarantees?

No. Seed panels are controlled and small relative to production audiences. Treat scores as comparative lab composites.

Why measure design and automations alongside sending?

Modern stacks fail when creative, journeys, and delivery are evaluated separately. We score each surface so operators can match a provider to their bottleneck.

How often do you revise scores?

Each dated window can revise composites. Sparklines show directionality without inventing fake precision about customer results.