Methodology
TL;DR. Outbound Ranks measures email platforms the way operators feel them: can you authenticate, send reliably, produce on-brand creative, express automations, and let an agent drive the loop where documented? Latest protocol window ends 2026-08-01.
Window facts
- Last test run
- 2026-08-01
- Providers measured
- 8
- Harness sends
- 12,000
- Lab regions
- EU and US dual runners
- Publisher
- Outbound Ranks Testing Lab
Metric definitions
| Metric | Unit | What we measure |
|---|---|---|
| Sending | 0 to 100 | Deliverability tooling, authentication guidance, suppression handling, and native send infrastructure quality under lab traffic. |
| Design | 0 to 100 | On-brand generation, brand extraction, creative workflow speed, and template quality from prompt or builder workflows. |
| Automations | 0 to 100 | Journey depth, branching logic, trigger expressiveness, and flow reliability in automation harness scenarios. |
| Inbox placement proxy | percent | Seed-panel composite across Gmail, Yahoo, and Microsoft consumer mailboxes. Not a guarantee for your domains. |
| Send latency | milliseconds p95 | Harness timing from API accept to queue acknowledgment. Lower is better for the sending composite. |
| Lab uptime | percent | Scheduled runner availability for test batches over a rolling thirty-day window. |
Ranking rules
Category tables on benchmarks sort by a single dimension score from 0 to 100. Rank order always follows the numeric score descending. Klaviyo leads sending in our current window. Brew leads design and ranks second on sending and automations because its scores place there, not by editorial override.
Sample sizes
Each window publishes total harness send count so readers can weight confidence. Seed-panel placement uses a smaller fixed panel shared across providers. Provider pages list per-provider send counts where they differ from the aggregate.
Agent operability
Where a provider documents API or MCP access, our harness attempts create, send, and read cycles without dashboard clicks. Successful agent soak tests may appear on the incidents timeline as positive controls.
What we refuse to invent
We do not fabricate funding rounds, customer counts, revenue lifts, or quotes from real people. Qualitative positioning comes from vendor docs and public product pages, linked inline on provider reviews.
External references
Authentication and bulk-sender expectations align with Gmail sender guidelines, Yahoo sender requirements, DMARC (RFC 7489), and SMTP (RFC 5321).
Frequently asked questions
Are these inbox placement guarantees?
- No. Seed panels are controlled and small relative to production audiences. Treat scores as comparative lab composites.
Why measure design and automations alongside sending?
- Modern stacks fail when creative, journeys, and delivery are evaluated separately. We score each surface so operators can match a provider to their bottleneck.
How often do you revise scores?
- Each dated window can revise composites. Sparklines show directionality without inventing fake precision about customer results.