Deliverability Monitoring: Under the Hood
What WarmySender actually measures: inbox-versus-spam placement on warmup mail, the 20-email minimum before any rate appears, the warmup score formula in the open, and the limits nobody can get past.
Most email tools show you "sent" and "delivered." Neither one answers the question you actually care about: did it land in the inbox?
This article explains, honestly, what WarmySender measures, how each number is produced, and - just as important - what nobody in this industry can measure, including us. If you have ever wondered why some platforms confidently show a per-campaign "inbox placement rate" and we do not, this is the answer.
"Delivered" Is Not "Inboxed"
When a campaign email leaves WarmySender, the receiving mail server either accepts it or rejects it. Acceptance is what "delivered" means - on our dashboard and on every other sending platform's dashboard. It is arithmetic: everything you sent, minus everything that bounced.
That is a real and useful signal. It tells you the address exists, the domain resolves, and the receiving server was willing to take the message.
It tells you nothing about folders.
The folder decision happens after acceptance, inside the recipient's mail system, and the result is never reported back to the sender. There is no header, no status and no notification that says "this one went to spam." From the outside, an email that lands in the inbox and an email that lands in spam look identical.
So there is exactly one way to know which folder a message reached: have permission to open the mailbox that received it, and look.
The One Place We Can Look: Warmup Mail
Warmup is the part of the product where both ends of the conversation are mailboxes that people connected on purpose. A warmup email leaves one connected mailbox and arrives in another. We have permission to open the receiving side over IMAP, find that specific message, and record where it is sitting.
The classification is deliberately narrow. Every warmup email we locate is recorded as one of two things: inbox or spam. There is no third bucket that quietly flatters the number.
From those two counts we derive the rates you see:
- Inbox rate = inbox / (inbox + spam)
- Spam rate = spam / (inbox + spam)
Notice what is missing from the denominator. Messages we could not locate, and messages that bounced, are excluded from the placement rate entirely - on purpose.
Why that exclusion matters more than it sounds. If a mailbox is briefly unreachable, or a message is still in flight, or a connection times out, that is a gap in our observation, not a spam hit. Folding "we could not check" into the spam side would produce a number that falls whenever our own visibility falls. A mailbox that is doing fine would look like it is burning down. Ours works the other way: if we could not see it, it does not get a vote.
Alongside placement, every warmup mailbox tracks:
- Bounce rate - bounced / sent
- Reply rate - how much of the warmup conversation gets answered
- Detection rate - (inbox + spam) / sent, meaning the share of what was sent that we were actually able to find and classify
Detection rate is the honesty meter on the other two. A high inbox rate resting on a low detection rate is a rate calculated from a handful of messages, and it should be read that way.
What this number is - and what it is not
This is the part most platforms blur, so we will be blunt.
The placement split measures our own warmup emails travelling between connected mailboxes. It is a reputation signal for the sending account: evidence about how receiving systems currently treat mail from that address and that domain.
It is not a measurement of where your campaign emails land. Your prospects' mailboxes are not ours to open, and never will be. Treat the warmup number as directional evidence about your sender's standing, not as a forecast for one campaign to one list.
The 20-Email Minimum: Why We Show "Building Data"
A new mailbox produces very little to measure in its first days, and a percentage calculated from very little is worse than no percentage at all.
Consider a mailbox with three classified warmup emails. One lands in spam. That reads as a 33% spam rate - a number that would send anyone into a panic. The next morning a fourth email lands in the inbox and the same mailbox reads 25%. Nothing about the sender changed. Only the arithmetic moved.
So we hold the number back. Until a mailbox has at least 20 verified emails - messages we actually located and classified - the interface shows Building data instead of a percentage.
This is a deliberate trade. It means a brand-new mailbox does not hand you a headline figure on day one. It also means that when a figure does appear, it rests on enough observations to be worth reacting to. A metric you can trust late beats a metric you can panic about early.
Day-by-Day History: Direction Beats Any Single Day
For every warmup mailbox we keep a per-day record: how much was sent, how much landed in the inbox, how much landed in spam, how much was rescued out of spam, how much bounced, how much was replied to, and how much stayed undetermined.
Those daily records are what make the two views in warmup analytics possible:
- A 7-day trend, for "is this getting better or worse right now?"
- A 30-day timeline, for "what does normal look like for this mailbox?"
How to read them. Receiving systems are not consistent day to day. Filters get tuned, volumes shift, one peer mailbox has a bad afternoon. A single dip is weather. Three consecutive days sliding the same way is climate, and climate is what deserves your attention.
The most common mistake we see is someone comparing today against yesterday, finding a few points of difference, and changing three things at once. Look at the shape of the line instead. If the 30-day view is flat and today is low, wait a day. If the 7-day view is stepping down, something changed and it is worth finding out what.
The Warmup Score, and Its Actual Formula
Each warmup mailbox gets a score from 0 to 100. We will show you the arithmetic, because a score whose formula is a secret is just a mood.
Score = 100 - (spam rate × 0.5) - (bounce rate × 2) - ((100 - inbox rate) × 0.3)
The result is clamped, so it never falls below 0 or rises above 100.
Three things are being penalised, at three different weights:
- Bounces are penalised hardest (× 2). A bounce means mail was sent to an address that does not accept it. Receiving systems treat repeated invalid sends as a strong negative signal - and it is also the single most fixable problem on this list.
- Missing inbox placement is next (× 0.3 of the shortfall). The further the inbox rate sits below 100, the more this subtracts.
- Spam placement is penalised directly (× 0.5).
One thing is deliberately absent: there is no complaint-rate term, because complaint feedback is not something we receive. We would rather leave a factor out of the formula than invent a value for it and let it move your score.
When there is no data, the score returns nothing at all. It does not default to 100. An untouched mailbox with zero classified emails is not a perfect mailbox - it is an unmeasured one, and the interface says so rather than handing you a green number you did not earn.
Spam Rescue: What It Does, and What It Cannot Undo
When a warmup email is found sitting in the spam folder of a receiving mailbox, we act on it:
- The message is moved to the inbox.
- It is marked as read, after a randomized delay of roughly 5 to 30 seconds.
- It is starred about a quarter of the time.
The randomization is the point. Real people do not open every message the instant it arrives, and they do not flag all of them or none of them. Handling that is uniform and instantaneous looks like handling that is automated.
Now the caveat, which you should hear from us rather than discover later:
Rescue does not erase the spam hit. The initial landing is recorded against the sending mailbox permanently. If a message went to spam and we then moved it, the spam count still went up. It has to - that landing is the honest signal about how the sender is currently perceived, and deleting it would make the metric worthless.
Rescue is also not a feedback loop with the provider. It is an action taken inside a mailbox we have permission to act in. We are not signalling to any provider that a message was misclassified, and any platform telling you it can do that is describing something the providers do not offer.
What rescue genuinely does is change the state of the receiving mailbox - a message that was in spam is now in the inbox, read, and sometimes flagged - which is exactly the engagement pattern warmup exists to build.
Authentication: SPF, DMARC, and the DKIM We Refuse to Guess At
When you connect a mailbox and run its connection test, we check the domain's DNS records:
- MX - is there anywhere for mail to this domain to go?
- SPF - is a sending policy published, and what does it authorise?
- DMARC - is a policy published, and what does it instruct receivers to do with mail that fails?
DKIM is deliberately not checked, and we say so in plain words instead of showing you a red mark.
Here is why. A DKIM public key is not published at a fixed, discoverable name. It lives at a name built from a selector that your provider chooses, and there is no lookup that enumerates the selectors in use under a domain. Without the exact selector there is nothing to fetch. We could guess at common names, get nothing back, and display a failure - but a failed guess is not evidence of a missing record.
That is the governing rule behind every check we show you: a check that could not complete is not a failed check. Every result carries whether it was checked separately from whether it passed. Unknown stays unknown.
It matters more than it sounds. A platform that renders an unverifiable DKIM as a red X sends people off to spend an afternoon fixing a record that was already correct - and teaches them to distrust the checks that were real.
Domain Blocklist Monitoring: Two Zones, Every Two Hours
Domain blocklists are one of the few deliverability signals that are genuinely public, genuinely checkable, and genuinely worth watching continuously.
We check exactly two, and both are domain-based:
- Spamhaus DBL
- SURBL
The check itself is a DNS lookup against each zone. No answer means clean.
We do not check IP blocklists, and we do not claim a list of "50+ blocklists." Most of that longer list is IP-based zones describing sending infrastructure rather than your domain, plus zones with unclear listing criteria and no working delisting path. Two zones that mean something beat fifty that generate noise.
It runs on a schedule, not on a button. The monitoring loop runs every two hours. Within it, each domain is rechecked every 12 hours while it is clean and every 2 hours while it is listed - so a delisting is picked up quickly, instead of leaving a stale warning on your dashboard for half a day.
There is no "check my domain now" control, and that is deliberate. Continuous background monitoring means the answer is already waiting when you look, and nobody has to remember to run anything.
Coverage includes customer sending domains plus platform and verified customer tracking domains - the hostname inside your links matters to filters just as much as the address in your From line.
Bounce Protection That Actually Stops a Campaign
Watching a number is worth far less than acting on one. Campaign bounce protection is where a metric becomes a control.
- Each campaign carries a maximum bounce rate, defaulting to 5%.
- The ceiling only applies after a minimum window of 50 sends, so three bounces among the first four addresses cannot trip it. Early noise is not evidence.
- Cross that ceiling with real volume behind it and the campaign pauses itself, instead of continuing to burn the sender's standing while you are asleep.
- When you clean the list and resume, a fresh baseline starts. Old history does not instantly re-pause a campaign that has already been fixed.
Separately, an address that bounces is only auto-suppressed once more than one account has independently seen the same address bounce. A single misconfigured sender should never be able to poison a shared list for everybody else.
If you want to keep bounces down long before that ceiling becomes relevant, that is what real-time email verification is for: check the addresses before the campaign sends to them, not after.
The Limits, Stated Plainly
Everything above is what we can see. Here is what we cannot, in the same plain language:
- Campaign "delivered" only means the receiving server accepted the message. It is not a placement signal and we will never present it as one.
- Nobody outside the recipient's mailbox can see which folder a message landed in. Not us, not anyone. A per-campaign inbox-placement percentage for cold prospects is an inference, not an observation.
- We do not detect Gmail's inbox tabs. Those categories are labels applied inside the recipient's own account, not folders anyone can read from outside. Our warmup classification is inbox or spam, and we would rather say that than report a tab we did not read.
- We publish no accuracy figure. We have not run a validation study against manual checks, so any percentage printed here would be decoration. A number with no measurement behind it is worth less than no number.
- The warmup placement rate describes the sending account, based on mail between connected mailboxes. It is strong directional evidence about your sender. It is not a prediction for one campaign to one list.
We would rather publish a short list of things we actually measure than a long list of things that merely sound impressive.
How to Use All of This
A practical order of operations for a mailbox you care about:
- Fix authentication first. SPF and DMARC are binary, checkable and entirely within your control. Nothing downstream compensates for missing them.
- Let warmup reach 20 verified emails before reading anything into the placement number. "Building data" means the number is not ready yet, not that something is broken.
- Read the trend, not the day. Use the 30-day view for normal and the 7-day view for direction. React to a slope, not a dip.
- Watch bounces harder than anything else. They carry the heaviest weight in the score for the same reason receiving systems weigh them heavily - and verifying addresses before sending removes most of them.
- Take a blocklist listing seriously and immediately. It is rechecked every two hours for a reason. Both zones we watch publish their own delisting process, and a listing that stays listed keeps costing you.
- Do not read a low detection rate as a good result. If little was classified, the percentage sitting on top of it is thin.
Measurement You Can Trust
It would be easy to build a dashboard full of confident percentages. Inventing a placement number for campaign mail is not technically hard; it is just not true, and a number you cannot trust is worse than a blank space, because you will act on it.
So this is the trade we made. Fewer numbers, each one traceable to something genuinely observed: a warmup email located in a real folder, a DNS record that either resolves or does not, a blocklist zone that either answers or does not, a bounce that either happened or did not. Where we cannot see, we say we cannot see.
Every mailbox you warm up in WarmySender gets all of it - the inbox-versus-spam split, the day-by-day history, the score with its arithmetic in the open, the authentication checks, and the two-hourly blocklist watch. Connect a mailbox, let it gather its first 20 verified emails, and read the trend from there.