Cold email teams still spend a lot of energy on surface tricks.
Add the first name. Mention the company once. Rewrite the opener with an AI model so it sounds more human. Remove the obviously spammy words. Vary the template enough that each message looks a little different.
That approach made questionable sense even before mailbox-provider filters got better at reading context.
Now it makes less sense.
Google has not published a sender-facing formula for an "AI cold email score," and other large mailbox providers do not publish their exact classifier internals either. But they have published enough guidance to show the direction of travel: they care about authentication, clear sender identity, non-misleading content, understandable links, user choice, low spam rates, and recipient expectations. Better AI just makes those signals easier to combine.
If the broader filtering shift needs context first, read Gmail's Gemini AI spam filtering. This post is narrower: why shallow personalization stops helping, and what providers are more likely to evaluate instead.
Shallow personalization fails because mailbox providers do not need to ask whether a message contains a recipient's first name. They need to ask whether the message behaves like wanted mail from a trustworthy sender.
For cold outreach, that usually means the important signals are more like these:
In other words, AI-era filters likely reduce the value of cosmetic uniqueness and increase the value of program-level coherence.
No major mailbox provider publishes a complete scoring recipe for cold outreach. The safest interpretation comes from their public sender rules: stronger classifiers make it easier to connect content, identity, links, recipient reaction, and sender history into one trust judgment.
Many cold-email programs call a message personalized when only the top layer changed:
Hi Sarah instead of Hi thereThat is not relationship evidence. It is template decoration.
From a mailbox provider's point of view, the deeper pattern can still look exactly like bulk solicitation:
If a classifier is better at recognizing those repeated structures, the added first-name token does not change much. It only makes the unwanted message sound slightly more tailored.
That matters because Google's public guidance already pushes senders away from deceptive or confusing framing. In Email sender guidelines, Google says message headers and content should be accurate and not misleading, links should be visible and easy to understand, and sender information should be clear and visible. In Top 10 Gmail sender issues, Google also warns against misleading subject lines and display names that imply a reply or thread that never existed.
The obvious implication is that filters are not only looking for bad words. They are looking for misfit.
Even though providers do not disclose their full models, the sender guidance they do publish is revealing.
Google explicitly emphasizes:
5322 formattingThat list is useful because it reveals what providers can reasonably score at scale:
Cold-email operators sometimes read AI filtering news and assume the lesson is "make the text more human." The public guidance points somewhere else. The lesson is closer to: make the whole mail program more trustworthy.
The sections below are not a leaked provider algorithm. They are the operational signals most consistent with the sender requirements mailbox providers keep publishing.
This is the big one.
A message can be beautifully written and still unwanted.
Mailbox providers do not need a philosophical answer to whether cold email is "legitimate outreach." They only need to observe whether recipients seem to welcome it. Public sender guidance keeps circling this same principle: send to people who asked for the mail, make unsubscribing easy, and keep complaint rates low.
That is why shallow personalization is weak evidence. Mentioning a recipient's city or company does not prove the recipient expected contact. It may even reinforce the opposite impression if the message still feels mass-produced.
Operationally, the likely higher-value signals are:
For the reputation side of that, Email sender reputation explained is the companion reference.
Providers likely care a lot about whether the visible identity and the real sending context match.
Examples of bad coherence:
From: identity names one brand, but the links land on another domainGoogle's published rules are unusually direct here. Headers and content should not be misleading. Sender information should be clear. Links should be understandable. Display names should identify the sender and should not imply fake continuity.
That strongly suggests a classifier can reward senders whose identity signals line up and penalize streams where the presentation keeps trying to smuggle marketing through a personal or transactional disguise.
If this failure pattern sounds familiar, Gmail display name mistakes that hurt deliverability goes deeper on the identity side.
A common cold-email belief is that enough rewriting defeats similarity detection.
That is a risky assumption.
A modern filtering stack does not need exact text matches to recognize that many messages belong to the same campaign. Similar subject intent, same sender infrastructure, same destination domains, same reply routing, same volume shape, same CTA, and same audience pattern can all reveal the shared origin.
So while AI can spin ten slightly different openings, the classifier may still see one repeated solicitation pattern.
This is one reason "write every message uniquely" is often the wrong optimization target. The more important question is whether the campaign still looks like a large-scale persuasion stream sent to people with weak prior expectation.
The clicked destination is part of the message, even if it is not in the body copy.
Providers likely score whether:
This follows naturally from Google's guidance that links should be visible and easy to understand. It also fits the broader reality that many bad campaigns use technically clean messages whose real problems only become obvious when the destination, identity, and call to action are considered together.
Mailbox providers have something much stronger than copy analysis: user reaction.
If recipients repeatedly mark the traffic as spam, delete it quickly, ignore it, or use unsubscribe paths as an escape hatch, that outcome says more than a clever opener ever will.
This is why teams that obsess over wording but ignore complaints usually lose. Better language may improve reply rate a little at the margin. It does not cancel out a stream that recipients broadly treat as unwanted.
Google's sender documentation keeps putting spam rate at the center for a reason. It says to keep spam rates below 0.1%, avoid ever reaching 0.3%, and use Postmaster Tools compliance dashboards to monitor whether Gmail sees the domain as well run.
That does not prove every provider uses the exact same threshold. It does show what kind of real-world outcome they care about.
AI filters do not replace the basics. They sit on top of them.
Mail from a new or poorly prepared domain, mixed-purpose stream, unstable warm-up, or sloppy authentication path starts from a weaker trust position. Then the content layer gets judged on top of that.
For cold outreach, the risky patterns are familiar:
Those issues are exactly why nearby posts like New bulk-sending domains after January 2024, New domain vs new subdomain vs new IP, and One-click unsubscribe in 2026 matter so much.
The practical shift is uncomfortable for teams that rely on volume plus personalization veneer.
If mailbox providers are getting better at scoring intent, fit, and recipient response, then the classic cold-email stack becomes less defensible:
The usual answer is to spend even more time on prompt engineering.
That is probably the wrong answer.
The stronger answer is to reduce the number of contradictions in the program.
If a team insists on sending outreach mail, the safer posture is not "make it sound personal." It is "make it behave honestly and predictably."
Do not imply a prior relationship if there is none.
Avoid fake Re: or Fwd: framing, fake thread continuity, or body copy that pretends a manual one-to-one conversation is already underway. Google's own display-name and header guidance points directly away from that pattern.
If a segment is weak, rewriting the first paragraph rarely fixes the underlying unwantedness problem.
Tighter targeting usually beats smarter wording because it improves the signal that matters most: whether the mail actually makes sense to receive.
The sender name, domain, signature, and landing page should tell one coherent story. If the recipient has to work to understand who is asking for attention, the trust score is already moving in the wrong direction.
Not every outreach stream is classified the same way as a standard newsletter, and Google's one-click unsubscribe requirement is specifically for marketing and subscribed messages. But from a deliverability standpoint, making people fight to stop the mail is a terrible idea. Friction increases complaint risk, and complaint risk is exactly the signal providers already tell senders to watch.
Do not let outreach, lifecycle mail, receipts, alerts, and newsletters share the same reputation story unless they truly belong together.
Consistent stream separation makes it easier for both recipients and providers to understand what each mail flow is supposed to be.
A cold-email program can look successful in a sales dashboard while quietly degrading domain trust.
Reply rate is not a mailbox-provider trust metric. Complaint rate, reputation, and compliance are much closer to the real scoring surface.
Cold email is getting squeezed not because mailbox providers suddenly discovered a magical AI rule about personalization, but because better classifiers can judge the whole pattern more effectively.
Shallow personalization fails when the rest of the message still signals bulk solicitation, weak recipient expectation, mismatched identity, or poor user outcomes. A first name, an AI-generated compliment, or a spun opener does not outweigh complaints, distrust, confusing links, or a sender history that looks like nuisance mail.
So the important question is no longer "how human does this email sound?"
It is closer to this: does the full sending program look like wanted, honest, well-run mail when a mailbox provider evaluates identity, content, links, recipient reaction, and history together?
That is the standard shallow personalization cannot meet by itself.