Campaigns¶
Generally available
Email Security is generally available. Subscribe to the Email Security extension in your organization to purchase and enable it. Paid usage costs \(1 per protected mailbox per month**, billed daily at **\)1/30 per mailbox-day. Free-tier organizations receive a limited 14-day trial. See security product billing for pricing, trial limits and upgrading.
CLI examples require LimaCharlie CLI 5.7.0 or later, which includes the
mailsec commands. Install or upgrade, then check:
Credential-file examples also require jq.
An attack that reached forty mailboxes is one thing that happened, not forty. A campaign is the cluster of messages the engine attributed to one attack, so it is triaged once and remediated once.
How clustering works¶
Every message contributes cluster keys at ingest time:
| Key | How it is built |
|---|---|
| Normalized subject | Lowercased, reply/forward prefixes stripped (Re:, Fwd:, AW:, SV: …), then the parts an attacker varies per recipient collapsed — runs of two or more digits, UUIDs and long hex tokens all fold to a placeholder. "Invoice 88213" and "Invoice 88214" are one campaign, and a key that distinguished them would defeat the feature |
| Link root domains | The sorted, deduplicated set of registrable root domains the message links to, fingerprinted. Root domains rather than URLs, because a phishing kit gives every recipient a unique path or tracking parameter while the domain is what the attacker had to register. Sorting means link order, which varies with templating, cannot split one campaign in two |
| Attachment hashes | The set of attachment SHA-256s |
| Body similarity | A fuzzy hash (TLSH) of the message body, normalized first — see Body similarity |
A message joins an open campaign when at least two keys agree with a candidate. One key is too easy to hit by accident — two unrelated messages with the same generic subject are not a campaign — and requiring three would miss real campaigns that vary one dimension deliberately.
An empty key never matches. Two messages with no subject, or with no links, are not evidence of anything, and treating empty as agreement would cluster the whole tenant into one campaign.
Candidates are considered newest first, bounded, and the first one reaching the threshold wins. Joining an existing campaign always beats seeding a new one.
A campaign takes two messages¶
A campaign is created on the second message, never the first. When a message agrees with an earlier one that is not yet in any campaign, the campaign is created around both of them in one step — the earlier message is the seed, so the campaign's identity and its sample subject come from the first message of the attack, which is what you expect to see when you open it.
A single message is not a campaign. member_count is the spread the campaign
list leads with and the number a campaign-wide quarantine is justified by, so a
cluster of one is deliberately not shown anywhere: not in the campaign list, not
in the Overview's active-campaign count, and not in the campaign detail. Setting
min_members to 1 does not reveal them — the filter can only narrow past the
two-member minimum, never below it.
member_count never goes down. A campaign row lives for 400 days while the
message index it draws on lives for 35, so the count is the historical spread of
the attack rather than a count of messages still in the index. A campaign whose
members have aged out still tells you how big the attack was.
Messages that join after the fact¶
Clustering runs while a message is being ingested. Two copies of one attack that arrive in the same instant therefore each look for a campaign-mate before the other has been written down, find nothing, and are both stored attributed to nothing — and a campaign that forms around a third copy later would leave one of them out.
A retro-join pass closes that, and what it looks at is narrow on purpose.
A campaign is created from exactly one earlier message — the seed. When a message forms or joins a campaign, any other ungrouped message it agreed with is one the engine has just proved belongs with that campaign and passed over. Those, and only those, are queued to be asked again.
| What is queued | A message another message agreed with, on at least two cluster keys, and did not take into the campaign it made. Not "anything that could conceivably cluster" — most mail satisfies the two-key rule and is in no campaign, and queueing all of it would leave the messages that matter waiting behind it |
| How long it stays queued | 24 hours from delivery. Past that it is dropped: mail arriving later still finds it through the ordinary lookup, since it remains a candidate |
| How often it is asked | At most once every ten minutes, and only while it is still ungrouped and inside that window |
| What you see | The message's campaign_id and cluster_reason fill in, and a campaign-wide sweep from that point reaches it — the sweep and the member list read the members directly, so neither ever misses a late joiner. member_count catches up within about a minute, because it is folded on its own schedule rather than written by the join. An EMAIL_VERDICT with campaign_joined_late: true is emitted so a rule can respond to the change |
Nothing is queued for an organization whose mail is not clustering, and a message that agreed with nothing is never re-asked — there is no answer waiting to change.
It applies from the day it is enabled
Messages that were already ungrouped before this existed are not revisited: the queue is built by the passes that leave a message out, so it starts empty. In practice the next message of a live attack re-derives the same set, so an ongoing campaign repairs itself; a campaign that finished before the feature was on stays as it was recorded.
Body similarity¶
The first three keys are all things an attacker can randomize. A kit that gives every recipient a unique subject line and a unique tracking path on every link defeats all three and sends one attack that looks like forty unrelated messages.
What such a kit cannot randomize is the pitch. The body has to read the same to every victim, because the body is the attack. So the fourth key is a fuzzy hash of the body — one that gets closer the more two texts have in common, rather than changing completely when one character does.
The body is normalized first¶
Hashing the raw body would key on exactly the bytes a kit varies. Before hashing, the body is reduced to what it actually says, in this order:
- Visible text. The text a reader sees is used: the displayed text of the HTML part when there is one, the plain part otherwise. Text hidden from the reader (zero-height divs, white-on-white) is dropped, because planting per-recipient noise there is the oldest way to defeat a similarity hash.
- Case and invisible characters. Everything is lowercased, and control characters and zero-width or other invisible formatting characters are removed — the same evasion as hidden text, one layer down.
- Links collapse to their scheme and host, in place. The path, query and
fragment are where the victim's identifier lives; the host is what the attacker
had to register, so the whole host is kept, subdomains included. A
credential-in-URL host such as
www.example.com@evil.exampleis kept too, with the@read as the word "at", because the deception is the point of that link. - Quoted replies and forwarded threads collapse from the start of the quote to the end of the message, so a kit that top-posts one pitch above each victim's own stolen thread is one message, not forty. Only unambiguous markers count: a forwarded-message or original-message separator, a client attribution line ("On …, name wrote:"), or a block of quoted lines running to the end.
- Email addresses collapse to a placeholder — the greeting and the "this was sent to …" footer are otherwise a per-recipient signature.
- Signature blocks collapse: the sign-off line ("Kind regards,") and the short block after it, which is where a templated message puts its rotating persona or case worker. Only a short block at the end of the message qualifies.
- Greetings collapse, salutation and all, so "Dear Riley," and "Good morning Morgan" are the same sentence.
- Every recipient's name collapses wherever it appears in the prose — the display names and addresses on the To and Cc lines as well as the mailbox's own — so a kit that writes "this notice was sent to Riley only" does not sign each copy. Names shorter than three characters and names that are also ordinary words are left alone.
- Hex, opaque and short letter-and-digit tokens collapse to a placeholder:
long hex strings and UUIDs, any 20-character run of letters, digits,
_and-, and any run of six or more characters that has at least four digits beside a letter. That covers invoice numbers, ticket ids, and unsubscribe and tracking tokens. - Numbers collapse to a single placeholder each, however they are punctuated — an amount, an account fragment, a date, a time. This runs after step 9 on purpose: collapsing digits first would hide the mixed letter-and-digit references that step 9 exists to catch.
- Whitespace collapses last.
Steps 8 to 10 never touch the host of a link. That protection is deliberate: the
host is made of exactly the shapes those steps collapse (an IP address is all
digits, office365 is a letter-and-digit run), and without the protection two
unrelated messages linking to different hosts could normalize to the same text.
How much this matters, measured rather than claimed: one phishing pitch templated over eight recipients — name, greeting, amount, account fragment, tracking token and signature all varying per copy — is 37 to 219 apart before normalization and 0 apart after it, across all three shapes such a kit takes (addressed to each victim, collected from a shared mailbox, or top-posted above a stolen thread). That is 84 pairs in all, and every one of them is 0 apart once normalized. The default join distance is 30, so without the normalization this key would not work at the length of an ordinary email.
The threshold¶
Two bodies count as the same body at a distance of 30 or less, which is a
policy knob (clustering — see the Policy Reference).
30 is measured. Across a corpus of several hundred pieces of ordinary business mail — newsletters, invoices, calendar invites, internal notices — the closest pair of unrelated messages is 40 apart, and at 30 (or at the ceiling of 35) the body key produces zero agreements across every pair. The policy ceiling is 35, below that closest pair on purpose: a setting above it is one you cannot have measured, and what it buys is a campaign-wide quarantine reaching mail that was never part of the attack. That margin is re-measured whenever the normalization changes, and it is the number a change has to justify itself against.
The normalization is what groups mail — the threshold is not a tolerance dial
At the length of ordinary email this hash is very sensitive, and the numbers above are what the normalization achieves, not what the threshold tolerates. Two bodies whose normalized text is byte-identical score 0; a single per-recipient word the normalization could not identify — a company, a city, a word of prose — costs a median of 20 to 30 points but exceeds 100 in the worst 5% of cases, measured across a whole corpus rather than on one pair.
So a residual per-copy word is close to a coin flip: in our measurement about 56% of such pairs land within the default distance of 30. Moving the distance to the ceiling of 35 raises that to about 71% while spending most of the margin against unrelated mail. Raising the threshold is not the lever it looks like. What the key reliably buys is the mass case — one pitch, randomized subjects, links, names, amounts, references and signatures — which is the dominant real shape. A kit that rewrites a word of prose per victim may still not group, and that failure is entirely on the recall side: it can never merge unrelated mail.
It still takes two keys¶
Body similarity does not join a campaign on its own. It is the strongest of the four keys and it is still one key, and a 40-point margin is a margin rather than a wall — two form letters from different vendors can read alike. What the key changes is not the threshold but how often it is reachable: a message whose subject and links were randomized now has a second key to agree on.
Why a message did or did not group¶
Every message records the keys that agreed when it joined, and the drawer and the API return them:
cluster_reason is the answer to the question a campaign-wide quarantine
provokes: why was my mail in that group?
body_tlsh is the digest itself. It is shareable — it is the standard TLSH form,
so it pastes into other tools — and the "similar messages" view ranges on it:
Every candidate carries matched_keys and, when both messages have a digest, a
body_distance next to the threshold your policy set. That view returns
candidates, not campaign members: the "two keys agree" decision belongs to the
clustering engine, and a list presented as "similar" without saying how similar
would be an unexplainable claim.
The window is wider for bodies¶
The other three keys look back 72 hours, the open-campaign window. Body similarity looks back seven days, because the attack this key exists to catch is also the one that trickles: a kit sending a hundred distinct-looking messages over five days is one campaign, and a 72-hour horizon would split it into two nobody would ever connect. A body-similar message arriving days later reopens the campaign, because a campaign still receiving mail is still live.
Messages with no usable body¶
A body under 50 bytes of normalized text — "ok, thanks" — or one with too little
variety to characterize produces no digest, and body_tlsh is absent. That is
a fact about the message, not a missing value: an empty key never matches, which
is what stops every short message clustering with every other one.
Campaigns close after 72 hours of silence. Only open campaigns absorb new members, so an attacker re-using a subject three months later starts a new campaign rather than resurrecting an old one.
A campaign's verdict is the strongest verdict among its members.
Working campaigns¶
limacharlie mailsec campaign list --min-members 3 --oid $OID
limacharlie mailsec campaign get <campaign_id> --oid $OID
Filters: state (open, closed), verdict, min_members, since/until,
all repeatable where it makes sense and keyset-paginated like every other list.
min_members narrows past the two-member minimum every campaign already meets;
values below 2 have no effect.
The detail view gives the campaign's span, its membership, its verdict and the
keys that bound its messages together. From a message, campaign_id is on the
index row; from a campaign, message list --campaign-id gives the members.
In the console, Campaigns has the list, the detail, and the sweep controls.
Sweeping a campaign¶
A campaign-wide action applies a per-message action to every member:
quarantine_message, trash_message or restore_message. Three protections
apply, and none of them is optional.
1. Preview, then confirm¶
A sweep with no confirmation token changes nothing. It returns exactly what
it would do — the member ids, the distinct mailboxes affected, and the counts —
plus a confirm token.
# Preview. Nothing is touched.
limacharlie mailsec campaign action <campaign_id> \
--action quarantine_message --oid $OID --output yaml
preview: true
campaign_id: <campaign_id>
action: quarantine_message
member_count: 38
mailbox_count: 31
confirm: <token>
mailbox_count is the number that matters. "38 messages" and "31 people's
inboxes" feel very different, and only one of them is the real blast radius.
# Execute exactly the set the preview described.
limacharlie mailsec campaign action <campaign_id> \
--action quarantine_message --confirm "<token>" --oid $OID
The token is derived from the member set, not from the campaign id
Passing the campaign id as --confirm is refused. The token is a
function of the exact members the preview showed you, so a campaign that
absorbed new messages while you were reading the preview fails the
confirmation rather than sweeping a set nobody approved. Re-run the preview
and confirm the current set.
2. A cap¶
A sweep refuses outright above 500 members rather than asking again. An operator confirming a four-thousand-message sweep from a dialog has not really consented to four thousand mailboxes changing; that needs a person deciding.
3. The executor still decides¶
Every member goes through the same remediation path as a single-message action,
so alert_only, the audit row and idempotency all apply unchanged. A sweep is
many ordinary actions, never a bulk write that skips them.
That includes alert-only mode, which withholds a sweep you run as much as one an
automation would. When any member comes back alert_only, the result carries
force_required: true; re-send the execute with force (--force on the CLI,
beside the same --confirm token — previews do not take it) to perform it. See
Forcing an action in alert-only mode.
Members are routed per message, so a campaign that spans a Microsoft 365 connection and a Google Workspace connection sweeps correctly across both.
The result¶
campaign_id: <campaign_id>
action: quarantine_message
attempted: 38
succeeded: 36
skipped: 4
alert_only: 0
force_required: false
failed:
<msg_uuid>: "<provider error>"
action_id: <action_id>
skipped counts members that were already where the action wanted them. It is
a subset of succeeded — the campaign is where you asked, and those members
cost no provider write — so a re-run of a sweep reads succeeded: 38, skipped: 38
rather than looking identical to the run that really moved 38 messages.
A sweep does not abort on the first error. Stopping halfway leaves a campaign half-remediated, which is the worst of both states: the attacker still has reach and the operator believes it is handled. Every member is attempted and every failure is named.
action_id is the sweep's own audit row — see below.
Say why¶
reason is your justification for the campaign-wide action, recorded against
your authenticated identity. It is stored on every member's audit row and on
the sweep's own record, so an analyst asking "why was my message
quarantined" gets the answer inline from the message, without having to find the
sweep it came from.
limacharlie mailsec campaign action <campaign_id> \
--action quarantine_message --confirm "<token>" \
--reason "INC-4471: reporter-confirmed credential harvest" --oid $OID
It is optional — an automation has no sentence to type — and bounded at 1024 characters. An over-long reason is refused, not truncated: a clipped justification is a corrupted audit record. The bound applies to the preview too, so you learn about it before the dialog asks you to confirm.
The reason is deliberately not part of the confirmation token. Rewording your justification after reading the preview does not invalidate it.
The sweep's own record¶
A sweep writes one audit row for itself, beside the one row per member. Its id
comes back as action_id, and it reads like any other action — though its
action is the campaign-level name (quarantine_campaign), with the
per-message action inside the request:
It carries who asked, when, why, and the counts — "quarantined 412 of 418, 6 failed". The campaign's own action history lists the sweep as one row — who asked and how it ended — and this is the read you expand it with: the counts and the justification live on the row's request, which the history strip does not select.
A second sweep of the same campaign upserts each member's row, so an already-swept
member's inline reason and actor become the second operator's. The first
operator's justification survives on their own sweep record, which is the other
reason each sweep writes one.
Repeating a sweep, and asking for a second one on purpose¶
Re-running a sweep is idempotent per message and action: clicking twice does not
produce two audit rows claiming two quarantines, and the provider re-checks each
message's placement, so a member that is already where the action wanted it comes
back skipped. This is the supported repair for a sweep that partly failed —
re-run the same action and the members that failed are attempted again.
The same repair covers a sweep that was interrupted. A collector that begins
shutting down during a sweep (a deploy, a rebalance) stops between two members —
the member in flight is always finished and recorded — and answers a retryable
error naming how far it got (N of M members were actioned) and the sweep's
action_id; the same happens when the caller's own request ends first. The
sweep's audit row then reads pending rather than ok, so a partial
campaign-wide action can never look complete. Re-run the same confirmation:
members already actioned come back skipped, the rest are attempted.
When you want the retry recorded separately — a re-run after a provider
outage, where the record of what failed matters as much as the record of the
retry — pass an attempt token:
limacharlie mailsec campaign action "$CAMPAIGN" --oid $OID \
--action quarantine_message --confirm "<token>" \
--reason "re-running after the provider outage" \
--attempt after-the-outage
Or straight against the API:
curl -X POST "https://api.limacharlie.io/v1/mailsec/$OID/campaigns/$CAMPAIGN/actions" \
-H "Authorization: bearer $JWT" -H "Content-Type: application/json" \
-d '{"action":"quarantine_message","confirm":"<token>",
"reason":"re-running after the provider outage","attempt":"after-the-outage"}'
Any new value mints a new audit row per member and a new record for the sweep, so
the retry lands beside the attempt it retried rather than over it. Repeating
the same attempt collapses onto the same rows, which is what makes a lost
response safe to re-send. attempt is not part of the confirmation token either,
so a token minted by a preview stays valid when you decide to record one.
It is an opaque handle, not prose: it is bounded at 128 characters and refused rather than truncated, because it is recorded on every member's audit row and a clipped idempotency token is a different token.
Permissions¶
Previewing needs mailsec.act, the same as executing — a preview reaches the
collector and enumerates members. Reading campaigns needs only mailsec.get.