Email overload is a modern disease. Hundreds of emails a day, most of them noise. Newsletters you forgot you subscribed to. Marketing emails disguised as personal messages. Bank statements. Gas bills. The occasional important email buried in the flood.
I tried fixing this with a simple ignore list. It didn’t scale.
This is the story of how I built a multi-layer email filtering system that replaced hardcoded bash patterns with intelligent scoring, security scanning, and adaptive learning.
The Initial Setup (That Didn’t Work)
My first attempt was straightforward: a bash script that checked Gmail via the gog CLI, applied a few regex patterns, and ignored obvious spam.
#!/bin/bash
# Old system: tmp_gmail_check.sh
# Hardcoded ignore patterns
IGNORE_SENDERS="newsletter@|no-reply@|comparethemarket.com|confused.com"
# Fetch unread emails
EMAILS=$(gog gmail messages search "is:unread" --max 30)
# Filter out ignored senders
echo "$EMAILS" | grep -vE "$IGNORE_SENDERS"
Simple. Effective at first. Then it broke down.
Why It Failed
- Manual tuning via code edits — Every new spam sender meant editing the bash script and redeploying.
- No prioritization — All non-ignored emails were treated equally. A bank statement had the same urgency as a work email.
- No security layer — Phishing emails sailed through if they didn’t match ignore patterns.
- No learning — The system never got smarter. Same mistakes every time.
- Brittle pattern matching — Regex in bash is fragile. Edge cases broke constantly.
After a few weeks, the ignore list had grown to 30+ patterns. Adding a new rule required reading through the entire script to avoid conflicts. This wasn’t sustainable.
The Redesign: A 4-Phase Pipeline
I scrapped the bash script and rebuilt the system as a multi-phase Python pipeline:
- Rule Engine — Pattern-based filtering (the good parts of the old system)
- Security Scanner — Phishing & spoofing detection (new)
- Priority Scoring — 0-100 score based on multiple factors (new)
- Learning — Adaptive filtering from feedback (new)
Each phase builds on the previous one. An email flows through all four before a notification decision is made.
Gmail → Rule Engine → Security Scanner → Priority Scoring → Notification
↓ ↓ ↓
Ignore Flag Risk Score (0-100)
↓
Immediate / Batch / Digest / Ignore
Phase 1: Rule Engine
The rule engine is the old ignore list, but structured.
Instead of hardcoded bash patterns:
{
"senders": {
"always_notify": [
"@company.com",
"boss@"
],
"ignore": [
"newsletter@",
"no-reply@",
"marketing@"
],
"digest_only": [
"github.com",
"medium.com"
]
},
"patterns": {
"ignore_subjects": [
"^Daily Digest",
"^Weekly Summary",
"Gas bill"
]
},
"keywords": {
"high_priority": [
"urgent",
"deadline",
"invoice",
"payment failed"
],
"low_priority": [
"unsubscribe",
"promotional",
"marketing"
]
}
}
Benefits:
- Easy to edit (JSON, not code)
- Supports regex patterns for subjects
- Whitelisting (always notify) + blacklisting (ignore)
- Keyword-based scoring
Phase 2: Security Scanner
Phishing emails are getting sophisticated. The old system had zero protection.
The security scanner checks for three red flags:
1. Sender Spoofing
Catches emails like:
From: "PayPal Security" <[email protected]>
The display name says “PayPal” but the actual sender is [email protected]. This is a spoofing attempt.
The scanner extracts both the display name and the actual domain, then checks if they match:
def check_sender_match(from_addr: str) -> Tuple[bool, str]:
# Extract display name and email
match = re.search(r'(.+?)\s*<(.+?)>', from_addr)
if not match:
return False, ""
display_name = match.group(1).strip()
actual_email = match.group(2).strip()
actual_domain = actual_email.split('@')[-1].lower()
# Check if display name contains a different domain
if any(domain in display_name.lower()
for domain in COMMON_BRANDS):
if domain not in actual_domain:
return True, f"Display name '{display_name}' doesn't match {actual_domain}"
return False, ""
Common brands checked: PayPal, Amazon, Google, Microsoft, Apple, Bank of America, Chase, Wells Fargo.
2. Suspicious Links
URL shorteners (bit.ly, tinyurl.com) and sketchy TLDs (.tk, .ml, .ga) are phishing red flags.
SUSPICIOUS_TLDS = ['.tk', '.ml', '.ga', '.cf', '.gq']
SHORTENER_DOMAINS = ['bit.ly', 'tinyurl.com', 'goo.gl', 't.co']
def has_suspicious_links(email_body: str) -> bool:
urls = re.findall(r'https?://[^\s<>"]+', email_body)
for url in urls:
domain = urlparse(url).netloc.lower()
if any(domain.endswith(tld) for tld in SUSPICIOUS_TLDS):
return True
if domain in SHORTENER_DOMAINS:
return True
return False
3. Phishing Keywords
Common phishing patterns:
PHISHING_KEYWORDS = [
"verify your account",
"unusual activity detected",
"suspended account",
"click here immediately",
"confirm your identity",
"reset your password",
"account will be closed",
"update payment information"
]
If any of these appear in the subject or body, score drops by 50 points.
Example flagged email:
Score: 10 (IGNORE)
From: "PayPal Security" <[email protected]>
Subj: Verify your account immediately
Why: -40: Sender mismatch | -50: Phishing indicator | -30: Suspicious link
The system caught this before I ever saw it.
Phase 3: Priority Scoring
Not all emails are created equal. Instead of binary (spam/not spam), the system assigns a 0-100 score.
Base score: 50 (neutral)
Adjustments:
| Factor | Score Change |
|---|---|
| Whitelisted sender | +40 |
| Gmail IMPORTANT label | +20 |
| High priority keyword | +30 |
| Low priority keyword | -30 |
| Reply to my email | +25 |
| Directly addressed to me | +20 |
| Phishing indicator | -50 |
| Sender mismatch | -40 |
| Suspicious link | -30 |
Example scoring:
From: [email protected]
Subj: Deadline for quarterly report - action needed
+40: Whitelisted sender (@company.com)
+30: High priority keyword ("deadline")
+30: High priority keyword ("action")
Final score: 100 (IMMEDIATE)
4-Level Delivery System
Instead of immediate notifications for everything, emails are batched by urgency:
🔴 Immediate (80-100 points)
- Checked: Hourly during awake hours (07:00-23:00)
- Delivered: Immediately when found
- Examples: Whitelisted senders, urgent keywords, direct messages
🟡 Batch (50-79 points)
- Checked: Same as immediate
- Delivered: Once daily at 20:00
- Examples: Gmail Important, medium priority, replies
🟢 Digest (20-49 points)
- Checked: Same as immediate
- Delivered: Weekly (Sunday evening with GTD ritual)
- Examples: Newsletters, FYI emails, low priority
⚫ Ignore (0-19 points)
- Delivered: Never
- Logged: Yes (available for review)
- Examples: Marketing, spam, explicitly ignored senders
Why this works: I get critical emails immediately. Medium-priority stuff waits until evening (when I check my inbox anyway). Low-priority gets batched weekly for review. Noise gets silenced completely.
Phase 4: Learning System
The system learns from feedback. Two interfaces:
CLI Tool
# Ignore a sender
python3 email_trainer.py ignore "[email protected]"
# Whitelist a sender
python3 email_trainer.py whitelist "[email protected]"
# Ignore by keyword
python3 email_trainer.py ignore-keyword "promotional"
# Review recent filtered emails
python3 email_trainer.py review --limit 20
# Show stats
python3 email_trainer.py stats
Natural Language Feedback
In Telegram (via the OpenClaw assistant), I can just say:
"ignore [email protected]"
"whitelist [email protected]"
"always notify from @company.com"
"never show comparethemarket"
When replying to an email notification:
"ignore this sender"
The assistant parses the intent, executes the command, and confirms:
✅ Now ignoring: [email protected]
All feedback is logged to a JSONL file for future analysis:
{
"timestamp": "2024-12-19T10:30:00Z",
"action": "ignore_sender_added",
"sender": "[email protected]",
"reason": "User feedback",
"email_id": "abc123",
"score": 25
}
This log will eventually feed a machine learning classifier, but for now it’s just structured history.
Implementation Details
Silent Wrapper for Cron Jobs
The email processor runs via cron every hour (for immediate checks), daily at 20:00 (for batch), and weekly (for digest).
Problem: How do you prevent spurious “no emails found” notifications?
Solution: A silent wrapper that only outputs when emails exist:
#!/bin/bash
# email_check_silent.sh
MIN_SCORE=$1
MAX_SCORE=${2:-100}
FORMAT=${3:-""}
# Run processor
OUTPUT=$(python3 /path/to/email_processor.py --min-score $MIN_SCORE --max-score $MAX_SCORE $FORMAT)
# Only output if emails were found
if echo "$OUTPUT" | grep -q "__IMPORTANT_COUNT__"; then
COUNT=$(echo "$OUTPUT" | grep "__IMPORTANT_COUNT__" | cut -d= -f2)
if [ "$COUNT" -gt 0 ]; then
echo "$OUTPUT"
exit 0
fi
fi
# Silent success (no emails)
exit 0
When the cron job runs:
- If emails found → output email data → assistant sends notification
- If no emails → complete silence → no notification
No more “I checked your email and found nothing” spam.
Structured Output Format
The processor outputs a special format for the assistant to parse:
__IMPORTANT_COUNT__=2
__IMPORTANT_1__
From: [email protected]
Subj: Important message
Score: 85 (HIGH)
Why: +40: Whitelisted sender | +20: Gmail marked important
__IMPORTANT_2__
From: [email protected]
Subj: Deadline tomorrow
Score: 90 (HIGH)
Why: +40: Whitelisted | +30: High priority keyword 'deadline'
The assistant reads __IMPORTANT_COUNT__, loops through __IMPORTANT_N__ blocks, and formats a Telegram message:
📬 2 important emails
1. Score: 85 (HIGH)
From: [email protected]
Subj: Important message
Why: +40: Whitelisted sender | +20: Gmail marked important
2. Score: 90 (HIGH)
From: [email protected]
Subj: Deadline tomorrow
Why: +40: Whitelisted | +30: High priority keyword 'deadline'
Results
Before:
- 30+ emails a day → all treated equally
- Manually filtering noise every time
- No phishing protection
- Hardcoded bash patterns breaking constantly
After:
- ~3 immediate notifications a day (truly important)
- ~5-10 batch emails at 20:00 (worth checking)
- ~20-30 digest emails weekly (skim or ignore)
- Phishing attempts flagged before I see them
- Easy tuning via CLI or natural language
Time saved: ~15 minutes a day (previously scanning inbox for important emails).
Security wins: 3 phishing attempts caught in the first week.
What’s Next
Short-term improvements
Weekly review digest — Auto-generate summary of filtered emails
“You ignored 247 emails this week. Top senders: [email protected] (23), [email protected] (18). Any of these need adjustment?”False positive detection — Flag when an ignored email gets a reply
“You replied to an email from [email protected], but it’s on your ignore list. Want to whitelist it?”Thread context — Boost priority for replies in active conversations
Long-term enhancements
Gmail API integration — Direct access (no CLI dependency), faster
Attachment analysis — Smart handling of invoices, receipts, PDFs
“This looks like an invoice. Auto-categorize as HIGH?”ML-based scoring — Train a classifier on implicit feedback:
- Emails you open → boost sender priority
- Emails you ignore → lower priority
- Emails you reply to → strong signal
Code Structure
Files created:
| File | Purpose |
|---|---|
config/email-rules.json | Filtering rules & scoring config |
scripts/email_processor.py | Main processor (all 4 phases) |
scripts/email_trainer.py | Training CLI |
scripts/email_feedback_parser.py | Natural language feedback parser |
scripts/email_check_silent.sh | Silent wrapper for cron jobs |
memory/email-training-log.jsonl | Training event log |
memory/email-processed-cache.json | Recent email cache |
Total code: ~45 KB
Lines of Python: ~500
Self-contained, no external dependencies beyond Python stdlib and gog CLI (Google Workspace tool).
Lessons Learned
1. Start simple, evolve based on pain points
The bash script worked until it didn’t. The redesign addressed specific failures (manual tuning, no security, no scoring).
2. Separate config from code
JSON config means non-technical users can adjust rules without touching Python. The assistant can also edit the config programmatically.
3. Transparency matters
Every email shows its score + reasoning. When the system makes a mistake, I understand why and can fix the underlying rule.
4. Silent success is a feature
Cron jobs that say “nothing to report” create noise. Silence when there’s nothing important is better than empty notifications. This principle — and its failure modes — is explored in depth in The Silent Killer in AI Automation.
5. Learning systems need feedback loops
The training CLI + natural language parser make it trivial to tune the system. No friction = faster convergence to good behavior.
Running a similar system? Share your filtering strategy in the OpenClaw Discord. I’d love to hear what works (and what doesn’t).