Email overload is a modern disease. Hundreds of emails a day, most of them noise. Newsletters you forgot you subscribed to. Marketing emails disguised as personal messages. Bank statements. Gas bills. The occasional important email buried in the flood.

I tried fixing this with a simple ignore list. It didn’t scale.

This is the story of how I built a multi-layer email filtering system that replaced hardcoded bash patterns with intelligent scoring, security scanning, and adaptive learning.

The Initial Setup (That Didn’t Work)

My first attempt was straightforward: a bash script that checked Gmail via the gog CLI, applied a few regex patterns, and ignored obvious spam.

#!/bin/bash
# Old system: tmp_gmail_check.sh

# Hardcoded ignore patterns
IGNORE_SENDERS="newsletter@|no-reply@|comparethemarket.com|confused.com"

# Fetch unread emails
EMAILS=$(gog gmail messages search "is:unread" --max 30)

# Filter out ignored senders
echo "$EMAILS" | grep -vE "$IGNORE_SENDERS"

Simple. Effective at first. Then it broke down.

Why It Failed

  1. Manual tuning via code edits — Every new spam sender meant editing the bash script and redeploying.
  2. No prioritization — All non-ignored emails were treated equally. A bank statement had the same urgency as a work email.
  3. No security layer — Phishing emails sailed through if they didn’t match ignore patterns.
  4. No learning — The system never got smarter. Same mistakes every time.
  5. Brittle pattern matching — Regex in bash is fragile. Edge cases broke constantly.

After a few weeks, the ignore list had grown to 30+ patterns. Adding a new rule required reading through the entire script to avoid conflicts. This wasn’t sustainable.

The Redesign: A 4-Phase Pipeline

I scrapped the bash script and rebuilt the system as a multi-phase Python pipeline:

  1. Rule Engine — Pattern-based filtering (the good parts of the old system)
  2. Security Scanner — Phishing & spoofing detection (new)
  3. Priority Scoring — 0-100 score based on multiple factors (new)
  4. Learning — Adaptive filtering from feedback (new)

Each phase builds on the previous one. An email flows through all four before a notification decision is made.

Gmail → Rule Engine → Security Scanner → Priority Scoring → Notification
            ↓              ↓                    ↓
          Ignore       Flag Risk           Score (0-100)
                                                ↓
                                   Immediate / Batch / Digest / Ignore

Phase 1: Rule Engine

The rule engine is the old ignore list, but structured.

Instead of hardcoded bash patterns:

{
  "senders": {
    "always_notify": [
      "@company.com",
      "boss@"
    ],
    "ignore": [
      "newsletter@",
      "no-reply@",
      "marketing@"
    ],
    "digest_only": [
      "github.com",
      "medium.com"
    ]
  },
  "patterns": {
    "ignore_subjects": [
      "^Daily Digest",
      "^Weekly Summary",
      "Gas bill"
    ]
  },
  "keywords": {
    "high_priority": [
      "urgent",
      "deadline",
      "invoice",
      "payment failed"
    ],
    "low_priority": [
      "unsubscribe",
      "promotional",
      "marketing"
    ]
  }
}

Benefits:

  • Easy to edit (JSON, not code)
  • Supports regex patterns for subjects
  • Whitelisting (always notify) + blacklisting (ignore)
  • Keyword-based scoring

Phase 2: Security Scanner

Phishing emails are getting sophisticated. The old system had zero protection.

The security scanner checks for three red flags:

1. Sender Spoofing

Catches emails like:

From: "PayPal Security" <[email protected]>

The display name says “PayPal” but the actual sender is [email protected]. This is a spoofing attempt.

The scanner extracts both the display name and the actual domain, then checks if they match:

def check_sender_match(from_addr: str) -> Tuple[bool, str]:
    # Extract display name and email
    match = re.search(r'(.+?)\s*<(.+?)>', from_addr)
    if not match:
        return False, ""
    
    display_name = match.group(1).strip()
    actual_email = match.group(2).strip()
    actual_domain = actual_email.split('@')[-1].lower()
    
    # Check if display name contains a different domain
    if any(domain in display_name.lower() 
           for domain in COMMON_BRANDS):
        if domain not in actual_domain:
            return True, f"Display name '{display_name}' doesn't match {actual_domain}"
    
    return False, ""

Common brands checked: PayPal, Amazon, Google, Microsoft, Apple, Bank of America, Chase, Wells Fargo.

URL shorteners (bit.ly, tinyurl.com) and sketchy TLDs (.tk, .ml, .ga) are phishing red flags.

SUSPICIOUS_TLDS = ['.tk', '.ml', '.ga', '.cf', '.gq']
SHORTENER_DOMAINS = ['bit.ly', 'tinyurl.com', 'goo.gl', 't.co']

def has_suspicious_links(email_body: str) -> bool:
    urls = re.findall(r'https?://[^\s<>"]+', email_body)
    for url in urls:
        domain = urlparse(url).netloc.lower()
        if any(domain.endswith(tld) for tld in SUSPICIOUS_TLDS):
            return True
        if domain in SHORTENER_DOMAINS:
            return True
    return False

3. Phishing Keywords

Common phishing patterns:

PHISHING_KEYWORDS = [
    "verify your account",
    "unusual activity detected",
    "suspended account",
    "click here immediately",
    "confirm your identity",
    "reset your password",
    "account will be closed",
    "update payment information"
]

If any of these appear in the subject or body, score drops by 50 points.

Example flagged email:

Score: 10 (IGNORE)
From: "PayPal Security" <[email protected]>
Subj: Verify your account immediately
Why: -40: Sender mismatch | -50: Phishing indicator | -30: Suspicious link

The system caught this before I ever saw it.

Phase 3: Priority Scoring

Not all emails are created equal. Instead of binary (spam/not spam), the system assigns a 0-100 score.

Base score: 50 (neutral)

Adjustments:

FactorScore Change
Whitelisted sender+40
Gmail IMPORTANT label+20
High priority keyword+30
Low priority keyword-30
Reply to my email+25
Directly addressed to me+20
Phishing indicator-50
Sender mismatch-40
Suspicious link-30

Example scoring:

From: [email protected]
Subj: Deadline for quarterly report - action needed

+40: Whitelisted sender (@company.com)
+30: High priority keyword ("deadline")
+30: High priority keyword ("action")

Final score: 100 (IMMEDIATE)

4-Level Delivery System

Instead of immediate notifications for everything, emails are batched by urgency:

🔴 Immediate (80-100 points)

  • Checked: Hourly during awake hours (07:00-23:00)
  • Delivered: Immediately when found
  • Examples: Whitelisted senders, urgent keywords, direct messages

🟡 Batch (50-79 points)

  • Checked: Same as immediate
  • Delivered: Once daily at 20:00
  • Examples: Gmail Important, medium priority, replies

🟢 Digest (20-49 points)

  • Checked: Same as immediate
  • Delivered: Weekly (Sunday evening with GTD ritual)
  • Examples: Newsletters, FYI emails, low priority

⚫ Ignore (0-19 points)

  • Delivered: Never
  • Logged: Yes (available for review)
  • Examples: Marketing, spam, explicitly ignored senders

Why this works: I get critical emails immediately. Medium-priority stuff waits until evening (when I check my inbox anyway). Low-priority gets batched weekly for review. Noise gets silenced completely.

Phase 4: Learning System

The system learns from feedback. Two interfaces:

CLI Tool

# Ignore a sender
python3 email_trainer.py ignore "[email protected]"

# Whitelist a sender
python3 email_trainer.py whitelist "[email protected]"

# Ignore by keyword
python3 email_trainer.py ignore-keyword "promotional"

# Review recent filtered emails
python3 email_trainer.py review --limit 20

# Show stats
python3 email_trainer.py stats

Natural Language Feedback

In Telegram (via the OpenClaw assistant), I can just say:

"ignore [email protected]"
"whitelist [email protected]"
"always notify from @company.com"
"never show comparethemarket"

When replying to an email notification:

"ignore this sender"

The assistant parses the intent, executes the command, and confirms:

✅ Now ignoring: [email protected]

All feedback is logged to a JSONL file for future analysis:

{
  "timestamp": "2024-12-19T10:30:00Z",
  "action": "ignore_sender_added",
  "sender": "[email protected]",
  "reason": "User feedback",
  "email_id": "abc123",
  "score": 25
}

This log will eventually feed a machine learning classifier, but for now it’s just structured history.

Implementation Details

Silent Wrapper for Cron Jobs

The email processor runs via cron every hour (for immediate checks), daily at 20:00 (for batch), and weekly (for digest).

Problem: How do you prevent spurious “no emails found” notifications?

Solution: A silent wrapper that only outputs when emails exist:

#!/bin/bash
# email_check_silent.sh
MIN_SCORE=$1
MAX_SCORE=${2:-100}
FORMAT=${3:-""}

# Run processor
OUTPUT=$(python3 /path/to/email_processor.py --min-score $MIN_SCORE --max-score $MAX_SCORE $FORMAT)

# Only output if emails were found
if echo "$OUTPUT" | grep -q "__IMPORTANT_COUNT__"; then
  COUNT=$(echo "$OUTPUT" | grep "__IMPORTANT_COUNT__" | cut -d= -f2)
  if [ "$COUNT" -gt 0 ]; then
    echo "$OUTPUT"
    exit 0
  fi
fi

# Silent success (no emails)
exit 0

When the cron job runs:

  • If emails found → output email data → assistant sends notification
  • If no emails → complete silence → no notification

No more “I checked your email and found nothing” spam.

Structured Output Format

The processor outputs a special format for the assistant to parse:

__IMPORTANT_COUNT__=2
__IMPORTANT_1__
From: [email protected]
Subj: Important message
Score: 85 (HIGH)
Why: +40: Whitelisted sender | +20: Gmail marked important
__IMPORTANT_2__
From: [email protected]
Subj: Deadline tomorrow
Score: 90 (HIGH)
Why: +40: Whitelisted | +30: High priority keyword 'deadline'

The assistant reads __IMPORTANT_COUNT__, loops through __IMPORTANT_N__ blocks, and formats a Telegram message:

📬 2 important emails

1. Score: 85 (HIGH)
   From: [email protected]
   Subj: Important message
   Why: +40: Whitelisted sender | +20: Gmail marked important

2. Score: 90 (HIGH)
   From: [email protected]
   Subj: Deadline tomorrow
   Why: +40: Whitelisted | +30: High priority keyword 'deadline'

Results

Before:

  • 30+ emails a day → all treated equally
  • Manually filtering noise every time
  • No phishing protection
  • Hardcoded bash patterns breaking constantly

After:

  • ~3 immediate notifications a day (truly important)
  • ~5-10 batch emails at 20:00 (worth checking)
  • ~20-30 digest emails weekly (skim or ignore)
  • Phishing attempts flagged before I see them
  • Easy tuning via CLI or natural language

Time saved: ~15 minutes a day (previously scanning inbox for important emails).

Security wins: 3 phishing attempts caught in the first week.

What’s Next

Short-term improvements

  1. Weekly review digest — Auto-generate summary of filtered emails
    “You ignored 247 emails this week. Top senders: [email protected] (23), [email protected] (18). Any of these need adjustment?”

  2. False positive detection — Flag when an ignored email gets a reply
    “You replied to an email from [email protected], but it’s on your ignore list. Want to whitelist it?”

  3. Thread context — Boost priority for replies in active conversations

Long-term enhancements

  1. Gmail API integration — Direct access (no CLI dependency), faster

  2. Attachment analysis — Smart handling of invoices, receipts, PDFs
    “This looks like an invoice. Auto-categorize as HIGH?”

  3. ML-based scoring — Train a classifier on implicit feedback:

    • Emails you open → boost sender priority
    • Emails you ignore → lower priority
    • Emails you reply to → strong signal

Code Structure

Files created:

FilePurpose
config/email-rules.jsonFiltering rules & scoring config
scripts/email_processor.pyMain processor (all 4 phases)
scripts/email_trainer.pyTraining CLI
scripts/email_feedback_parser.pyNatural language feedback parser
scripts/email_check_silent.shSilent wrapper for cron jobs
memory/email-training-log.jsonlTraining event log
memory/email-processed-cache.jsonRecent email cache

Total code: ~45 KB
Lines of Python: ~500

Self-contained, no external dependencies beyond Python stdlib and gog CLI (Google Workspace tool).

Lessons Learned

1. Start simple, evolve based on pain points

The bash script worked until it didn’t. The redesign addressed specific failures (manual tuning, no security, no scoring).

2. Separate config from code

JSON config means non-technical users can adjust rules without touching Python. The assistant can also edit the config programmatically.

3. Transparency matters

Every email shows its score + reasoning. When the system makes a mistake, I understand why and can fix the underlying rule.

4. Silent success is a feature

Cron jobs that say “nothing to report” create noise. Silence when there’s nothing important is better than empty notifications. This principle — and its failure modes — is explored in depth in The Silent Killer in AI Automation.

5. Learning systems need feedback loops

The training CLI + natural language parser make it trivial to tune the system. No friction = faster convergence to good behavior.


Running a similar system? Share your filtering strategy in the OpenClaw Discord. I’d love to hear what works (and what doesn’t).