Case Study Breakdown
Teardown: the inbox that drafts itself
What we learned building an inbox system for one of our ecommerce clients. Why the noise filter is plain rules, why the model gets a written rule instead of judgement, and why the draft, not the send, is the product.
DashboardLim · Systems desk
2026-06-246 min readWe built an inbox system for one of our ecommerce clients and it has been running ever since. Every screen below is that system on a real month of their mail. Here is what they came to us with, what broke early, and what it does now.
The system, on a real month
The board
What arrived, what the rules threw out, and what still needs a person.
The problem
What their inbox was actually costing them
One of our ecommerce clients came to us because their mail had stopped being a queue and started being their day. A founder's personal inbox, the shared team addresses, and messages arriving from their store and their social accounts on top.
Around fifty customer emails a day in peak season, and nothing converged anywhere except on a person. Someone was clearing the backlog late at night.
What they asked us for was not an AI that answers customers. It was a way to stop having to read everything.
The pain in a busy inbox is rarely the volume. It is the attention every message takes on the way past.
What we built
Rules first, model second
The obvious build points a model at the inbox and lets it sort. We did the opposite.
Plain rules run before anything else. Blocklists, allow-lists and pattern matches the team edits themselves from the dashboard. Newsletters, noreply notifications and system mail never reach the model at all.
In the month on these screens that removed 5,536 messages, and every one is still counted against the rule that dropped it. The model only reasons over mail a person would actually care about, which made everything after it both cheaper and sharper.
What to look at
- The rules belong to the client. They add and remove them from this screen and the change is live.
- Nothing is deleted quietly. A dropped message stays in the ledger with the rule that dropped it.
The cheapest step in an AI system is usually the one where no AI runs.
What broke
The early failures were ours, not the model's
It read the newest message in a thread and politely asked a customer for something they had already sent us. The fix was not a smarter model, it was a sharper rule. Read the whole thread, not the latest message.
It guessed at prices, and it once offered a discount that did not exist. So pricing came out of every drafted reply. Drafts point at the website now instead of quoting a number.
A spam subject carrying the right keyword got flagged urgent with high confidence. Internal forwards between team mailboxes were dropped as noise until the rules learned what a forward looks like. Each one of those was a rule we had not written yet.
Escalation stayed deterministic because a wrong guess there is expensive. Drafting uses the model because the bottleneck was a slow first draft, not a wrong one.
What we changed
Money, legal and safety never reach the model
Anything touching money, legal or safety is caught before the model has a chance to be clever about it. Refunds over an agreed threshold, chargebacks, legal correspondence and anything diet critical go to a named person instead.
They land in the alert queue with the reason they landed there, the mailbox they came into and how long they have been waiting.
Opening one shows the classification, the recommended action and the whole thread, so whoever picks it up is not starting from a subject line.
What to look at
- A refund is never answered by the system. It is routed, with everything a person needs to answer it.
- Rules beat judgement in the places where a wrong answer costs real money.
The skill is knowing which messages belong to a rule and which belong to a model.
How it works now
The draft is the product
Nothing sends itself. Every reply the system writes waits in that mailbox's own Drafts folder, marked as a draft, until a person reads it and presses send.
In the month on these screens it drafted 1,052 replies and flagged 1,204 more for a human. Their founder works two queues and a morning digest instead of a midnight inbox.
There is a path to auto sending and it is strict on purpose. A category has to run in shadow mode for two weeks with almost nothing edited before anyone will discuss it, and money, legal, dietary and complaints never qualify at all. Nothing has graduated yet.
Nobody was promised an empty inbox. What they got was a first draft waiting on every message that deserves one.
What I would take from it
None of this needed a new tool. It needed a few decisions made in the right order.
Filter with plain rules first
Drop the obvious noise deterministically and log every drop with its reason. Save the model for the mail that deserves one.
Give the model a written rule, not judgement
Write down the call the founder already makes in their head, then hold the model to it.
Make it read the whole thread
Most embarrassing replies come from answering the newest message instead of the conversation.
Keep money, legal and safety out of its reach
A deterministic check and a named human beat a confident model everywhere that being wrong is expensive.
Draft, do not send
Let it prepare the reply and keep a person on the send. Trust is earned one draft at a time.
The proof behind this issue
The full case study, with the architecture and the message flow, is on the proof feed.
Case study
· Inbox automation
Founder Inbox OS
Seven inboxes and customer channels were feeding one overloaded team. We built a human-reviewed system that sorts every message, drafts the next reply and escalates anything that needs judgement.
7
mailboxes triaged
81
noise messages filtered in one morning
0
replies sent without human approval
Freckleberry
2026-06-18OpenClawn8nRailwayRead the case study →Get the next issue in your inbox
Real systems we've shipped, with practical lessons you can use today.
Where is your inbox costing you?
If someone on your team is still clearing the backlog at night, that is the part worth building first. Tell us what your mail looks like and we will show you what this build looked like from the inside.
Keep exploring the Intelligence Library
Workflows
Three systems to install before Q4 planning

2026-07-01
Operations
What we shipped last quarter, and what it changed

2026-09-04
Sales & CRM
AI Sales Call Memory

2026-06-05
OpenClaw

