I had three months of uncategorised expenses sitting in my accounting software like an unread inbox of pure shame. I know the rule — categorise as you go, every week, do not let it pile up. I know it the exact same way I know I should floss daily. Then a busy quarter happened, life happened, and by the time I finally opened the screen there were 187 transactions staring back at me with a tax deadline exactly two weeks out. The backlog had quietly achieved sentience, and it was deeply disappointed in me.
The old way, which I will never do again
Manually categorising 187 transactions is a uniquely soul-flattening kind of tedium. You read a merchant name, decide the category, hunt for it in a dropdown, click it, move to the next line, and repeat the whole loop until your consciousness gently leaves your body somewhere around transaction number 40. The error rate climbs steadily as the boredom deepens. I always miscode a few, always miss something important, and always end up doing the whole thing on a Saturday I will never get back. It is the financial-admin equivalent of untangling a box of holiday lights.
What I did instead
I exported all 187 transactions to a CSV, trimmed it down to three clean columns — date, description, amount — and pasted the entire thing into the Bulk Expense Categorisation Assistant prompt, running it on Grok 4.1 Fast Reasoning. I included my exact accounting category list so nothing got invented, and I set a manual-review threshold of $1,000. Grok's Fast Reasoning model is precisely the right pick here, because the task is high-volume and pattern-heavy rather than deeply analytical — it tore through the whole list and returned a fully categorised table in about fifteen seconds flat. Then the real work, which turned out to be almost no work at all, began.
The fifteen-minute breakdown
Here is how the 187 transactions actually shook out, because the confidence tiers are what made this fast instead of just slightly-less-slow.
- 165 of 187 transactions came back at High confidence — I accepted those without touching them at all
- 18 came back Medium — each took about ten seconds to eyeball and confirm against my memory
- 4 came back Low — all genuinely ambiguous, and one turned out to be a real miscoding I had been repeating for months
- Total review time: 14 minutes for the entire three-month backlog, start to finish
Two of the prompt's flags more than earned their keep. The $1,000 review threshold caught a $2,200 software contract I had completely forgotten about, which needed its own dedicated line in our profit-and-loss rather than being buried in a lump under 'subscriptions.' Without that flag it would have quietly distorted the whole category.
I want to be honest about the trust question, because handing your real financial transactions to a model deserves a moment of skepticism. The reason this workflow is safe rather than reckless is the confidence column. The prompt never silently decides for me; it tells me how sure it is, line by line, and it pushes anything it is unsure about into a dedicated review list instead of quietly guessing. That design means the model handles the 88% that is genuinely obvious — the coffee shops, the recurring software, the obvious office supplies — and hands the genuinely ambiguous 12% straight back to the only person qualified to judge it, which is me. It is not pretending to be a bookkeeper. It is being an extremely fast first-pass sorter that knows what it does not know, and that distinction is the entire reason I trust it with a CSV of my actual spending.
The pattern it caught that I never would have
The Ambiguous Items section caught something genuinely useful that I had been missing for half a year. There were two charges from a vendor named 'Metro Services' that could plausibly have been office cleaning or could equally have been a transport expense. The prompt marked them Low confidence and explained both interpretations side by side. When I looked it up, I discovered I had been coding charges from that vendor inconsistently across multiple months — sometimes one, sometimes the other. The Low-confidence flag made me actually stop, investigate, and pick one category permanently. That is the kind of small, compounding error that never shows up until an accountant frowns at your books.
I now run this monthly, so a three-month backlog can never ambush me again, and when the categories are clean I hand the whole month straight to my Monthly Budget Review from Raw Transaction Data prompt for the actual analysis. One honest caveat: this is bookkeeping, not tax advice, and the prompt is firm about never commenting on deductibility — that conversation belongs entirely to my accountant, who is noticeably happier with what I now hand her. Grab the Bulk Expense Categorisation Assistant prompt on Prompt Dock, paste in your messiest month, and reclaim your Saturday from the dropdown menu.