I Put a Small Model in Charge of My Bank Statement
I have about 1,500 transactions and no interest in sorting them by hand. So I set up a model that sorts all of them, tells me which ones it isn’t sure about, and remembers what I say when I correct it.
Recently I moved nine months of bank history into myqi, the private app I run for myself. It’s about 1,500 transactions, encrypted, and only readable on my own machine. The page showed the feed and a couple of charts, but nothing said what any of the money was actually for.
Every budgeting app I’ve tried made me do that part by hand, and honestly that’s why I stopped using each of them. I didn’t want another place to sort transactions. I wanted something that would look at every row and decide, tell me how sure it was, and remember what I said when I disagreed with it. That turned out to be a pretty specific ask, and the usual chat models weren’t the right tool for it.
A different kind of AI
The models everyone knows write text (ChatGPT, Claude, Gemini). You can ask the same question five times and get five different answers. The one I ended up using, Jev from TypeSafe, is different. It doesn’t write anything. You give it something to look at, you set the guardrails for it, and it picks one answer and tells you how sure it is. If you want to look at it yourself, it’s at typesafe.ai.
The guardrails are the list of answers it’s allowed to give, and what each one means. For my transactions, that’s sixteen categories, with a sentence for each one explaining what belongs there. Groceries means supermarkets. Transfer means money moving between my own accounts. Jev can only pick from that list. It can’t come up with a seventeenth category, and it can’t answer with a paragraph. If nothing fits, the only thing it can say is “other,” because I put that on the list too.
What comes back is the pick, plus a number for how sure it is. If a row is clearly a supermarket, the number is close to one. If it could be two or three things, the number drops, and that’s my signal to look at it myself. The same row, sent again, gets the same answer, because it isn’t writing from scratch each time. It’s choosing.
At first this seemed like a downgrade. It can’t chat, it can’t explain itself, it can’t write a summary. But for sorting 1,500 transactions, none of that matters, and the limits are what make it useful.
It can’t give me an answer I didn’t put on the list, so I never have to clean up something it made up. It’s fast and cheap enough that I can run it on every single transaction instead of a sample: all 1,500 took under six seconds and cost a few cents.
Every answer also comes with a number, and Jev picks that number, not me. It’s Jev’s way of saying how sure it is. A 0.96 means “I’m almost certain this is groceries.” A 0.5 means “this could be a couple of things.” I don’t touch those numbers. The only thing I decide is the cutoff. I set it at 0.8: if Jev’s number is 0.8 or higher, I take its answer. If it’s lower, the row goes on a list for me to look at. If I want to see more rows, I raise the cutoff. If I trust it more, I lower it.
So the app decides when to ask and what to do with the answer. Jev only answers. That split is the whole reason I was comfortable letting it loose on my money.
Putting it to work
For every transaction, Jev gets five things:
- the description
- the amount
- whether the money came in or went out
- which bank it’s from
- the date
It never sees account numbers, account names or balances. Then it answers five questions about each one:
- Which of my sixteen categories is this?
- Is this money moving between my own accounts?
- Is this a recurring charge?
- Is this a need or a want?
- Is this worth a second look?
The first run went through all 1,508 transactions in under six seconds. Jev was sure about 1,144 of them, meaning its number was 0.8 or higher. It was unsure about 117, and it flagged 13 as surprising. The surprising ones were a reasonable “you might want to look at these” from something that has never seen my budget: a few big one-off purchases, a couple of large transfers, a trip.
The 117 it was unsure about had one thing in common: the description didn’t say much. Some were a payment app’s name with no store behind it. Some were transfers with just a reference number. Looking at them myself, I couldn’t tell either. So Jev isn’t guessing on those. It’s handing them to me, and it’s handing me the right ones. In myqi, those rows go to the top of a “Needs review” list, and that list is the only place I spend time. I don’t go through the other 1,400.
How I Keep Improving Jev’s Criteria
This is the part I care about most. Jev doesn’t learn on its own. It works from a set of criteria I control: sixteen category definitions, plus every answer I’ve ever given on a row. When Jev gets something wrong, the fix is never “train the model.” It’s “sharpen the criteria.” There are two ways I do that.
The first is answering a row. When I change a category on a row, three things happen:
- That row keeps my answer.
- Every other row from the same merchant switches to my answer too, and so will new ones that show up later. I only have to say it once.
- The page confirms it. Something like “Saved Groceries. 59 other rows of that merchant now follow it.”
That answer also gets added to the criteria. The next time Jev looks at anything from that merchant, “this one is Dining” is right there in what it’s working from.
The second is rewriting a definition. Early on I noticed refunds were being filed as income, because my definition of income said “refunds.” That wasn’t a Jev mistake, it was doing what I wrote. I changed one sentence so a refund keeps the store’s category, then hit “Re-judge all.” Jev went back through every row using the new criteria, and it kept the old answers so I could see what changed.
Either way, Jev only ever works from the latest version. The page keeps score, and right now Jev agrees with what I’ve taught it 99 percent of the time. That’s the number that made this worth it. Out of 1,500 transactions, I only have to give real attention to the 1 percent that actually need me.
If Jev got a row right and I want to say so, there’s a “Looks right” button that does the same three things without changing the category. Either way, one tap handles the whole merchant. Corrections queue up in the background, so I can go through thirty rows on my phone without waiting on each one.
How Could Jev Be Used in Businesses
My bank statement is 1,500 rows and one person checking them. A lot of businesses have a bigger version of the same thing, and it usually looks like bookkeeping. Someone spends hours every week logging orders, quantities and shipments into spreadsheets, not because the logging is useful on its own, but because it’s the only way to catch it when something doesn’t line up. Most weeks nothing is wrong, and you did all that work for the few rows that were.
Take a small clothing manufacturer. An order comes in by email. Someone types it into a work-in-progress sheet, makes a cutting ticket, and sends it to a cutter. The cutter sends back a report of what was actually cut. Someone compares it to the ticket, updates the quantities, writes purchase orders for the sewers, and logs what shipped. The same numbers get typed into five or six spreadsheets so that a mismatch will eventually get noticed.
Every one of those checks is a question with a short list of answers:
- Is this email a new order, a change, or a cancellation?
- Does this line on the cutter’s report belong to this ticket?
- Do the report’s totals match the ticket? Match, quantity differs, sizes differ, or date differs.
- Does this carton on the packing list match an order line?
Jev would answer those the moment an email or a report lands, with a number for how sure it is. The ones it’s sure about get logged on their own. The ones it isn’t sure about, and the ones that don’t match, go on a short list for the owner. When the owner answers one, say “this cutter always lists sizes on separate lines, that’s not a mismatch,” it goes into the criteria, and the next report from that cutter is read with that in mind. The owner still decides what a mismatch means. They just stop retyping numbers to find out whether there is one.
How this benefits businesses
Every business owner I talk to says the same thing: the one thing they wish they had more of is time. Most of the time they’re losing goes to sorting through noise. Logging rows that were fine. Reading reports that matched. Looking for the one thing that’s off in a pile of things that aren’t.
This flips that. You skip the haystack and go straight to the needle. The model does the sorting, and the only thing that reaches you is the short list that actually needs a decision. If it’s not right the first time, you fix it once, that fix goes into the criteria, and it stays fixed for every row like it. The list gets shorter every week.
Start with the one check someone already does by hand every week. Write down the question they’re really asking, and the short list of answers it can have. Let the model do all of them, and only look at the ones it isn’t sure about. Then take the hours you get back and spend them on something that actually moves the business, or on something you enjoy. That’s what the time was for in the first place.