14 ms·
GLM 5.2 is nearly as accurate as a human book keeper
- malfist 3mo ago> nearly as accurate as a human book keeper Anything to avoid using the metric system. Though seriously, what is this metric? Why would I care if an LLM is accurate as a human bookkeeper? Humans aren't exactly known for perfect recall.
- wat10000 3mo agoYou'd care if you had a human bookkeeper and you were considering replacing them with this company's AI bookkeeper.
- altruios 3mo agoWhy would I want both a less accurate book keeper and to incur all of the liability of doing the books myself?!
- murderfs 3mo agoPresumably your current book keeper is not your slave, and you have to pay them...
- altruios 3mo ago...again: liability is the key issue. Cost savings a not exactly an isolated issue here.
- murderfs 3mo agoThen why only have one human bookkeeper? Surely two would be better, since you can compare their results. But then, perhaps you should hire three, so you can figure out which one is right.
- CamperBob2 3mo agoGot some bad news for you about that "liability" thing: it's always been on you. Read the fine print on your tax forms sometime.
- onraglanroad 3mo agoIt's not less accurate. As commented above, the control knew they were being tested against the machine, so made sure to be super careful. In everyday life the human is less careful, and the machine costs 1% of the human.
- wat10000 3mo agoBecause it's way cheaper. Not saying it's a good tradeoff, but that appears to be their pitch.
- adamkurkiewicz 3mo agoHey, author of the blog post here. I was one of the human book-keepers for this benchmark (the preparer; my co-founder verified the VAT submission once ready), and given that at the time of doing this I knew I was eventually going to use this data for evaluating the models, I was super careful. So I guess this is a "good book-keeper". In the previous company our book-keepers made lots of mistakes; some serious enough that we had to restate our company's accounts.
- infecto 3mo agoIt’s a service that provides an AI bookkeeper so it’s a pretty relevant metric.
- helterskelter 3mo ago> They've done studies, you know. Sixty percent of the time, it works every time.
- raesene9 3mo agoInteresting write-up. Having been a bookkeeper a long time ago, I'm not too surprised at this being susceptible to automation by an LLM backed system. It seems also that the classes of error they encountered could be handled by improved skills/knowledge base access on the fine points of relevant tax legislation. The important part for their software ofc is, will they take responsibility for the output if HMRC come calling? Without that users are adopting the risk which they may not be keen to do (dealing with HMRC is not fun), with that it could be a very nice saving for a lot of small companies (and bad for the employees of a lot of accountancy firms)
- traverseda 3mo agoThis doesn't surprise me at all. You can really constrain this problem, give very narrow context, and get pretty reliable and reproducible results. I've gotten very good results with some vibe-coded deepseek book keeping. https://github.com/traverseda/beansync https://github.com/traverseda/beansync Parses emails or other sources, extracts numbers, correlates different transactions, web search, asks questions, stores notes (regex based, very simple). The hard part is getting good data, I'm sure that lexus nexus or whoever can get API access to my bank account and all my credit cards, but I can't. Email turned out to be the best way for most of my providers. Managed to avoid 2factor auth so far, but it will suck when I need it.
- adamkurkiewicz 3mo agoWe've got integrations with major UK banks. Curious if you'd like to use a polished product or be more interested in bank-feed-as-an-API type of use case? What banks do you use?
- traverseda 3mo agoCredit union atlantic. The possibility of you being able to profitably support small regional credit unions is pretty much 0. Maybe with a general browser use AI and an SMS portal. Also you'd need to bypass anti-bot protections, maybe solve captchas. In my opinion that is the thing holding back almost all of these accounting products that make consumers lives better. I've solved it because all my cards and bills happen to support email, so I can use that as my source of truth for all my credit cards.
- krupan 3mo agoNote to self: traverseda doesn't have 2factor auth on his email and his LLM seems to have full access. Hmmm
- traverseda 3mo agoThe token is stored on a LUKS encrypted drive on kwallet. You'd have just as much luck getting my password from my firefox installs sqlite DB. This is also how email clients in general work. Also this isn't really an agent. At least not a long lived one. Each email or transaction gets it's own session, the llm can make a few tool calls but must emit json as the final result. Very very short context lengths, very predictable results. The bot does not have access to my passwords, there are pre-defined scripts that fetch the data that the run an LLM call for each transaction discovered.
- petesergeant 3mo agoOh, I'm actively doing this at the moment. FreeAgent grabs my transactions from Wise already, and then I give it [Claude Code, in fact] a folder of PDFs to attach to my invoices, including figuring out VAT, and it's uploading what it found using the FreeAgent API. My accountant hasn't complained yet, and it seems considerably more accurate than when my wife was doing it. Quiet plug for https://github.com/pjlsergeant/byre https://github.com/pjlsergeant/byre which I use for all my little projects like this.
- adamkurkiewicz 3mo agoVery cool! We've used the following CLI to do the freeagent upload: https://github.com/anjor/freeagent-cli https://github.com/anjor/freeagent-cli How are you dealing with finding the receipts? Would you like to try a receipt finder that grabs them from your mailbox/ google drive?
- Havoc 3mo ago>My accountant hasn't complained yet They're not going to unless it's obviously and egregiously wrong - the risk on quality of input remains yours. It's the tax version of garbage in, garbage out. They're just guaranteeing the processing step.
- Shitty-kitty 3mo agoThe problem with LLM's is that they could work correctly for months and years and then do something egregious which will will go unnoticed because of the misplaced trust one develops on a system that "just seems to work." Get flagged for an expensive audit and there go all the savings and then some.
- petesergeant 3mo agoisn't that true for humans too?
- Shitty-kitty 3mo agoHumans are more likely to make small mistakes but the internal consistency check is pretty good at catching large errors. On top of that, fudging numbers to make everything add up is not something humans do (not unintentionally at least)
- aerhardt 3mo agoI'd be scared shitless to even try something like this. There is just a pretty website, a video, and a blog post. No info on the founders, I can't find anything on LinkedIn, just a company Vineyard Finance LTD that was incorporated last year. We're all unhinged about the data we're giving LLMs but here I'd draw the line. I'd rather keep paying the small amount I pay to have my accounts done.
- adamkurkiewicz 3mo agoInfo on the founders coming soon -- we're just going public with this. For slightly out of date founder bios (both Adam and Iva) were also co-founders here: https://www.biomage.net/our-team https://www.biomage.net/our-team
- phildenhoff 3mo agoThe company I work for, Digits, has been regularly updating our AI-vs-human bookkeeper benchmark. Look at page 8 -- many models are nearly as accurate as a human bookkeeper https://digits.com/downloads/beyond-the-hype-evaluating-llms-vs-digits-agl.pdf https://digits.com/downloads/beyond-the-hype-evaluating-llms...
- ivababukova 3mo agoThanks, that's useful! We did use ChatGPT 5.5 for some time and it did perform pretty well too (of course more expensive than GLM 5.2). We tried Claude 4.7 and 4.8, but we found both models to be "lazy" and very expensive. Claude would always rather prefer the route of saying that the evidence was not found or something is incomplete, rather than put more effort into finding/repairing the particular issue.
- senordevnyc 3mo agoI use Digits for my solo saas and it’s been great so far!
- cs702 3mo agoIt's not hard to imagine that will be able to do as good a job as a human accountant in the not too distant future. It's also not hard to imagine tax authorities using AI to audit everyone's tax returns every year. We sure live in interesting times.
- Obscurity4340 3mo agoThe real test to see if AI is just a rich person thing will be to see how the tax authorities treat it, even for more complex returns. They can save humans for the really complex edge case stuff but at the end of the day, the tax code is just checkboxes and input forms that get boiled down into Integers, Floats/Doubles and enumerated choices with some Strings for deductions
- adamkurkiewicz 3mo agoI've submitted my German taxes this year using a mix of Claude 4.6 and Claude 4.7, with lots of manual checking. The German Finanzamt granted most of the things I listed in the tax return (they send you an official letter by post) -- I did have to appeal for one of the items though (again using Claude, this time 4.8 ). The most important thing I've found is to ask Claude to thoroughly audit the reply (to find all hallucinations). I usually ask it to give me an enumerated list of all facts and all legal cases quoted, and then I give it to a new instance to carefully validate each one. Newer models are getting much better at not hallucinating German case law though :)
- markdown 3mo agoHave you considered using using two completely different models and comparing their output in order to catch hallucinations?
- quickthrowman 3mo agoBookkeeping is not tax preparation, just FYI. The reports generated by an accounting system are used on tax returns for companies, but they’re distinctly different things.
- arjie 3mo agoI just have a folder on my computer where I keep things in beancount. Then I have mercury CLI access with a read token to my business bank account, and I have my emails fully synced in there as well via IMAP. Claude Code with Opus just seamlessly hooks everything up so my accounts are up to date. At the end of the year, I used that information to prepare my tax returns for the business and then later the part that flowed to me as the owner. I had a fairly complicated tax return in 2025 involving a couple of change of business tax consideration and some money that was accidentally sent to me as a 1099 instead of to the business and I did everything with tax software with Claude Code advising. The end result was pretty damned good. I was unsurprisingly audited (or at least carefully reviewed) and the only error was in some way where I allocated a small amount of my wife's tax-free disability payments (disability is the mechanism that California uses to provide maternal benefits pay protection; she's not actually disabled). The IRS told me about it, I paid that bit (it was meant to be claimed back from the employer, not the US government) and everything was hunky dory. To be honest, the sum was so small I did not investigate (and haven't yet followed up with getting reimbursed by her employer). Honestly, almost all of it could have been avoided if I'd paid an accountant and a tax lawyer and they'd told me things and I'd done as they did, but in the end the combination of the fact that the IRS is very reasonable when you explain things and a modern agent means that the entire process was quite simple. In the end, I preferred the interactive mechanism of working with software because most accountants and lawyers will prefer to get all of your documentation all at once and then work on it rather than do it incrementally. In my case, I was able to work on the return incrementally and then have everything plugged in. I could ask a bunch of questions and get clarification. I think I will probably do all this the same way this year (though of course my taxes will be simpler).
- rgbrgb 3mo agoWhy do you think you were audited?
- arjie 3mo agoAs in, "what makes you think you were audited?". They sent me paperwork saying they were reviewing my return and that they found a discrepancy (the disability thing) and it took many months after the usual time for it to process. If you mean "for what reason could you have been audited?" it was because a client I had previously worked for accidentally reported paying me as an individual instead of my LLC.
- zerobees 3mo agoThis is a prime example of a problem space where accuracy matters, but it also matters who ultimately goes to prison. I'm going to go out on a limb and guess it's not the LLM. If you're acting in good faith and your accountant does something crazy or evil, your liability is limited to some extent. You may get a tax bill but you're probably not gonna end up behind bars. But if your LLM decides to do a little bit of tax fraud, you're in uncharted waters. In the end, the gun did it, but you were the one holding the gun. A lot of jobs are like that. You're not as much buying the service as you're buying not having to worry about the service.
- zitterbewegung 3mo agoI agree with you on every point but it is interesting to see real world benchmarks like this. Showing the standard benchmarks that all LLMs use is not only boring but at this point likely gamed or even has issues (according to OpenAI) by every LLM.
- andai 3mo ago>has issues (according to OpenAI) What is this referring to? I've been testing various big and small models for years and, about a year ago, switched my attitude from "biggest model is best!" to "best depends on task". For example, I had a simple coding task which required making 3 trivial changes in 3 source files. Biggest Model completed task perfectly and took 90 seconds. Smaller Brother also completed task perfectly, but took 30 seconds and cost 5x less.
- deleted 3mo ago[deleted]
- andy_ppp 3mo agoIt’s almost impossible to be in a situation where your taxes end up this wrong you’re accused of fraud. I honestly do not believe my accountant does anything better than AI or AI generated code that does calculations to see what you should be paying.
- 3mo ago
- Diogenesian 3mo agoThis shouldn't be ignored in the discussion here: The job performed by the humans was broader than what was requested of the model in this benchmark: humans also had to find the relevant invoices (searching through mailboxes, or requesting them from providers) and reason through any circumstances which cannot be inferred from the bank feed and invoices/receipts on their own. In the benchmark these circumstances are presented to the model as “user notes." This is precisely the kind of fine print on white-collar AI capability that companies keep running into: pretty much any non-entry office job worth having involves a lot of undocumented (even undocumentable) problems requiring judgment and experience. And I would be pretty nervous about asking any of the frontier LLMs to retrieve invoices: "cool, Claude logged that it found the May 6th bill from the paper supplier, I am sure it didn't just make something up arbitrary, then compound on the error by agentically iterating over the made-up invoice lurking in its reasoning traces. I checked the first 30 times and there were no problems!"
- adamkurkiewicz 3mo agoHey, the author of the benchmark here. The benchmark data was prepared in April 2026 (when I was manually doing our VAT return with my co-founder). The invoices were indeed found manually. Currently we're using a custom "invoice searcher" built on Kimi 2.6 (in our testing several weeks ago it outperformed Opus 4.7; it was just more persistent). Ultimately, I still verify everything manually after the model is finished fetching invoices for the month -- but it's a great help to have all the invoices already found (usually correctly).
- breadislove 3mo agoadam, i'd like to get in touch and would love to run the benachmark with mixedbread as a search backend. we are doing this right now with a lot of compliance companies. would be very curious how it improves quality/cost e2e
- adamkurkiewicz 3mo agoSure, I'm at adam@vineyard-finance.com
- Gander5739 3mo agoUnrelated, but I feel it's unfair to rob the word "bookkeeper" of its peculiarity of having three subsequent double letters by inserting a space in the middle.
- lowsong 3mo agoThat "nearly" is doing an awful lot of heavy lifting. It doesn't matter if your AI model is 99% or 99.99% accurate. For a tax return it has to be perfect every time or someone is at best getting a fine or at worst going to prison. Sure, human error happens too, but humans take accountability. That's why accountants are a regulated profession. Until an AI company CEO is willing to go prison if the output of their model is wrong, these tools are worthless. But don't take my word for it, head on over to Toot's own terms of service https://toot-books.pages.dev/terms#ai-not-advice https://toot-books.pages.dev/terms#ai-not-advice Toot uses automated and AI systems to generate classifications and reconciliation suggestions. Output may be incomplete or wrong and must be reviewed by you. Toot is a software tool. It does not provide accounting, tax, legal, audit, or financial advice, and nothing it produces is a substitute for a qualified accountant or tax adviser. You are responsible for checking Output before approving it or relying on it, and for any decision you make based on it. To the extent permitted by law, we are not responsible for outcomes arising from automated Output you approve without review. Comparing this to a human book keeper is farcical.
- rustcleaner 3mo agoWhat about errors resulting in overpaying taxes?
- senordevnyc 3mo agoThat "nearly" is doing an awful lot of heavy lifting. It doesn't matter if your AI model is 99% or 99.99% accurate. For a tax return it has to be perfect every time or someone is at best getting a fine or at worst going to prison. This is wildly out of touch with reality, at least in the US. The IRS barely audits anyone anymore, and if you make a mistake and they catch it, they often just correct it and send you a letter. You might pay a little in penalties, but it’s not that big of a deal. And I know literally dozens of small business owners who pretty blatantly dodge taxes in a hundred different ways and none of them have ever gotten so much as a slap on the wrist. And some of them have been audited!
- lowsong 3mo agoI don’t think “I know plenty of people who commit fraud anyway so who cares if it’s right” is the ringing endorsement of AI you think it is.
- krupan 3mo ago"The VAT return prepared by the model was essentially correct: the most important number in the return, which is how much VAT the company was owed by the tax agency, was off by only 7 pence relative to the human-prepared return." I don't know how taxes work in Europe, but in the US being "essentially" correct is not good enough for the IRS. The paragraph after this one goes on to explain other mistakes the LLM made? Yikes
- BeetleB 3mo ago> I don't know how taxes work in Europe, but in the US being "essentially" correct is not good enough for the IRS. It often is. On multiple years they've found errors in my return, and they just fix it for me and bill/refund the difference (usually a refund). And obviously, the IRS is not going to quibble over 7 pence, given that you round up/down to the nearest dollar anyway.
- markdown 3mo agoThe accounting concept of materiality doesn't exist in the US? What about petty cash?
- asdff 3mo agoAsk most people in the US how they file their taxes, most will probably tell you they clicked around turbotax until the biggest return number showed up. Is that indicating correct tax filing? I'm not sure. I would guess not.
- shh_labs 3mo agoThis is interesting to see, but surely any non-trivial business [edit: who needs to file a VAT return] is already entering its invoices into a finance system which can automatically generate a VAT return in a deterministic way.
- __MatrixMan__ 3mo agoI'm not so interested in having an LLM do my bookkeeping for me. But I'm very interested in whether LLM's can unravel the accounting obfuscations that billionaires use to avoid paying taxes on their wealth.
- rustcleaner 3mo agoOne idea I've had for reform is to ban central banks from inventing new currency except to pay every citizen a UBI debt free, and only tax limited liability entities (not humans). This would give citizens, instead of borrowers, first spend on new money and orient limited liability entities to serving citizens in the market to get their taxed revenues. It would also be a boon for financial privacy and reduce the burden on the average man.
- __MatrixMan__ 3mo agoI'm not so jazzed on state-issued UBI myself. It makes the state feel like they're special and I think that leads to bad behavior. We should use money issued by the most trustworthy entity around, whatever it may be at the time, and the state should have to compete for that spot.
- fragmede 3mo agoCorporations, government, or billionaires. Those are the three options, pick one.
- __MatrixMan__ 3mo agoThere's also your friends and family and neighbors--anyone that you know in meatspace and maybe have a reason to trust explicitly.
- luciana1u 3mo ago[flagged]
- josefritzishere 3mo agoI can't file "almost" with the IRS. That's not going to end well. Sorry Mr. auditor, this is a non-deterministic filing.
- fibers 3mo agothats not true. maybe for simple 1040s yes, but depending on what you have you could take aggressive positions bc they could be subject to valuation differences.
- petercooper 3mo agoI wouldn't be surprised if it were more accurate based on the errors I've seen. I always eyeball the books and was confused when a £15k building popped up on our asset sheet. It turns out a "workshop" had been categorised as a building we had purchased, rather than the training session it actually was. This is the importance of having layers and multiple sets of eyes on things, though. Even if it had got past me, my accountant would have surely queried it at year end, but that could be true of an LLM mistake too.
- xiaodai 3mo agoi suspect the key word here is nearly.
- Fno44 3mo ago[flagged]
- deleted 3mo ago[deleted]
- VerityLayer 3mo ago[flagged]
- snthpy 3mo agoI was wondering what a book keeper is, thinking it would be some kind of librarian. Isn't the accounting sense written as a single word, i.e. bookkeeper?
- classified 3mo agoSo it nearly didn't cause bankruptcy? That's nearly great. Don't get me wrong, I'm all for open-weight models, but you still need to be careful what you use LLMs for.
- ryss20 3mo ago[flagged]
- josefrichter 3mo agoAs someone who practised tax law for global companies for a few years, I can say that the correct classification of some invoices from the tax law perspective requires a full-blown legal analysis. That involves full interpretation of the law, which includes "teleology" - that is, understanding the purpose and goals of the law. We are probably (?) not yet at the point where frontier models can do this better than humans. In other words, it's the same question as when AI will replace judges. Not lawyers, but judges themselves.