---
title: The real reason AI fails in accounting
source: https://www.tech-japan.jp/blog/accounting-ai-failure/
updated: 2026-08-29
published: 2025-02-20
facts: https://www.tech-japan.jp/facts.json
---
# The real reason AI fails in accounting

Accounting looks like the ideal first target for automation. It is rule-bound, high volume, and repetitive. It is also where a surprising number of AI projects quietly stop, and the reason is not model accuracy.

## The pitch, and where it breaks

The pitch is straightforward: feed the system your historical journal entries, let it learn the patterns, and have it code new transactions. Vendors demonstrate ninety-something percent accuracy and the demo is genuinely impressive.

Then it meets a real ledger, and accuracy on the entries that matter collapses. The entries it gets right were the ones a rules engine could already have handled. The ones it gets wrong are the ones a human spent thirty seconds thinking about — and thirty seconds of thinking is exactly what was never recorded.

## The context was never written down

Consider a payment of 120,000 yen to a company called Yamada Kogyo. Which account does it belong in?

You cannot answer that from the transaction. It depends on what was purchased, whether it was capitalized or expensed, which project it belongs to, whether it was a deposit against a larger contract, and whether last quarter’s identical-looking payment was coded the way it was for a reason or by mistake.

A human in the accounting department answers it by knowing things: which project the invoice number belongs to, that this supplier does two different kinds of work, that the controller decided in March that this category is capitalized. **None of that is in the data.** It is in a person, and it is being reconstructed from scratch each time.

So the model is asked to infer, from amount and payee, a decision that was actually made from context it has never seen. It will produce an answer. The answer will be confident. It will be right about as often as the context happened to be inferable, which is a number nobody can predict in advance.

## What actually works

### Capture the reason at the moment of the decision

The single highest-value change is usually not an AI feature. It is adding structure at the point of entry — a project code, a contract reference, a capitalization flag — so the reason is recorded alongside the transaction. That is unglamorous, it takes a quarter, and it is what makes everything downstream possible.

### Encode the rules that are actually rules

A large fraction of coding decisions follow a deterministic rule that nobody has written down. Write them down. A rule you can read is auditable, correctable, and free to run; a model that has learned the same rule statistically is none of those things.

### Point the model at the residue

What remains after the rules are explicit is genuinely ambiguous work, and that is where a model earns its place — ranking candidate accounts, surfacing the three prior entries that most resemble this one, and flagging the transaction that does not resemble anything. Assistance on judgment, rather than a claim to have replaced it.

> A useful test: ask whether a new hire in the accounting department, given only the data the model receives, could code the entry correctly. If they could not, the model cannot either, and the problem is upstream.
