Can Benford's Law Catch Fraud in Your Ledger? We Tested It
Real money has a fingerprint: ones are common, nines are rare. We invented two companies and checked whether the fingerprint could tell which one we'd been careless with.
The fingerprint of real money
In 1881 the astronomer Simon Newcomb noticed that the early pages of his book of logarithm tables, the ones for numbers starting with 1, were far more worn than the later ones. In 1938 the physicist Frank Benford tested the idea on 20,229 numbers, from river areas to street addresses, and found the same thing everywhere: in naturally occurring data, the leading digit is 1 about 30% of the time, and 9 less than 5% of the time.
The reason is scale. Money grows by percentages, not by fixed steps, so an amount spends far longer working its way from £1,000 to £2,000 (a doubling) than from £9,000 to £10,000 (an 11% rise). Across thousands of transactions, that tilt shows up as a predictable curve. Invented numbers rarely follow it, because people inventing figures spread them evenly, avoid repeats, and lean on the digits that feel random.
Quick answer: Benford's law says that in real financial data about 30% of amounts start with 1 and under 5% start with 9. Testing a ledger against that pattern can flag amounts that were invented, capped or bent around a limit, but it can't prove fraud, and a careful fabricator can pass it. In our test it caught the company whose numbers we generated carelessly, and half the reason turned out to be ordinary payroll.
How conformity is measured
The usual measure is the mean absolute deviation (MAD): the average gap between how often each digit actually leads and how often Benford says it should. Mark Nigrini, whose work brought Benford into forensic accounting, set the bands most auditors use.
| Verdict | First digit MAD | First two digits MAD |
|---|---|---|
| Close conformity | up to 0.006 | up to 0.0012 |
| Acceptable | 0.006 to 0.012 | 0.0012 to 0.0018 |
| Marginal | 0.012 to 0.015 | 0.0018 to 0.0022 |
| Nonconformity | over 0.015 | over 0.0022 |
The first-two-digits test (10 to 99) is the sharper tool on a large ledger. It is the one that picks up a spike at 48 and 49, which is what it looks like when somebody keeps claims just under a £500 sign-off limit.
The experiment: two companies we invented
We build demo companies to test our own software, which gave us an unusual opportunity: two sets of books where we know exactly how every number was made.
Brackenfield Outdoor is a garden and tool supplier with two years of books. Its generator was written to pass this test: wherever an amount spans more than a decade, it is drawn on a logarithmic scale across whole decades, precisely so the leading digits come out the way real ones do. We took its 5,291 posted journal lines of £10 or more.
Airedale Joinery is a simulated joinery business we use to test bank feeds. Its simulator wasn't written with Benford in mind. Most amounts are drawn evenly between a lower and an upper limit: a fuel stop somewhere between £38 and £96, a timber order between £180 and £2,400. We took its 717 bank lines of £10 or more.
Brackenfield
Acceptable, MAD 0.0073Ones slightly over, nines slightly under, everything else close. On the first two digits it scored 0.0020, just into marginal. LedgerIQ, run on its own export of the same books, gave the same verdict on the first two digits.
Airedale
Nonconformity, MAD 0.044Nearly three times past the nonconformity line. Too few ones and twos, and a striking bulge at 4 and 5: 17.3% and 19.5% against an expected 9.7% and 7.9%.
So the test did its job. It told the careful fake from the careless one, at a glance, without knowing anything about either business.
Investigating the spike
A failed Benford test is the start of an investigation, not the end of one. So we did what an examiner would do and asked which transactions made the 4s and 5s.
The first answer was innocent. Airedale pays four staff every week, between about £430 and £610 each. That is a narrow band, and every one of those payments starts with 4, 5 or 6. Payroll, rent, subscriptions and anything else confined to a tight range are well-known exceptions to Benford, because nothing about them grows by percentages.
The second answer was the guilty one, and it was us. With the wages taken out, the remaining 561 lines still scored a MAD of 0.0152, just over the nonconformity line: too few ones, too many sevens, eights and nines. That flatness is what you get when amounts are drawn evenly between fixed limits, which is exactly what our simulator did, and exactly the habit Benford catches in a person inventing figures.
The lesson in one line: Benford was right that something was wrong, and it could not tell us which part was innocent. Only looking at the transactions could.
What Benford can't do
- Prove fraud. It shows that amounts don't behave like natural data. Payroll, price lists, fixed fees and VAT-inclusive round prices all fail innocently.
- Work on small samples. A few hundred transactions can drift from the curve by chance. The test needs a ledger of real size.
- Catch a careful fraudster. Brackenfield passes because we built it to. Anyone who knows the test can do the same.
- Read assigned numbers. Invoice numbers, account codes and phone numbers have no reason to follow it.
Used properly, it is a fast first pass that tells you where to look. The useful work is what comes next: duplicate payments, round amounts, entries just under approval limits, postings at weekends and by unusual users, and journals that hit unusual account pairs.
Running it on your own ledger
You can do the basic test in a spreadsheet: take the first digit of every amount, count them, and compare with the expected shares. It takes an afternoon to set up properly, and the follow-up tests take longer.
LedgerIQ runs the whole sequence from a general ledger export from any accounting platform. It uses the first-two-digits test once a population is large enough, gives the Nigrini verdict, and runs it alongside duplicate, round-number, weekend-posting and threshold tests and anomaly detection, then sets out what it found in a working paper. A full analysis costs 1,000 credits and every account starts with 1,000 free. For charity and independent examination work, our journal entry testing toolkit sets out the full procedure.
Test a real ledger
Upload a general ledger export from any platform. LedgerIQ runs Benford's law alongside duplicate, round-number and weekend tests and anomaly detection, and sets out what it found.
See LedgerIQFrequently Asked Questions
Benford's law describes how often each digit appears as the leading digit in naturally occurring numbers. In financial data, about 30.1% of amounts start with 1, 17.6% with 2, and the share falls steadily to 4.6% for 9.
It can flag amounts that don't behave like natural data, which is what invented, capped or manipulated figures often look like. It can't prove fraud: payroll, price lists and fixed fees fail innocently, and a careful fabricator can pass it. Treat a failure as a reason to look at the transactions.
Enough that chance doesn't swamp the pattern. A few hundred transactions can drift from the curve for no reason, so the test is most useful on a full year or more of a ledger, and the first-two-digits test on larger sets still.
On Mark Nigrini's widely used bands, a first-digit MAD up to 0.006 is close conformity, 0.006 to 0.012 acceptable, 0.012 to 0.015 marginal and over 0.015 nonconformity. For the first two digits the bands are 0.0012, 0.0018 and 0.0022.
Amounts confined to a narrow range or set by a price list, such as wages, rent, subscriptions and fixed fees, and very small amounts. Assigned numbers such as invoice numbers should never be included.
Yes. Extract the first digit of each amount, count each digit, convert the counts to shares and compare them with the expected Benford shares. LedgerIQ does the same from a general ledger export, along with the follow-up tests.