Transaction ML limitations

Where the categorizer is weak: unseen brand names, missing amounts, ambiguous merchants, and the gap between synthetic and real statements.

What it does

The categorizer is a demonstration of a method on invented data. It shows how a transparent text model behaves on merchants it has never seen, how to measure that honestly, and how confidence can be used to decide when to ask a person.

It is weakest on well-known brand names it has not seen, because a brand name often says nothing about what the brand sells. Local businesses are easy for the opposite reason: their names usually contain the category ("... family dental", "... auto repair").

Why it is used

Every model has a region where it should not be trusted. Publishing where that region is, with numbers, is more useful than a single headline accuracy.

Inputs

  • The same test merchants as the evaluation page, grouped by merchant type.
  • The model's accuracy with and without the amount.
  • Coverage and accuracy at confidence thresholds from 30% to 99%.

Assumptions

  • A person reviewing low-confidence predictions is available and correct. Coverage figures describe how much work is left for them.
  • The descriptor is the only text. Real systems also use the merchant category code (MCC) sent by the card network, which settles most of these cases and is not modeled here.

How to read the results

The first table shows accuracy by merchant type; the gap between brands and local businesses is the main weakness. The coverage table reads as a trade-off: at a higher confidence threshold, fewer rows are categorized automatically, and those that are are more often right.

Limitations

  • Synthetic data. The model has never seen a real bank statement, and real descriptors are messier and far more varied.
  • Unseen brands are often wrong. The model can only use the characters in the name, and many brand names carry no category information.
  • One category per transaction. A supercenter receipt that is half groceries, half household goods gets one label.
  • No personalization. People disagree on categories (is a coffee subscription dining or a subscription?); the model learns the generator's labels.
  • US English descriptors only.
  • Educational only. Do not use it for tax, accounting or credit decisions.

Where it can fail

  • Descriptors that are mostly reference numbers or processor boilerplate.
  • Names that collide with another category's vocabulary.
  • Amounts far outside the training ranges (a $40,000 grocery bill) push the amount flags into rarely seen bins.
  • Very long or unusual input is rejected rather than guessed: over 200 characters, or text with no letters or digits, returns an error.

Validation on current data

Accuracy by merchant type, the effect of a missing amount, and the coverage-accuracy trade-off, on the test merchants.

References

  • Chow, C. K. (1970). On optimum recognition error and reject tradeoff. IEEE Transactions on Information Theory 16(1).
  • Mitchell, M. et al. (2019). Model cards for model reporting. FAT* Conference.

Try it in the categorizer