A Small AI Beat GPT-4 at Sorting Court Documents
A small AI beat GPT-4 at sorting legal papers — and it costs far less.
Legal paperwork is a nightmare to sort automatically. Cases look almost identical on the page, but belong in completely different categories, and getting it wrong has real consequences. So a researcher compared three families of AI on ten categories of Korean sexual offense rulings: old-school machine learning, general-purpose chatbots like GPT-3.5 and GPT-4, and small models given extra training on legal language.
The surprise: the small, specialized model won. KLUE-BERT hit 99.3% accuracy, beating the famous, much larger GPT models. The lesson is that for narrow, expert tasks, training on the right material matters more than raw size. It's the difference between hiring a famous generalist and hiring someone who has read nothing but case law for a year. Smaller models are also far cheaper to run, which matters if a court system processes thousands of documents a day.
But here's the catch. The researcher used explainable AI — tools that show why a model made a decision — and found it leans on obvious word clues and misses subtle context. When tested on records that look like genuine case files, with details implied rather than spelled out, accuracy dropped. So that 99.3% is real, but it's a lab number, not a guarantee.
None of this means robots replacing judges. It means AI as a librarian and first-pass assistant: sorting files, pulling up similar past cases, flagging documents for review. For the rest of us, faster document handling could mean shorter case backlogs and cheaper legal research. The honest risk is over-trust — treating a confident-sounding AI as correct when it has missed the context a human would catch. Sensitive case records also raise obvious privacy questions about who holds that data.
- A small AI trained specifically on Korean legal text sorted court documents with 99.3% accuracy — beating GPT-3.5 and GPT-4.
- Specialized training mattered more than model size, and small models are much cheaper to run at scale.
- 'Explainable AI' showed the model relies on obvious word clues and struggles with implied context in realistic case files.
Why It Matters
Faster, cheaper sorting of legal documents could speed up case backlogs — if humans keep checking the AI's mistakes.