New AI Can Grade Translations Without a Human Cheat Sheet
This could make the translated websites, apps, and manuals you read far more accurate.
Machine translation has gotten remarkably good, but someone still has to check the output. Normally that check is slow and expensive: a human translator creates a "reference" translation, and software compares the machine's version against it. FACET, a system from researchers Ahrii Kim, Chanjun Park, and Seong-heum Kim, skips that entirely. It judges a translation on its own merits, which means quality checks could happen automatically, instantly, and at almost no cost.
The clever part is how it splits the work. Different translation mistakes need different evidence. Whether the meaning was preserved can only be judged by looking at the original sentence. But whether the result reads naturally, or whether a company name is spelled the same way every time, can be judged by looking at the translated text alone. FACET runs three passes — Fluency, Accuracy, and Consistency — and gives each pass only the information it actually needs. One fixed AI model is simply prompted three times, with no special training involved, which keeps the whole thing cheap to run.
There is a companion version, FACET-C, that drops the Consistency check. The researchers found that check nudged about one in ten sentence scores, but barely changed the overall ranking of competing translation systems. In other words, consistency matters for individual sentences more than for the big picture.
When the team ranked real translation systems, professionally edited human translation came out on top — a reassuring sign the tool recognizes quality. The honest catch: the researchers couldn't fully verify their scores because they had no human ratings to check against. It's a research result from a competition, not a product you can use today. But it points toward automatic translation checks that make global content faster and cheaper to produce.
- FACET grades translations without needing a human-made "correct" version to compare against, which makes checking fast and cheap.
- It asks three separate questions — is it fluent, is the meaning right, are names consistent — instead of one vague quality score.
- The team found consistency checks changed about 10% of individual sentence scores but barely shifted the overall rankings.
Why It Matters
Cheaper, faster quality checks mean better translated apps, manuals, and websites reach you sooner.