I link LLM AI to game localization, and prove it with eval. 14 years in game studio localization, 8 of them inside the publisher.
Existing MT metrics cannot see what breaks in game text. In two of the three games I measured, the human-written golden scored worse than a plain machine draft on a generic metric. So I build the evaluation per house, in the client's own terms, and use a high-low model mix so the cost lands where it should.
Cross-posted from LinkedIn, archived here.
Not a screenshot. An interactive table you can sort and inspect, with the numbers baked in at publish time.
Real results land here after the first measured run.
No placeholder numbers go on this page.
Email goes live once Workspace is wired to the domain.