Back to all posts
Article

Best AI Translation Models in 2026: Quality, Terminology, and API Cost

Author:

Compare GPT, Gemini, Claude, DeepSeek, HY-MT2, DeepL, and Google Translation for text workflows. Budget for translation and review without confusing tokens with characters.

An open book with aligned abstract text lines on facing pages and a small tab on its spine

“Do not delete the backup” is a short sentence. A fluent translation that loses “not” is still a failed translation. For product copy and documentation, choosing a model means balancing meaning, terminology, formatting, and the time someone will spend correcting the output.

Our budget specialist pick to test is HY-MT2-7B through the Tencent Cloud endpoint listed on OpenRouter, for text that fits its 8K context. GPT-5.6 Luna and DeepSeek V4 Flash are our general-model alternatives. For a dedicated translation workflow, we would also compare DeepL's quality-optimized option and Google Cloud Translation rather than assume a general chatbot is always better. None is an all-language accuracy winner established by this article.

Documentation and prices were checked on September 4, 2026. This is a public-source comparison with worked budgeting examples, not a multi-model translation experiment. The practical examples focus on English–Chinese product content. Speech translation, interpreting, and translation-app subscriptions are outside the comparison.

Start by choosing the kind of service

OptionWhy it belongs on a shortlistWhat needs checking
GPT-5.6 LunaLow listed token cost for a first translation passMeaning, placeholders, and actual billed output
DeepSeek V4 FlashLow-cost general model with cheaper off-peak usageTiming, reasoning configuration, and review burden
Gemini 3.8 FlashAlternative within Google's general-model APIIntroductory price and full-run token usage
Claude Sonnet 5A higher-priced comparison for instructions and revisionWhether it actually reduces editing on your content
HY-MT2-7BLow-priced hosted specialist or downloadable weightsHosted length limits, deployment cost, and language pair
DeepL quality-optimizedTranslation-specific controls and glossary workflowLanguage support, plan, and requested model behavior
Google Cloud TranslationNMT and Translation LLM endpoints under a translation APIWhich endpoint is used and which characters are billed

These are different delivery models. A downloaded model does not come with a free hosted service. DeepL's API option is not a public weight checkpoint. Google Cloud Translation is not the same endpoint or billing scheme as Gemini. Keep those distinctions when comparing costs.

What token-billed translation costs

All figures below are USD per million tokens, using standard uncached text input. The last column assumes 2 million input and 2 million billed output tokens, spread across requests within the quoted tier. This is an independent volume assumption, not a conversion from a fixed number of pages or Chinese characters.

ModelInput / 1MOutput / 1MExample cost
HY-MT2-7B · Tencent Cloud via OpenRouter$0.074$0.295$0.738
GPT-5.6 Luna$0.20$1.20$2.80
DeepSeek V4 Flash, peak$0.44$1.32$3.52
Gemini 3.8 Flash$0.75$3.75$9
Claude Sonnet 5$2$10$24

Luna uses the short-context tier. DeepSeek's example falls to $1.76 when all requests qualify for off-peak pricing; weekday peak windows are 01:00–04:00 and 06:00–10:00 UTC. Gemini's introductory rate ends December 31, 2026, after which the announced rate doubles the example to $18. Tax, additional tools, retries, and human review are excluded.

The table explains our low-cost starting shortlist; it does not prove translation quality. A model that ignores a terminology instruction can erase the entire API saving with a few minutes of correction. Likewise, sending an entire manual just to translate one button can cost more than sending the button with the few sentences that explain it.

Dedicated translation changes the calculation

DeepL: a glossary is a workflow feature, not a quality guarantee

DeepL's text API exposes quality- and latency-oriented model options, context, and glossary controls. Its context parameter is not billed, while a general LLM usually counts supplied context as input tokens. A glossary must match the requested language pair, and using it requires the source language to be specified.

That makes DeepL a useful comparison when short strings depend on surrounding product information. Check support for the exact language and option: a control available for one target language is not automatically available for another. Use the rate in your API plan rather than importing a consumer subscription price or an unrelated region's quote.

We would choose this style of service when an established glossary workflow saves integration and review effort. We would not call it the most accurate English–Chinese translator without a test on the actual material.

Google Cloud Translation: characters, not tokens

Google's translation pricing lists NMT at $20 per million input characters after the applicable monthly allowance. Its standard Translation LLM text endpoint charges $10 per million input characters and $10 per million output characters. Adaptive translation and formatted documents have different prices.

For a hypothetical one million input characters and one million output characters, the Translation LLM usage is $20. NMT usage is also $20 before applying its eligible credit. That is not the same workload as two million input tokens and two million output tokens in the previous table.

Google counts Unicode code points, including submitted whitespace and untranslated characters. A ten-page file is not a billing unit you can compare directly with an LLM's token price. Check the endpoint and count the submitted content before making a budget.

HY-MT2: a cheap hosted specialist, or your own deployment

OpenRouter lists a Tencent Cloud endpoint at the rates above, with an 8,192-token context and up to 4,096 completion tokens. That makes HY-MT2-7B worth testing for short, repeated translation jobs. A long manual must be divided carefully; glossary and context overhead still count. The example excludes any applicable platform charges and is not a quote for self-hosting.

Tencent's HY-MT2-7B model card provides a dedicated translation model and examples for terminology, style, and structured content. It is worth evaluating when deployment control matters and the team can maintain inference.

The model card's own comparisons are provider-reported evidence, not an independent verdict on your documents. Include compute, idle capacity, upgrades, and review in the budget. Smaller or quantized variants also need their own checks; a result from another checkpoint does not transfer automatically.

How we would decide whether a translation is usable

WMT's translation evaluations are organized around specified tasks and language pairs. The MT Metrics Eval toolkit provides human and automatic scores, including MQM ratings for some language pairs. Neither justifies combining different years, languages, and model versions into a single universal ranking.

For a product team, a smaller acceptance check is often more useful than an unsupported headline score. Consider this invented English UI string:

Do not delete the backup until the export is complete. Your trial ends on September 30. Contact {support_email} for help.

Give each candidate the same glossary and product context. A usable Chinese translation must preserve the prohibition, the order of events, the date, and the exact placeholder. It must not turn “trial ends” into “your subscription renews.” That would invent a billing event not present in the source.

Then check a paragraph with repeated product terms and a structured file containing keys, links, and visible strings. Parse the result, compare placeholders, and ask a bilingual reviewer to inspect meaning. Grammar, typography, and preferred style matter, but do not let them hide a changed amount or missing condition.

Ask the reviewer to record critical meaning errors separately from minor edits. Hide model names where practical. A single attractive example is not enough; keep failed and awkward results in the review set rather than selecting only the best output.

Budget for revision before choosing a winner

Use AILesson's model pricing calculator to compare the token-billed options: 2 in input and 2 in output reproduces the assumed volume, using the corresponding catalog rates. Match the hosted endpoint, not just the model name. It does not estimate DeepL's plan fees, Google Translation's character bill, or the cost of self-hosting HY-MT2.

Start with a sample and read actual token usage before extrapolating to all documents. Different tokenizers and translation directions change the count. If a second model reviews the translation, its input usually includes both the source and first output, plus instructions; review is not a free output-only step.

At the example volume, Sonnet costs $21.20 more than Luna. With an assumed reviewer rate of $30 per hour, 42.4 minutes saved across the whole batch would cover that difference. This is a break-even calculation, not a measured Sonnet advantage. The cheaper candidate remains the better purchase if both need the same review.

For low-stakes short drafts, compare hosted HY-MT2 with one budget general model. For public product text, compare a dedicated service and a general model on the same glossary and error checklist. Keep human approval for consequential meaning. Choose the lowest total cost that preserves what the original actually says—not the output that merely sounds most polished.

References