AI Translation and Content Localization - Running Multilingual Content Without a Translation Agency
RETURN_TO_BLOG
AI & Automation 14 min

AI Translation and Content Localization - Running Multilingual Content Without a Translation Agency

Paweł Wiszniewski
Paweł Wiszniewski
SEO & GEO Specialist · AI Engineer

The human evaluation results from the WMT25 competition (Workshop on Machine Translation), published in late 2025 following earlier automatic-metric results from August, showed something that would have sounded exaggerated a few years back: on most of the 30 evaluated language pairs, general-purpose models - Gemini 2.5 Pro, Claude 4 and the top commercial engines - beat classic machine translation in human linguists' judgment. That doesn't mean DeepL has lost its reason to exist, though: in independent blind tests on European language pairs it still wins most comparisons against general-purpose models, while for Asian languages (Chinese, Japanese, Korean) and Arabic the LLM advantage is clear. The real question isn't "which engine translates better", it's "which engine for which job" - and that's what this post is about.

The human evaluation from the WMT25 competition showed that Gemini 2.5 Pro, Claude 4 and the top commercial engines beat classic machine translation on most of the 30 evaluated language pairs - yet DeepL still wins most blind tests on European language pairs. The real question isn't "which engine translates better", it's "which engine for which job". I cover when to reach for an LLM versus DeepL, how not to lose brand terminology across languages, how localization differs from translation, what it costs per 1,000 words - and how I run this exact blog in Polish and English without a translation agency.

Running content in several languages without a translation agency is technically easier today than ever, but ease of generating text doesn't solve three real problems. Terminology drifts between publications unless someone systematically enforces it. Literal translation differs from localization, which accounts for currency, date format and local regulations. And cost and quality depend on whether you pick the tool for the content type, or the other way around. Here's how to build a process that keeps all of this under control - ending with how I run this exact blog in Polish and English.

Machine translation (NMT) or LLM - when to use which

/// WHICH ENGINE FOR WHICH JOB

High volume, European languages
→ DeepL (NMT)
Asian languages and Arabic
→ LLM (Gemini/GPT/Claude)
Marketing content, brand voice
→ LLM with a style prompt
Technical documentation
→ GPT-4o/5
Very long, multi-file context
→ Gemini

DeepL is a neural machine translation (NMT) engine trained for exactly one task: translating text. That's the source of its edge - in European languages, where it has the deepest training data, it wins most comparisons against general-purpose models in independent blind tests. It's also cheap and fast, because you're paying for a narrowly defined operation, not a general-purpose language model. The trade-off is inflexibility: you can't instruct DeepL to "translate this in a casual tone, like talking to a friend" - you get one deterministic output.

An LLM (GPT-4o/5, Claude, Gemini) wins wherever more than sentence-level correctness matters. It holds the context of an entire document, so it won't lose the thread between paragraphs. You can steer it with a style prompt to preserve brand voice. And it wins decisively for non-European languages - for Chinese, Japanese, Korean, Arabic and Hindi, independent tests consistently put LLMs ahead of classic MT. Benchmarks also reveal a split within the LLM category itself: Claude tends to handle nuance-heavy marketing copy best, GPT-4o/5 does well with technical documentation, and Gemini holds up best for consistency across very long context (translating dozens of files at once, for instance). Treat this as a shortlist signal, not a verdict - the differences depend on language pair, domain and evaluation method, so it's always worth validating on your own content sample.

A workflow that doesn't lose brand terminology

/// A WORKFLOW THAT DOESN’T LOSE TERMINOLOGY

A prompt alone gets you probabilistic compliance, not a guarantee

01
GLOSSARY INTO THE DRAFT
Approved terms and "do not translate" flags feed the prompt generating the first draft
02
AUTOMATED CHECK
The output is checked against that same term list before it reaches a human
03
TRANSLATION MEMORY (TM)
Identical segments are reused, similar ones billed by match percentage
04
HUMAN REVIEW WITH THE GLOSSARY OPEN
The reviewer fixes nuance, not rebuild terminology from scratch

Telling a model in a prompt "use these terms" gets you probabilistic compliance, not deterministic accuracy - the model will follow it most of the time, but not every time, and at volume "most of the time" means dozens of inconsistencies. An LLM can sound perfectly fluent while drifting on product naming: it might translate a feature as one word in one batch and a different synonym in the next, breaking UI parity and confusing a user reading documentation in two languages at once.

The fix is structure, not another prompt. A glossary - a list of approved terms with their translations and a "do not translate" flag for product proper nouns - gets fed directly into the drafting step, and once a draft is produced, the output is automatically checked against that same list before it ever reaches a human. Translation memory (TM), in turn, is a database of previously translated segments: a segment identical to one already translated is reused rather than retranslated, and a similar segment (a fuzzy match) is billed at a percentage of the full rate - typically 40-75%, depending on the match level. A side effect worth its own section: this is also the single biggest cost lever in the whole process.

Localization is not translation

/// TRANSLATION CHANGES WORDS, LOCALIZATION CHANGES THE WHOLE EXPERIENCE

04/03 vs 03.04
the same date written in US and PL format reads as two different days
date format
$1,234.56 vs 1 234,56 zł
the decimal separator and currency symbol position differ by market
currency and numbers
imperial vs metric
units of measurement need conversion, not just a translated label
units
sales tax vs VAT
different tax regimes and cookie-consent patterns per market
legal requirements

Translation changes the words. Localization adapts the whole experience - at levels that are easy to forget precisely because they're invisible until they fail. Currency isn't just a symbol swap: the conversion rate has to be current, and the notation itself (1,234.56 USD versus 1 234,56 zł) differs by market. Date format is the classic trap - 03/04/2026 in the US and 04.03.2026 in Poland are two different dates written with the same digits. Units of measurement (metric versus imperial), the decimal separator (comma versus period), and legal requirements (VAT in the EU, sales tax in the US, different cookie-consent patterns) form another layer that no translation engine resolves on its own - because it's not a language problem, it's a product and legal one.

OptionPrice levelCharacterBest for
DeepL (NMT)~$5.49 per million characters (~150-200k words)purpose-built for translation, strong on European pairs, fast, low control over toneUI strings, product listings, high-volume European-language content
LLM (GPT-4o / Claude / Gemini)token-based billing, in practice roughly single-to-low-double-digit dollars per 1,000 words depending on model and directionholds document-level context and brand voice, best for non-European languages and where nuance mattersmarketing copy, blog content, technical docs where tone matters
Human translator (full translation or MTPE)typical agency rate $0.05-0.15 per wordthe only option for legal text and final quality controlcontracts, regulated content, a final pass on public-facing pages
Own translation memory + glossarynear-zero marginal cost on repeat content thanks to 30-70% savings from fuzzy/exact matchesrequires upfront setup and maintenanceteams publishing regularly in the same subject domain

Our own PL/EN pipeline - a case study from this blog

This blog is exactly that case: every post exists in Polish and English, with no translation agency involved. A few rules that matter most for the final quality in practice. First, content in each language is written natively rather than translated sentence by sentence - in the AI/SEO niche, terms like "AI visibility" or "zero-click" don't have one natural equivalent in the other language, so a literal rendering reads as stiff, while rewriting from scratch reads as native. Second, internal links are rewired per language - the English post links to the English slug of the target post, not to the Polish URL with a note that it's a translation. Third, meta title and meta description are written separately per language rather than translated, because search intent and keyword phrasing differ between PL and EN queries even for the same topic. Fourth, a fixed, automatically checked style rulebook applies - among other things, no em/en dashes in favor of a plain hyphen, and a list of banned, generic-sounding stock phrases in both languages - before a post goes live. The result: both language versions read as if written natively in that language from the start, at the cost of more editorial work than dropping text into DeepL and pasting the output back.

The traps vendors mention more quietly

  • Fluent and wrong at once. A model can produce a grammatically perfect sentence with an inverted meaning (a dropped negation, for instance) - a native reviewer catches it in a second, a spell-checker never does.
  • Forgotten legal pages. Privacy policies, terms of service and cookie banners often fall out of a localization project because they're treated as technical debt, not marketing content - and they're exactly the pages carrying the most legal exposure if something's wrong.
  • No glossary means brand drift. Without an approved term list, an LLM will pick a different synonym for the same product feature in every batch, breaking interface consistency.
  • Automation with no human in the loop on public pages. Fine for internal documentation. Risky for content that represents the brand externally without a review pass.
  • Forgotten hreflang after adding a language. A new language version without correctly wired hreflang tags (covered separately in hreflang architecture) means a search engine may serve the wrong version to the wrong audience - all the translation effort never converts into visibility.

How to choose - four scenarios

  1. 1.High volume, European languages, tight budget → DeepL plus your own translation memory.
  2. 2.Marketing content, tone matters, non-European languages → an LLM with a style prompt and a glossary wired into the pipeline, not just the instructions.
  3. 3.Legal and regulated content → human translator only; the machine at most produces a first draft for review, never the final version.
  4. 4.You publish regularly in the same niche → build a glossary and translation memory from day one - the cost pays back by the second or third large batch through match discounts.

A step-by-step implementation plan

  1. 1.Inventory your content by type (marketing, legal, UI, blog) and assign each type an engine from the table above.
  2. 2.Build a glossary of 30-50 key brand and product terms before translating a single sentence.
  3. 3.Pick a tool that supports glossaries and translation memory natively, not only through prompt text.
  4. 4.Decide who does human review and on which content types it's mandatory versus optional.
  5. 5.Plan localization separately from translation - map currency, date, units and legal requirements per market before you start translating.
  6. 6.Measure cost and time on the first batch, then compare against the second - the drop from translation memory should already show on the second publication in the same domain.
  7. 7.Wire multilingual content into a correct hreflang architecture, so the translation effort actually converts into search visibility in each market.

What translation and localization cost per 1,000 words

/// COST PER 1,000 WORDS, BY OPTION

Human translator ($0.05-0.15/word)$50-150
LLM (GPT-4o/Claude/Gemini, per token)single-to-low-double-digit $
DeepL API (~$5.49/1M characters)single-digit $
TM - repeated segment30-70% off

* Bar length = cost position relative to the most expensive option, not an exact numeric ratio (different billing units: per word, per token, per character).

The range is wide because billing models differ. Raw translation APIs (DeepL, Google Cloud Translation, LLM-based engines) typically fall between single cents and a few dollars per 1,000 words, depending on provider and language direction - DeepL bills per character (roughly $5.49 per million characters), while LLM providers bill per token, which for typical English text lands in a comparable order of magnitude. A human translator prices differently: $0.05-0.15 per word, or $50-150 per 1,000 words - tens of times more than the machine alone, which is why the MTPE model (machine translation post-editing) is popular: the machine produces a first pass, and a human edits it for a lower rate than a full translation from scratch. A third factor, easy to miss in a spreadsheet: with regular publishing in the same subject domain, translation memory cuts the cost of repeat content by 30-70%, so the real cost per post drops with every month of work in the same topic - provided someone actually maintains that memory instead of starting from zero on every job.

---

I help build a translation and localization process that doesn't lose brand terminology across languages - from choosing the right engine, through glossary and translation memory, to wiring it into a correct hreflang architecture. I do this as part of AI automation and content marketing SEO. I teach it in the SEO & GEO course. Get in touch - I'll start with an audit of your current multilingual content and where the terminology has already drifted.

Worth reading next:

/// RELATED_RECORDS

AI & Automation

Copilot, Gemini, or ChatGPT Business - Which AI Package Should Your Company Choose (2026 Comparison)

Microsoft 365 Copilot is really $69-90 per seat a month once you add the required base license, not the $30 from the ad. ChatGPT Business costs $20 today - $5 less than earlier this year. Claude Enterprise dropped from a $40-200 range down to a flat $20 per seat. The prices alone are enough to get lost in - and that's just one of four variables that should actually decide which package goes to the whole company. I break down the three ecosystems into real cost, data protection, admin controls, and integrations - and show how to design a pilot before you sign a 300-seat contract.

13 min
AI & Automation

Data Readiness - How to Prepare Your Company's Data for AI Before You Spend a Dollar on Implementation

Gartner projects that by the end of 2026, organizations will abandon 60% of AI projects specifically because they lacked AI-ready data - not because the model was bad. MIT NANDA's July 2025 report went further: with $30-40 billion invested by companies in generative AI, 95% of pilots delivered no measurable return. The common denominator in both cases isn't technological - it's the data mess nobody cleaned up before signing the vendor contract. Here's how to audit your data sources, assess their quality, and clean them up BEFORE the rollout, not while putting out fires afterward.

13 min
AI & Automation

AI in Accounting Firms - From Invoices to Tax Filings: What to Automate in 2026

73% of accounting firms worldwide have already rolled out some form of AI automation, and among tax advisory firms adoption jumped from 9% in 2024 to 41% in 2025. At the same time, starting February 1, 2026, every business in Poland must be able to receive invoices through KSeF (the national e-invoicing system), and from April, issue them too. That's not a coincidence: mandatory e-invoicing and AI automation reinforce each other - just not the way most accounting firms assume. Here's the process map, what AI already automates well, what has to stay with a human signing the filing, and how to calculate ROI before you sign a vendor contract.

14 min
/// AUTHOR
Paweł Wiszniewski – AI & Web Engineer

Paweł Wiszniewski

SEO & GEO Specialist & AI Engineer

SEO/GEO specialist (10 years) and AI engineer (3 years). I build search visibility, AI systems and automations that reduce costs and improve operational efficiency.

Signal received?

Terminate
Silence

Initiate protocol. Establish connection. Let's build something loud.

> WAITING_FOR_INPUT...