Visibility in English AI answers does not carry over to other languages, and the drop is usually larger than teams expect. SQSEO treats each language as its own prompt set for that reason: the questions differ, the competitors differ, the citable sources differ, and the retrieval layer is working from a different and usually much thinner pool of documents. A brand cited in nine of ten English answers can be absent from all ten German ones without anything being wrong with the brand.
Why the transfer is weaker than it looks
Large models are multilingual, which creates a reasonable but mistaken expectation that knowledge transfers cleanly. Model knowledge does transfer to a degree. Retrieval does not.
The retrieval layer searches documents, and documents have languages. A query in Dutch retrieves Dutch pages far more often than English ones, so the citable pool for a Dutch question is the Dutch web, which for most B2B categories is a fraction of the size of the English one. Your English content is not in that pool no matter how good it is.
Benchmarks built to measure this gap are explicit about its size. mMARCO, a multilingual version of a standard passage ranking dataset, exists precisely because retrieval quality measured in English does not predict retrieval quality in other languages, and the broader embedding benchmarks collected in MTEB show meaningful variation in how well models represent different languages. Weaker representation means noisier retrieval, which means less reliable citation.
What actually differs between languages
| Layer | Transfers from English? | Why | What to do |
|---|---|---|---|
| Model background knowledge | Partially | Multilingual training carries entity knowledge across | Build entity presence once, benefit broadly |
| Retrieval pool | No | Queries retrieve same language documents | Publish native content per language |
| Question phrasing | No | Buyers ask differently, not just translated | Research prompts per language |
| Competitor set | No | Local players dominate local pools | Track local rivals, not your English ones |
| Citable third parties | No | Review sites and media are market specific | Earn local mentions |
| Schema and technical setup | Yes | Markup is language neutral | Do it once, apply everywhere |
The pattern is that anything grounded in documents is local, and anything grounded in the model’s weights is partially global. Since generative answers lean heavily on retrieval for anything commercial or current, the local half dominates.
Machine translated content is the trap
The obvious response to a thin local pool is to translate the English site. It rarely works well, for three reasons that compound.
First, translated content answers translated questions rather than real ones. A German buyer does not ask the German translation of the English question; they ask a different question shaped by a different market, different regulations and different competitors. A page answering the wrong question precisely still does not match the query.
Second, translation flattens the specific vocabulary that makes retrieval work. Industry terms, product category names and the way practitioners actually talk are frequently not the dictionary translation. Retrieval compares meanings, and a page written in slightly off register language represents its meaning less sharply.
Third, quality is visible. Research on generative engine optimization found that source characteristics change how generative engines use content, and thin translated pages tend to lack exactly the properties that get content used: concrete statistics, quotable statements and specific claims.
None of this argues against translation as a starting point. It argues against treating translation as the finish line. A translated draft rewritten by someone who works in that market is a different artefact from a translated page.
The competitor set changes completely
This is the finding that most surprises teams running their first non English audit. The brands appearing in German, French or Japanese answers for your category are frequently not the brands you compete with in English.
Local players have local content, local review coverage and local press, which means they occupy the local citable pool. A global leader with only English content can be genuinely absent from a market where a smaller regional competitor is cited constantly.
The practical consequence is that tracking your English competitor set across languages measures the wrong thing. Run an open prompt in each language, see who actually appears, and build the tracked competitor list from that. This is the same principle as tracking competitor citations, applied per market rather than globally.
Query volume is not the right selection criterion
Choosing which languages to invest in by search volume imports a habit from classic SEO that fits poorly here. In AI search the useful question is not how many people search, it is how thin the citable pool is.
A market with modest query volume and almost no quality content in the category is a market where a handful of well written native pages can dominate answers quickly. A market with high volume and dense existing coverage is expensive. The thin pool is the opportunity, and it is invisible if you select markets by volume.
The way to measure it is direct: run your key questions in each candidate language, look at what gets cited, and judge the quality. If the answers are citing forum posts, vendor marketing and auto translated pages, the pool is thin and winnable. If they are citing substantial native industry publications, it is not.
Regional variants matter more than expected
Language is not the same as market. Spanish for Spain and Spanish for Mexico retrieve different documents, involve different competitors, and use different terminology for the same product category. The same applies to Portuguese, French, German across Germany, Austria and Switzerland, and English itself across the US, UK and India.
Where regional differences are commercial rather than merely linguistic, and they usually are once pricing, regulation or availability enter, separate content per region beats one page per language. Where they are purely linguistic, one well written page per language is usually enough.
The cheap diagnostic is to run the same question with a regional qualifier and compare the cited sources. Different sources means different pools means separate content is worth it.
Technical setup that does carry over
Some work genuinely is do once, benefit everywhere. Schema markup is language neutral in structure, so a correct Product or Organization block serves every locale. Crawlability applies across the board: if your site renders client side, it is invisible in every language equally. Clean URL structure, hreflang and sitemaps help engines find the local versions that do the work.
Entity consistency is the other global asset. A brand described consistently across languages, with the same category framing and the same product names, is easier for a model to treat as one entity. Inconsistent naming across locales splits the entity and weakens it everywhere, which is the multilingual version of the problem behind how AI engines categorize your business.
Measuring across languages without fooling yourself
Equal run counts per language, always. Comparing twenty English runs against six Dutch ones measures your sample sizes. Use the same number of runs per prompt in every language, and expect the absolute numbers to be lower outside your home market.
Track the citable pool, not only your own appearance. Record which URLs get cited per language. A market where your answers cite three shaky sources is a market where one good page changes the picture; a market citing established publications needs a different strategy.
And set separate baselines. Comparing German visibility against English visibility produces a number that always looks like failure. Comparing German visibility this quarter against German visibility last quarter produces a number that means something.
Which engines your buyers use also changes by market
Engine share is not uniform, and assuming a global ChatGPT plus Perplexity picture will mislead you in several large markets. Copilot has meaningful presence in enterprises standardised on Microsoft, which skews by country and by sector. Gemini distribution follows Android and Google Workspace adoption. Regional assistants matter in markets where global providers have limited reach or where local platforms dominate messaging and search.
The practical effect is on where you spend run budget. Measuring German visibility exclusively through ChatGPT, in a market segment that lives inside Microsoft 365, produces a number that is accurate and irrelevant. Ask a handful of customers in each market which assistant they actually use before deciding what to track, because the answer is cheap to obtain and expensive to guess wrong.
It also affects what content wins. Engines differ in how heavily they retrieve versus rely on model knowledge, so a market dominated by a retrieval heavy engine rewards publishing native content faster than one dominated by an engine leaning on its weights, where entity presence matters more and moves slower.
Third party coverage is the hardest part to localise
Publishing your own pages is the controllable half. The other half is being mentioned by sources that already sit in the local citable pool, and that is slower.
Every market has its own review sites, its own comparison directories, its own trade press and its own active communities. The equivalents of the English language sources you already appear on frequently exist, staffed by fewer people and easier to reach than their English counterparts. Getting listed on a local directory or covered by a local trade publication can move answers in a thin market more than several of your own pages, because a third party source carries independence that your own site cannot.
Local communities matter for the same reason they matter in English. Where practitioners in a market gather, whether that is a national forum, a professional association or a language specific subreddit, those discussions frequently end up in the retrievable pool. Being genuinely present there is slow and it compounds, which is the same dynamic behind how digital PR trains language models on your software.
Sequencing a multilingual rollout
The mistake is doing all languages at once, thinly. Six markets with four translated pages each produces twenty four pages that win nothing. One market with six native pages produces a position.
A sequence that works: pick the thinnest promising pool first, publish six to ten native pages answering the highest intent questions in that market, measure for a quarter with equal run counts, and only then decide whether to deepen or move on. The first market teaches you how much content it takes to move the number, which is the figure you need to budget every subsequent market.
Keep the prompt set stable per market once set. Every time you reword a prompt you reset that market’s baseline, and with fewer runs per language the noise is already higher than it is in English.
A worked example: a B2B tool entering two markets
An English language B2B tool appeared in 45 percent of English answers across twenty buying intent prompts. The team ran the same prompt set, translated, in German and Dutch, twenty runs each.
German: 4 percent. Dutch: 0 percent. Neither result reflected anything about the product.
The audit found the causes quickly. In German, answers cited two established local industry publications and three German competitors, none of which the team had ever tracked. The pool was dense and the barrier real. In Dutch, answers cited almost nothing credible: a forum thread, a machine translated aggregator, and one competitor’s thin landing page. The pool was close to empty.
The decision followed from the pools rather than from market size. Dutch, a smaller market, got six natively written pages answering the six highest intent questions, written by someone who works in that market rather than translated. Within a quarter the tool was appearing in roughly a third of Dutch answers. German was deferred, because competing there meant earning coverage in two publications rather than publishing six pages, which is a different and longer project.
Key takeaways
Model knowledge partially transfers across languages; retrieval does not, because queries retrieve same language documents. Translate as a draft, never as a deliverable, since local buyers ask different questions in different vocabulary. Expect a different competitor set in every language and build the tracked list from what actually appears. Choose markets by how thin the citable pool is rather than by query volume, because a thin pool is cheap to win. Keep run counts equal per language and set separate baselines, or every non English number will read as failure.