AI Voice for Chinese Tour Commentary

Synthetic narration handles the scripted part of a day well. What it cannot do is answer, and the failures specific to Chinese are pronunciation, not fluency.

A diagonal shaft of pink light cutting through magenta and violet haze, with dust suspended in it

An operator who has worked through whether the talking is part of the product usually lands on a middle option: recorded Chinese commentary, on a handset or inside the Mini Program, without hiring anybody. Synthetic voice makes that cheap enough to consider for the first time.

The split that decides whether it works is not about quality. Current Chinese voice synthesis is fluent. The split is between the part of a day that was always a script and the part that was always a response.

The short version: synthetic narration covers scripted, repeated, factual material well, and it covers it in every language at once, updatable one sentence at a time. It cannot answer a question, shift register for bad news, or carry authority when something goes wrong. The failures particular to Chinese are pronunciation rather than fluency — place names, your brand, and characters with more than one reading — so the work is a pronunciation list, not a better model. And synthetic speech falls under the same labeling regime as generated images.

The scripted half

Some of what a guide says is identical on every departure. The history of the building. What you are looking at across the valley. Why the water is that colour. How long the walk is and where the toilets are.

That material was a script before anybody recorded it, and reading it aloud is not where a guide's value sits. A synthetic voice reading a well-written Chinese script outperforms a human reading a badly translated one — the comparison that actually applies for most overseas operators.

Two further advantages are real and often underweighted. You can update one sentence without re-recording anything, which matters because details change mid-season. And you can produce the same commentary in several languages from one source script, which no human arrangement gives you at that cost.

The half it cannot take

Answering. The reason people value a guide is that they can ask. A recording cannot know that this group has asked twice about the ferry, or that the question underneath the question is whether the walk is manageable.

Bad news. A synthetic voice delivering a route closure in the same even tone it used for the geology is worse than silence, because the register says nothing is wrong while the words say something is. Register is where these systems remain weakest in Chinese, and it is the moment a traveler is judging whether you are being straight with them.

Safety. Understanding is the requirement, and it is not negotiable. A recording can carry the briefing where the briefing is fixed and the guest is attending to it. It cannot verify that it landed, and it has no way to adapt when somebody visibly has not understood.

Authority when something breaks. When a vehicle fails or weather turns, the group needs a person who is deciding. A recording, however fluent, communicates the opposite.

Where each of those falls for your product is the question we worked through in do you need a Mandarin-speaking guide. Synthetic voice changes the cost of one option in that decision. It does not change the decision's shape.

The failures specific to Chinese

This is where operators get caught, because the audio sounds fine to them and wrong to the listener.

Place names and loanwords. Your region, your meeting point, the mountain, your own brand. Transliterated foreign names are exactly where a Chinese voice engine has least to go on, and a mispronounced place name inside an otherwise fluent recording is instantly audible. It reads as a business that did not check.

Characters with more than one reading. Chinese has plenty of characters pronounced differently depending on the word, and engines resolve them from context. Context fails most often on names, the worst place for it to fail. The same character in your company name can come out two ways in two sentences.

Numbers, times and prices. These are usually technically correct and hard to follow — pacing and grouping in spoken Chinese numbers do not survive default handling well. Anything a listener has to act on should be checked by ear.

Which Mandarin. Voices differ audibly in regional register, and one that reads as not-mainland to a mainland audience is a signal, in the way a European brand using an obviously American voice would be. Pick deliberately.

The pattern in all four: the fix is a pronunciation list and a listening pass, not a better engine. Supply the readings for every proper noun before generation, then have somebody who speaks Chinese listen to the whole thing. Reading the script is not the same check — the script was probably fine.

It counts as synthesised content

Synthetic speech sits inside the same Chinese labeling regime as generated images, and the same declaration applies when you publish it. The declaration you have to make, and the part of the duty that falls on the platform instead of you, are both in China's AI-content labeling rule.

Worth separating from the ethics of the earlier argument: a synthetic voice reading true facts about a real place asserts nothing about how the place will look. It is a delivery mechanism, not a depiction, and that puts it in a different category from a generated photograph of your site.

A workable arrangement

The version that holds up in practice is mixed, and the division is by content type rather than by budget.

Recorded and synthetic: the fixed commentary at each stop, the standard briefing, the background material, anything you would have written down anyway. Human and live: the welcome, the answering, anything delivered when a plan has changed, and the close.

That arrangement also survives the thing that kills pure-recording setups, which is the day when a stop is skipped. A recording assumes the itinerary. A person handles the one that actually happened.

Before you commit to it

Write the Chinese script first, as a script, with somebody who writes Chinese. A synthetic voice cannot rescue a translated-sounding script — it delivers the flatness faithfully, and the boundary around what translation tools handle safely is in where AI translation stops being good enough.

Then generate one stop, not forty. Put it in front of somebody Chinese who has not read the text, and ask them what they heard rather than whether it sounded good. Names come back wrong at that stage, cheaply.

What is still moving

Voice synthesis is improving quickly, and the pronunciation problems described here are the kind that get solved. Some of this will read as dated within a year or two.

The division of labour will not move, because it does not come from the technology. A recording cannot answer a question that has not been asked yet, and that limit sits underneath every version of this that will ever ship.

Where CN1X fits

We write the Chinese script, supply the pronunciation list for your proper nouns, and put the output in front of a native listener before it reaches a guest. Where a client wants forty stops generated in an afternoon, the listening pass is the step that cannot come out.

We do not supply guides or interpreters, so nothing here is us selling you the alternative. If you have a commentary script and want to know which parts of it should never be a recording, send us the running order.

More from the blog

5 min read Guides

Korean and Japanese Operators Selling into China

A short flight changes the customer. What long-haul advice gets wrong for operators in Korea and Japan, and which barriers you have already cleared.

5 min read WeChatGuides

What You Supply Before a Mini Program Build Starts

The quote is signed and then three weeks pass with nothing visible happening. Almost always the build is waiting on things only you can hand over.