Using AI to Read Your Chinese Enquiries and Reviews

Translating messages one at a time is triage. The move that pays is reading a whole season and counting what recurs — carefully, the tool agrees with you.

Five people at one wooden table, each reading from their own laptop screen

Most operators already run individual Chinese messages through a translator, and that is a settled, sensible habit. We said as much when we drew the boundary around what these tools handle safely in where AI translation stops being good enough.

This is the other use, and almost nobody does it: taking every enquiry and every review from a whole season and reading them as one body of text. The volume sits in an awkward middle — too much to read properly, too little to feel like data — so it usually goes unread. That gap is where the tool earns its keep.

The short version: the value comes from counting, not summarising. Label a season's messages by the question each one is really asking, then count, then go read the actual lines behind the top three. Three habits keep it honest: never ask a leading question, make it quote rather than paraphrase, and treat three mentions as a reason to go look rather than as a finding. And decide where the messages are allowed to be pasted before you start, because they contain customers' names and numbers.

Counting beats summarising

Ask a model to summarise two hundred enquiries and you get a paragraph that sounds right and tells you nothing you can act on. Ask it to label each one with the question actually being asked, then count the labels yourself, and you get a ranked list of the things your pages fail to answer.

The ranking is the product. Somewhere near the top there will be a question you answer in the second paragraph of a page nobody reaches, or one you answer in English only, or one you have never answered anywhere. Those are page edits, and they are cheap.

Two hundred messages is enough for this, and that is the surprising part. You are not looking for statistical confidence. You are looking for the four questions that arrive every week.

The valuable review is the one said once

Reviews repeat. Twenty people mention the scenery, eleven mention the guide, and you already knew both.

What hides in a set that size is the specific thing said once, usually in passing, often inside a positive review: the pickup that was fifteen minutes late, the toilet stop that did not happen, the thing the guide said that landed badly. A person scanning for problems skips those, because the review reads as happy overall. Given the whole set and asked for every operational detail mentioned regardless of sentiment, a model surfaces them.

Those single mentions are worth more than the repeated ones. The repeated things you know about. The thing said once, by somebody who otherwise liked the day, is a fault nobody has told you about properly.

Three habits that keep it honest

Do not ask a leading question. "What are customers complaining about?" produces complaints, whether or not the set contains many. The failure mode here is your own question, not the model. Ask it to categorise before you have told it what you expect to find, and look at the categories before you look at the answers.

Make it quote, not paraphrase. Require the original Chinese line for every claim it makes. This does two things: it stops the invented theme, and it lands you on the actual sentence, where the useful specificity lives. A paraphrase of a complaint is much less usable than the complaint.

Treat a small count as a pointer, not a result. Three mentions of the same thing means go and read those three. It does not mean eleven per cent of your customers think that. The temptation to turn labels into percentages is strong and the numbers do not support it — and where a figure matters, say which part you measured and which part you inferred — the discipline we set out in measuring when attribution is broken.

Two failures to expect

It will manufacture themes on request. Ask for five themes and you will receive five, in confident prose, whether the set holds five or two. Ask instead how many distinct themes it can support with quotes, and let the count come out at whatever it is.

It flattens intensity. Tone is what these tools handle worst in Chinese, so a set read this way tells you accurately what was said and unreliably how strongly. Mild disappointment and real anger arrive looking similar. For anything you plan to act on emotionally — a reply, an apology, a refund decision — go back to the original with somebody who reads Chinese, for the reasons we went through in answering a bad review in Chinese.

Who is actually in your sample

The people who wrote to you are not your customers. They are the ones who had a question your pages left open, or a complaint big enough to type out. Both groups are unrepresentative, and in opposite directions.

That does not reduce the value — an unanswered question is worth fixing no matter how few people asked it — but it does rule out one conclusion. You cannot read this set and learn what most of your customers thought. You can only learn what some of them needed.

Decide where the messages may go first

This is the part that gets skipped. Enquiries contain names, phone numbers, WeChat IDs and sometimes passport details. Pasting a season of them into whatever consumer tool is open in a browser tab is a data-handling decision, and making it by accident is how operators end up somewhere they would not have chosen.

Strip the identifying fields before anything leaves your systems, or use a tool your business has actually approved for this. Personal data, practically covers what an overseas operator genuinely has to handle here.

A workable cadence

Once a season, not weekly. The whole point is seeing what recurs, and a week does not contain enough to recur.

Export the raw Chinese without pre-filtering — filtering first removes exactly the quiet mentions you are looking for. Label, count, read the top three properly, and write down what you changed. Next season, the comparison between the two sets tells you whether the change worked, which is the only cheap outcome measurement available in this channel.

Where this stops being reliable

Nothing here produces a number you could defend to a board. It produces a list of specific things customers needed and did not get, in their own words — more useful than a number for deciding what to change, and no substitute for one.

Whether any of it holds also depends on your volume. Under fifty messages a season, read them yourself; a tool adds nothing but a layer between you and four paragraphs of text.

Where CN1X fits

We do this as part of the monthly reporting, and the output is a list of page edits with the Chinese lines attached, so you can see what prompted each one. Sometimes the list is short, and a short list is a real result rather than a failure of effort.

Where your customers' data may legally go is a question for somebody qualified on it and not for us, and we will not turn counts into percentages to make a report look more certain than the underlying set supports. Tell us roughly how many enquiries a season brings you — the volume decides whether this is worth doing at all.

More from the blog

5 min read Guides

Korean and Japanese Operators Selling into China

A short flight changes the customer. What long-haul advice gets wrong for operators in Korea and Japan, and which barriers you have already cleared.

5 min read WeChatGuides

What You Supply Before a Mini Program Build Starts

The quote is signed and then three weeks pass with nothing visible happening. Almost always the build is waiting on things only you can hand over.