Onboarding an agent on your brand voice in one afternoon

Not a prompt. A training set: your approved replies, your no-go list, your terms, and a calibration run where you decline things and say why. Half of consumers are put off by generic review replies; this is the afternoon that keeps the agent from writing them.

EDEditorial·30 Jul 2026·7 min read·1483 words
In short

An agent learns a brand voice from examples, not adjectives. Onboarding in Business Profile Agent takes about four hours: 30 approved replies, a written no-go list, the brand's terminology, per-location facts, and a calibration run of 40 real reviews where a person approves, edits or rejects each draft and says why. A proof-reading agent on a different model checks language, terminology and promises before anything is published.

Every brand we onboard says the same three words about its voice — friendly, professional, personal — and every brand means something different by them. An agent cannot work from adjectives. It can work from examples, and the research on this is unusually clear. Anthropic's prompting guidance: "Examples are one of the most reliable ways to steer Claude's output format, tone, and structure. A few well-crafted examples … improve accuracy and consistency" — three to five per case, relevant, diverse enough that the model "doesn't pick up unintended patterns". OpenAI's guide says the same about few-shot examples and adds that fine-tuning is the step after prompting, not instead of it.

So onboarding is not a form with a tone slider. It is an afternoon in which the brand shows the agent what it has already approved, tells it what it must never say, and then corrects it forty times.

What the agent learns from

Approved replies. Thirty is enough; fifty is better. Real replies a person at the brand signed off — ideally across star ratings, languages and locations, including two or three to complaints. These become the examples the drafting agent is shown for each new review, chosen by similarity.

The no-go list. Written by the brand in plain sentences: "We never promise a callback." "We do not discuss pricing in replies." "Never mention a competitor by name." The list is a hard filter, not a suggestion; the proof-reader rejects a draft that breaks it.

And two things people forget

Terminology. The brand says "workshop", not "garage"; "Kestrel", never "Kestrel Autohaus"; "branch", not "store". A short glossary, with the wrong forms listed, is worth more than a page of tone description.

Per-location facts. Hours, parking, booking link, languages spoken at the counter. The agent may only state facts that exist in this record. That rule is what keeps the afternoon from becoming a liability.

The afternoon, hour by hour

1 pm

Collect the approved replies

Export the last year of owner replies from the profiles; a person marks thirty as "this is us". Ten minutes of reading, twenty of choosing.

1:45 pm

Write the no-go list and the glossary

Usually eight to fifteen sentences. The most useful prompt for the room: "What has a new employee written in a reply that made you wince?"

2:30 pm

Calibration run

The agent drafts replies to forty real, recent reviews. The person approves, edits or rejects each one and says why in a few words. Every edit and every reason is recorded as feedback — the same feedback loop the Logbook's flag uses.

4 pm

Read the results, set the threshold

The approval rate by star band tells you where to start. Positive reviews usually pass at 90% or more on the first afternoon; replies to complaints need a second round. The confidence threshold is set from these numbers, not from a default.

4:45 pm

Language check

Ten reviews in every language the locations receive, including two-line ones. This is where wrong-language replies would happen; it is cheaper to find them here.

Why "facts from the record only" is not negotiable

In February 2024 a Canadian tribunal ruled against Air Canada after its website chatbot told a customer he could apply for a bereavement fare retroactively — a policy the airline did not have. Air Canada argued the chatbot was "a separate legal entity that is responsible for its own actions". The tribunal's reply is the sentence every brand should read before onboarding an agent: "It should be obvious to Air Canada that it is responsible for all the information on its website." The damages were small, CAD 812.02. The principle was not. A review reply is published under your name, in public, indefinitely. The drafting agent is allowed to be warm; it is not allowed to be generous with facts it does not have.

The agent may say "we are open until six on Saturdays" because the profile says so. It may not say "we will call you back" unless the brand has written that promise into the record.Onboarding rule 1

What consumers said in the blind tests

58%
preferred the AI reply

BrightLocal's 2024 blind test of an AI-written vs a human-written review response; the 2025 result was "almost identical".

46%
suspicious of AI-looking text

Consumers who would suspect a review was fake if it read as AI-written (2025).

50%
put off by templates

Consumers unlikely to choose a business whose replies are generic or templated (2026).

Read together, the three numbers say one thing: people like a good reply and dislike a reply that sounds like nobody wrote it. The tell is not that a machine helped. The tell is genericness — the reply that could sit under any review at any business. Thirty approved examples and a glossary are the cheapest known cure.

The proof-reader on a different model

Two findings from the research shaped the last piece. Models "struggle to self-correct their responses without external feedback", and LLM evaluators "recognize and favor their own generations". So the agent that checks a draft is not the agent that wrote it, and it does not run on the same model. It checks four things: is the reply in the reviewer's language; does it use the glossary's terms and none of the forbidden ones; does it make a promise the record does not contain; does every fact in it exist in the location's profile. Language is the check that fails most often on short reviews — language identification is measurably harder on short texts, which is why a two-word review in Polish can pull an English reply out of a careless system. The proof-reader sends such a draft back with a note; the Logbook shows the round-trip.

After the afternoon, the learning does not stop. Every edit a manager makes to a draft, and every flag in the Logbook, goes into the agent's next round. The thirty examples become a hundred without anyone scheduling a second workshop.

Questions we get on this

How many examples does an AI agent need to learn a brand voice?

Around thirty approved replies across star ratings, languages and locations are enough to start; fifty is better. Anthropic's guidance recommends three to five relevant, diverse examples per type of case, and OpenAI's guidance recommends few-shot examples before considering fine-tuning. The examples grow automatically as managers edit and approve drafts.

Can AI review replies sound like my business rather than generic?

Yes, if the agent is given real approved replies, a glossary of the brand's terms and a no-go list, and is restricted to facts in the location's record. BrightLocal's blind tests found most consumers preferred an AI-written reply, while 50% are put off by generic or templated replies — the problem is genericness, not authorship.

What should be on the no-go list for review replies?

Promises the brand does not make (callbacks, refunds, discounts), topics it does not discuss in public (pricing, staffing, legal matters), names of competitors and employees, and wording that argues with the reviewer. Each rule is one plain sentence. The proof-reading agent rejects any draft that breaks a rule.

Why did the agent reply in the wrong language?

Language identification is less reliable on very short texts, so a two-line review can be misidentified by a system without a second check. In Business Profile Agent the proof-reading agent verifies the reply's language against the review's before publication and sends mismatches back to the drafting agent; the round-trip is visible in the Logbook.

Does the agent keep learning after onboarding?

Yes. Every edit a manager makes to a draft and every flagged line in the Logbook is recorded with a reason and reviewed in the agent's next learning round, lowering its confidence on similar cases in the meantime. The onboarding afternoon sets the starting point; the approval workflow keeps extending it.

Sources

  1. Anthropic — Prompting best practices (examples / few-shot)
  2. OpenAI — Prompt engineering guide
  3. OpenAI — Model optimization guide (prompting before fine-tuning)
  4. McCarthy Tétrault — Moffatt v. Air Canada (2024)
  5. Dentons — Airline ordered to compensate for chatbot misinformation (2024)
  6. BrightLocal — Local Consumer Review Survey 2024
  7. BrightLocal — Local Consumer Review Survey 2025
  8. BrightLocal — Local Consumer Review Survey 2026
  9. Huang et al. — Large language models cannot self-correct reasoning yet (ICLR 2024)
  10. Panickssery, Bowman & Feng — LLM evaluators recognize and favor their own generations (2024)
  11. Baldwin & Lui — Language identification: the long and the short of the matter (NAACL 2010)
brand voiceonboardingreview repliesfew-shot
Published 30 Jul 2026
ED
EditorialNotes from the people who write the agents' rules — and read the logbook every morning.