Brand voice briefs fail in six recognisable ways. Each has a different cause and a different fix, so this guide is arranged by the symptom you are actually seeing rather than by the order you would build a brief.
Diagnose before you add When output comes back wrong, the instinct is to write more brief. That is usually the wrong move. A longer brief is followed less closely than a short one, because instructions start competing with each other, so adding text to fix a specific fault often makes every other instruction weaker. Each of the six faults below produces a distinctive signature in the output. Read the symptom, apply the one fix, and confirm it landed before touching anything else. |
Find the line that matches what you are seeing. Each entry below carries the diagnosis, the exact lines to add, and a test to confirm the fix worked.
| What you are seeing | What is actually broken | Fix |
|---|---|---|
| Reads like every other brand, nothing specifically wrong | No lexical constraint, so the model uses the most probable phrasing | Symptom 01 |
| Uses our words but does not sound like us | No mechanics specified, so sentence architecture defaults | Symptom 02 |
| Follows every rule and still feels lifeless | No examples, or examples that are weak or too similar | Symptom 03 |
| Starts right, drifts after a few paragraphs | Instruction position and session length, a documented effect | Symptom 04 |
| Invents figures, dates or details we never supplied | No instruction for what to do when information is missing | Symptom 05 |
| Works for blog posts, breaks on emails and error messages | Voice and tone treated as one thing when they are two | Symptom 06 |
Symptom 01It reads like every other brand ![]() |
This is the most common fault and the most misdiagnosed. Teams read output like the sample below, conclude the tone is wrong, and rewrite the tone section. The tone is usually fine. The problem is smaller and more literal.
Typical output with this fault We are thrilled to announce the launch of our new automatic filing feature. Designed with you in mind, this powerful addition streamlines your workflow and eliminates the hassle of manual work. Say goodbye to tedious paperwork and hello to peace of mind. |
Nothing in that passage breaks a tone rule. It is friendly, professional and approachable, exactly as most briefs request. What makes it generic is six specific strings: thrilled to announce, designed with you in mind, streamlines, say goodbye to, hello to, peace of mind.
A language model produces likely continuations. A brief made of adjectives rules almost nothing out, because nearly every piece of business writing ever published satisfies friendly and professional. The model therefore lands on the highest probability phrasing available, which is the phrasing everyone else also landed on.
The fix is not more description. It is prohibition, which is unambiguous in a way that description is not.
ADD TO THE BRIEF
|
Build the list from writing you already dislike rather than from memory. Take three competitor pages or three internal drafts that were rejected, and extract the exact strings you would not sign your name to. Fifteen to twenty five entries is the working range. Past that, adherence falls off and the list stops being read carefully, so rank by how often a phrase actually appears rather than by how much it irritates you. Those are different orderings.
CONFIRM THE FIX WORKED Generate five pieces and search the output for every banned string. Zero hits means the list is being applied. Consistent hits on the same two or three entries usually means the list has grown too long rather than that the model is ignoring it. |
Symptom 02It uses our words but does not sound like us |
The banned list worked. The vocabulary is right, the jargon is gone, and it still reads as someone else. This means the fault has moved from word choice to sentence architecture, which is where a large share of voice actually lives.
Two writers can use an identical vocabulary and sound nothing alike. Rhythm, sentence length variation, how paragraphs open and close, whether the reader is addressed directly, how often a short sentence lands after a long one. None of that is covered by a lexicon, and none of it is covered by adjectives either.
This is the least glamorous fix in the guide and produces the largest visible change after examples. Specify in numbers wherever a number is possible, because a number is checkable and a description is not.
ADD TO THE BRIEF
|
The opening line rule fixes more than any other single instruction Generated prose is most patterned at the start, where it reliably restates the topic or sets a scene before saying anything. Requiring the first sentence to carry the most useful fact in the piece removes more generic texture than any other line you can add, and it takes ten words to write. Compare the two openings. We are thrilled to announce the launch of our new automatic filing feature. Against: automatic filing is now on for every pay run. The second one has already told the reader something. |
CONFIRM THE FIX WORKED Measure rather than read. Take ten generated sentences, count the words in each, and check the average and the spread against what you specified. Uniform sentence length is the signature of this fault surviving. |
Symptom 03It follows every rule and still feels lifeless |
Everything you specified is being obeyed and the result is still flat. This means you have hit the limit of what can be written down. The remaining difference between your voice and the output is real but not articulable, and the established fix is to stop describing and start showing.
Providing examples inside the prompt, usually called few shot or multishot prompting, is the standard technique for exactly this situation: a target that is hard to describe and easy to recognise. Brand voice is the textbook case, since you know it when you read it and struggle to specify it.
The instinct is to supply many samples. The evidence points the other way. In one study of style matching, a single high quality demonstration outperformed ten lower quality ones, so the work sits in selection rather than volume. Three to five excellent samples beat fifteen adequate ones.
Published prompting guidance asks for two properties in examples. They should be relevant, mirroring the actual use case rather than a different format. And they should be diverse, varied enough that the model does not lock onto an unintended pattern shared by all of them. Five samples that happen to open with a customer quote will teach the model that every piece opens with a customer quote.
• P
• O
• A
• N
The single highest value addition is contrast. Pair one on voice sample with one rejected piece and a line explaining the difference. A positive example marks a point. A pair marks the boundary, which carries more information.
ADD TO THE BRIEF
|
More examples is not monotonically better Examples consume context that could hold your actual source material. Separately, on some reasoning oriented models heavy example loading has been observed to degrade rather than improve output, with published guidance recommending a leaner prompt for those models instead. Start with three, measure, and add only if output is still drifting. |
CONFIRM THE FIX WORKED Run a blind mix. Put three generated pieces among three human written ones and ask a colleague who did not build the brief to sort them. Sorting at roughly chance means the samples are doing their job. |
Symptom 04It starts right and drifts after a few paragraphs ![]() |
The first two paragraphs are on voice and the piece degrades from there, or a long working session starts well and is unrecognisable by message forty. This is the one fault on the list that is not really about your brief. It is a documented property of how these models use long inputs.
Research published in Transactions of the Association for Computational Linguistics examined how well models actually use information placed at different points in their input. Performance follows a U shaped curve. It is highest when the relevant information sits at the very beginning of the input, a primacy effect, or at the very end, a recency effect, and it degrades significantly when the model has to use information buried in the middle.
The finding is stark enough to be worth stating precisely. In one tested setting, placing the relevant information mid context produced performance below the closed book baseline, meaning the model did better with no documents at all than with the right document in the wrong position. The effect replicated across several model families, and models with extended context windows were not reliably better at using their context, so a larger window does not solve it.
What this means for a brief specifically A long brief is not read evenly. Whatever you bury in the middle carries the least weight, which explains the common experience of a carefully written brief where one section is consistently ignored while the rest is followed. So position deliberately. Put the constraints that matter most at the top and the bottom, and put the material that is merely reference in the middle. If one instruction is repeatedly ignored, try moving it before rewriting it. |
ADD TO THE BRIEF Structure the brief so that:
TOP voice position, reader, the two or three rules that matter most MIDDLE full lexicon, reference material, samples BOTTOM output format, edge handling, and a restatement of the single most violated rule
At the end of a long brief, repeat the one constraint that slips most often. Repetition at the boundary is cheap and it is where attention is strongest. |
For session drift the fix is procedural rather than textual. A brief given once at the start competes with everything generated since. Restate only the constraint that is slipping rather than the whole brief, or start a fresh session for each new piece. Where your tool supports a persistent system instruction, use it, because that places the brief in a position that does not decay as the conversation grows.
CONFIRM THE FIX WORKED Run the same brief in a completely fresh session with no history. If output is noticeably better than it was at the end of your last long session, the brief is fine and session length was the fault. |
Symptom 05It invents details we never gave it ![]() |
The voice is right and the copy contains a percentage, a date, a customer name or a capability that came from nowhere. This is the fault with the highest cost, because it is the one that reaches customers and is hardest to catch in review, precisely because the surrounding prose reads well.
The cause is a gap in the brief rather than a flaw in the writing. Every brief so far describes the normal case. None of them says what to do when a required fact is missing, and in the absence of an instruction the model does what it always does, which is produce the most plausible continuation. A plausible continuation of we reduced filing time by is a number.
ADD TO THE BRIEF
|
The uncertainty list is the cheapest line in any brief Asking for a list of assumptions after the deliverable costs nothing, does not affect the copy, and converts silent guesses into a checklist for whoever reviews the work. It is the single most useful line to add if you are handing generated copy to someone else to approve. |
CONFIRM THE FIX WORKED Deliberately withhold a figure the piece needs and generate anyway. A brief that is working returns a marked gap. A brief that is not returns a confident number, which tells you exactly what has been happening in the pieces you did not test. |
Symptom 06It works for blog posts and breaks on emails |
One brief produces good long form copy and unusable error messages, or a warm onboarding email and a tone deaf outage notice. This means voice and tone have been written as one thing when they are two, and the distinction is not academic.
Voice is constant. It is what stays the same across everything you publish, and it is the thing a reader would recognise. Tone varies with the reader and the situation. The same brand should not address a prospect and a customer whose payment just failed in the same register, and a brief that fixes both produces writing that is either inappropriately cheerful in a crisis or needlessly grim in an announcement.
The established framework for making this concrete comes from usability research at the Nielsen Norman Group, which reduced tone to four measurable dimensions rather than a list of adjectives.
| Dimension | What it measures | Example setting |
|---|---|---|
| Funny to serious | Whether humour is attempted at all | 7 of 10 toward serious, because errors here cost money |
| Formal to casual | Register, contractions, vocabulary | 6 of 10 toward casual, contractions yes and slang no |
| Respectful to irreverent | Attitude toward the subject and reader | 2 of 10, firmly respectful, never at the reader expense |
| Enthusiastic to matter of fact | Energy and expressed excitement | 8 of 10 toward matter of fact, the product is the excitement |
Two things make a scale work where an adjective does not. Each dimension names an explicit opposite, so the model knows what it is moving away from as well as toward. And a number forces a decision that a word lets a team dodge, which is usually where internal disagreement surfaces and gets settled.
ADD TO THE BRIEF
|
One research finding worth carrying into your settings Across the tested samples, casual, conversational and moderately enthusiastic tones performed best overall, though not universally. A conversational but serious tone worked well for a bank, and the more casual of two tested banks was rated friendlier by around 0.7 points on a five point scale and slightly more trustworthy. The uncomfortable finding is the useful one. In one tested pair, a tone made the brand seem friendlier without making people more likely to choose it. Likeable and persuasive are separate axes, so set your positions against what the writing is for rather than against what sounds pleasant. |
CONFIRM THE FIX WORKED Ask for the hardest thing you publish, usually an apology, a price increase or an outage notice. A brief that only handles announcements will produce something inappropriately upbeat, which is visible immediately. |
Applied together and ordered according to the position finding in symptom 04, with the highest weight material at the top and bottom and reference material in the middle. Sections are tagged because delimiters help models keep parts of a long instruction distinct, and because it lets you replace one section without rewriting the rest.
ADD TO THE BRIEF
|
Between 400 and 800 words of instruction plus samples is the working range. Shorter and it stops ruling things out. Longer and adherence drops as instructions compete. If yours is growing past that, the usual cause is describing in prose what a sample would demonstrate in a paragraph.
Most teams do not have a written voice, and the ones that exist are usually three adjectives on a slide. You can derive every input the six fixes need from your own back catalogue in about ninety minutes.
• Collect twenty published pieces. Real work across your actual formats, not aspirational drafts.
• Sort into three piles. This is us, this is not us, and this is fine but forgettable. The middle pile is the informative one because it marks the boundary.
• Count what differs. Sentence length, contraction use, opening moves, whether the reader is addressed directly. Count on a sample rather than estimating, since intuitions about your own writing are unreliable.
• Place the four dimensions using the good pile only, then check the placement predicts the difference from the bad pile.
• Harvest the banned list from the not us pile, taking exact strings rather than themes.
• Pick the samples from the good pile, one per format, including one difficult piece.
A shortcut with a real limit You can paste the good pile into a model and ask it to describe the voice it detects, which produces a serviceable first draft and is a legitimate way to save an hour. The limit is that it describes what you have written, including the habits you were hoping to drop. Treat the output as a mirror rather than a recommendation, and edit it against the writing you want to produce next year. |
A brief that feels thorough and a brief that works are different things, and the gap is measurable in an afternoon. Each test isolates a different fault so a failure tells you which symptom to return to.
| Test | Method | What a failure points to |
|---|---|---|
| Banned list audit | Search five outputs for every banned string | Symptom 01, or a list that has grown too long |
| Sentence measurement | Count words per sentence across ten sentences | Symptom 02, mechanics not specified or not followed |
| Blind mix | Three generated pieces among three human written, sorted by a colleague | Symptom 03, samples weak, too few, or too similar |
| Cold start | Run the brief in a fresh session with no history | Symptom 04, the conversation was doing the work |
| Withheld fact | Ask for a piece that needs a figure you never supplied | Symptom 05, no instruction for missing information |
| Hard case | Request an apology, a price increase or an outage notice | Symptom 06, no tone overlay for difficult content |
| Edit distance | Count edits needed before publication across ten pieces | The only number that connects to time actually saved |
| Repeat run | Generate the same piece five times | Whether output is consistent or one result was luck |
Applied together, these fixes reliably transfer register, vocabulary, rhythm and structure. That is most of what people mean by brand voice and most of the practical work.
What they do not transfer is judgement. Knowing that this particular week is the wrong week for a light touch. Knowing which detail your specific customers will fixate on. Knowing that a sentence is technically accurate and will still be read as a promise. Those depend on context that is not in any document, which is why the final edit stays with a person.
There is also a structural ceiling worth naming honestly. These systems produce probable continuations, and a distinctive voice is by definition the improbable choice. A brief can pull output away from the average, which is the entire reason the six fixes work. It cannot make a model reach for the odd word a good writer reaches for, because reaching for the improbable is the opposite of the mechanism. Expect a reliable house voice rather than a signature one, and treat that as the correct target rather than a disappointment.
If you only fix two of the six Fix symptom 01 and symptom 03. Between them they carry most of the effect. The banned list removes the specific strings that make generated writing recognisable and takes about twenty minutes to build. The samples transfer everything you could not have articulated and take about an hour. The other four make the result consistent rather than lucky, which starts to matter once more than one person is generating copy, or once volume rises to the point that nobody reads every piece closely. But a brief with a sharp banned list and three excellent samples already outperforms a page of adjectives by a wide margin. The underlying move never changes across all six. Stop describing how you want it to sound and start ruling out how you do not. |
Share your thoughts about this article.
Be the first to post a comment!