Research Whitepaper

Why AI doesn't
sound like you

Getting a consistent tone of voice from AI: what the evidence actually shows

Executive summary

Generative AI can write fast. Whether it keeps sounding like you is the question that matters.

In a market where every competitor prompts the same models with the same adjectives, AI does not just fail to capture your voice. It quietly erases the distinctiveness you spent years and budget building, and pulls your writing toward the bland middle of everything it has ever read. Brand is one of the few defensible advantages an organisation has. An "average" voice, produced at scale, liquidates it.

Most advice on fixing this is weak and lacks evidence. So we did the work properly: we synthesised the field through a structured search across six writing professions, graded every claim by the quality of its evidence, then tested the strongest techniques on our own published copy – scored blind, with the failures kept in.

Finding 01 · Show it

Give the system contrastive good and bad examples of the voice. The highest-return technique available to everyone.

Finding 02 · Bake it in

Fine-tune a model on a large set of examples of the voice. The only technique backed by a controlled experiment – and the least used.


The multiplier · Keep teaching it

No technique keeps working on its own. The system has to capture human corrections and turn them into permanent rules, so it gets better over time. Fine-tuning is out of reach for many organisations, so for most the smarter bet is a system of machine drafting and human correction that improves with use.

If you read one page

AI cannot automate a nuanced voice, and no tool delivers a perfect one – treat any vendor who claims otherwise with suspicion. What works is a system you build and improve: show the model examples, keep them loaded, and feed every human correction back in. The strategic prize is bigger than tone of voice – this is the template for every judgement-heavy task you will hand to AI. The leadership job is to resource the system and protect the human at the point where taste is applied.

01 · Why this matters, and who it's for

The strategic case before the technical one

This paper is written first for the people who carry the strategic risk – chief marketing officers, brand and communications leaders, and the executives deciding how AI runs across their organisation. It is also built to be used by the people who implement: heads of content and brand, comms leads, and the AI specialists who build the systems. Leaders should read the framing and the verdict; practitioners will want the method and the stacks. Both are here.

There are two reasons a leader should care, and neither is about phrasing.

The cost of average is paid twice

When AI drifts to the generic, you pay once to generate a draft that sounds like everyone else, and again in the human hours spent editing it back into something that sounds like you. At scale, that is real money spent making your brand less distinctive. The risk is not a clumsy sentence; it is the slow erosion of the thing that makes customers choose you over a competitor who looks identical on a spec sheet.

Tone of voice is the canary

It is the clearest, most testable example of a whole class of work organisations are now trying to hand to AI – judgement-heavy, taste-driven, and hard to write down. Get the approach right here and you have the template for the harder cases: legal judgement, design taste, strategic argument, anywhere expertise resists a simple instruction. The lesson generalises cleanly: AI does not replace the expert. It shrinks the number of decisions the expert has to make, and raises the value of the ones that remain.

The reframe

Voice consistency is not a phrasing problem. It is an architecture problem – decided by where you put the voice (the prompt, the persistent context, the model's weights, or a checking step after generation) and by whether the system keeps learning from human corrections.

What this paper does differently. Most writing on this subject is confident and unevidenced. We took the opposite approach: we graded the strength of every claim, including the weak ones; we ran a blind test on our own copy; and we have kept our own failure in the paper because it taught us something. Treat that as the standard to hold any AI claim to – including ours.

02 · The average is the enemy

Left alone, every model writes toward the middle of everything it has read

What we mean by tone of voice

Your voice is who you are – it stays constant. Your tone flexes by audience and moment: you do not write to the board the way you write to customers. And style is the mechanical layer – sentence case, "who" not "whom", en dashes not em dashes. The catch: you do not have one tone to capture. You have one voice and a hundred tones, most of which have never been written down. That is the first reason this is hard – it looks simple from the outside and turns out to be all taste and tacit knowledge on the inside.

Ask a marketing team for a metric on how well their AI produces drafts in their tone of voice, and the answer is often a speed metric: drafts in minutes, editing time down. Faster drafting tells you nothing about whether the tenth draft will sound like the first.

The default approach fails for a mechanistic reason. To a model, a word like "punchy" is not a feeling; it is a coordinate in a vast mathematical space, and the model has no fixed definition of it. It reinterprets the word on every run, and where the instruction is ambiguous it falls back on the statistical centre of its training data – the average of the average. The words also mean different things to different people: ask for "wry" and you are picturing Margaret Atwood while the model reaches for Jeremy Clarkson. The result is the register most readers now recognise on sight – competent, smooth and interchangeable. When every brand asks for "confident but approachable", every brand gets the same prose.

×3

The average of the average of the average

Call it the average – the pull every model has toward the middle of everything it has read. Left alone, a model gives you the least surprising, least specific, least defensible version of the thing. Defeating that pull is the whole task.

The business cost, named plainly. A distinctive voice is an asset on the balance sheet of attention – it is how people recognise you before they have read the logo. Sounding average is not always failure: a Santander sounds like a bank and does perfectly well, because it does not compete on brand. But where brand is part of how you win – a Monzo, an Innocent Drinks, a Brewdog – average is exactly what you cannot afford. Spend a year letting AI regress that voice to the mean and you have not saved money; you have diluted an asset while paying for the privilege, twice over.

It is worth being honest about how much is in play. Every answer a model gives is shaped by a stack of variables at once – its training data, the cut-off date of that training, the files you have loaded, the context window, and whatever you have said to it before – and it is, in effect, guessing how much weight to give each one. Ask the same thing twice and you will get two different answers. Consistency is not the default; it is something you have to engineer.

The target itself is hard to name. Writing in a voice is usually filed as a "soft" skill, but it is one of the hardest things people do – built over years, carried largely in the unconscious, applied by feel. Ask an expert reader why a draft is wrong and you often get no metric and no consistent reason: they just know. Taste is real knowledge that resists being written down. (The philosopher Michael Polanyi called this tacit knowledge: we know more than we can tell.) This is why showing the model examples beats describing rules to it – examples carry the signal that rules flatten.

Define the goal precisely

Voice consistency is the systematic reduction of variance: the same brief, run ten times, should produce drafts you and your readers would recognise as you, and not an AI system. Some techniques raise the quality of a single output; a different set narrows the variation between outputs. The two effects are routinely confused – including by the techniques' own advocates.

03 · The framework

Shrinking the decision space

Voice is a property you can encode at four layers of an AI system – prompt, context, weights, output gate – and the layer determines how durable that encoding is. Seen separately, our seven tactics look like seven tactics. Seen together they do one thing: each drags the model off the average into a narrow region where its output is close to what you would have written. A de-averaging machine.

Approach Layer Evidence What it's for
A · Describe itadjectives in the prompt Prompt Mechanistic – against The baseline to beat. A control condition, not a method.
B · Show itexamples & contrastive pairs Prompt Near-experimental The highest-return single technique.
C · Specify itstructured guide, rules, bans Prompt Testimonial, strong The workhorse: cheap, transparent, writer-controlled.
D · Persist itauto-load the spec each session Context Mechanistic, tight Delivery, not voice: makes B and C survive sessions.
E · Ground itretrieve your own prior words Context Testimonial Highest ceiling for fidelity; least proven.
F · Learn itfine-tune the weights Weights Experimental The rigorous answer for high-volume, single-voice systems.
G · Police itaudit or judge after output Output gate Vendor / mechanistic Measures drift at scale; doesn't create voice.

Four layers

Prompt · Context · Weights · Output gate. Where you put the voice decides how durable it is.

One job

The model still writes the first draft. The stronger the technique, the narrower the space, and the less the human has to do to finish.

03 · The seven approaches

Each approach, and what it actually does

B Show it

Pair on-brand examples with off-brand counterexamples. Models are pattern-matchers before instruction-followers, so the contrast fences off the right territory and the bad half marks the edges – it is the same instinct as a design mood board that says "like this, not like that". Latitude validated it by running prompts 10–15× at low temperature: good-plus-bad beat good alone. Silvestri builds it into every rule of a brand guide. This is the single highest-return move available to everyone.

C Specify it

The structured guide in three evidenced parts: annotated rules (a rule with an example sticks); quantified tone (NN/g's four dimensions written as 1–5 targets); and banned lists (Francis's ~56 terms). A document the model paraphrases is lossier than examples it pattern-matches – which is why B belongs inside C.

A Describe it

Adjectives in the prompt. Its only role in a serious system is as the control condition – the baseline every other technique has to beat.

D Persist it

Put the spec where the model cannot miss it – a Claude Project, a custom GPT, a Skill, a file read before every draft. Most between-piece drift is humans re-assembling or forgetting the prompt. Persistence is delivery, not voice. Do it regardless of what else you choose.

E Ground it

Stop describing the voice and retrieve it: condition each draft on the writer's actual prior words, queried at draft time. The newest ground and the most direct answer to "it never sounds like me". Testimonial, but the highest ceiling. Cheapest test: sentence-seeding.

F Learn it

Fine-tune so the voice becomes the model's default rather than an instruction it re-follows. The only approach backed by a controlled experiment. Trade-offs: cost, skill, and rigidity – a baked-in voice is hard to vary by channel.

G Police it

Score outputs against the spec after generation – an audit pass, a compliance score, or an LLM-as-judge. It catches drift humans miss at volume and produces a trackable number; it does not create voice, and judge models need calibrating against human labels. Most valuable for teams running many writers through one voice.

04 · How we did the research

A two-part method: synthesise the field, then test it on ourselves

We wanted to answer a single question – how do writing professionals get a consistent tone of voice out of AI, to a standard that passes expert human taste? – without relying on our own opinion. So the research has two parts. The first maps what the field already believes and grades how well it is evidenced. The second puts the strongest claims to a practical test on our own copy. This page covers the first; the demonstration follows below.

Part one · The synthesis

A structured, multi-agent literature synthesis

Rather than search ad hoc, we ran a fan-out workflow: a single human-written brief directed an orchestrating agent, which spawned specialist agents to investigate each domain in parallel, gather primary sources, and report in a common format (who · technique · why it works · source · verdict). We then synthesised across the domains by hand.

Six domains, chosen for range

We looked beyond marketing, because the people who fight hardest for a consistent voice are spread across very different trades:

  1. Journalism and newsrooms
  2. Comms, PR and exec ghostwriting
  3. Fiction and screenwriting
  4. Technical and UX writing
  5. Cross-domain tooling & evaluation
  6. Brand copywriting & voice teams

Adversarial verification

Gathering is where most "research" stops and where confident myths enter. So every named claim was checked by a separate verifier agent, working against the primary source rather than the summary, and given a verdict. Where a claim could not be corroborated, it was downgraded or cut. Two findings were kept deliberately as debunked records – a cautionary layer rather than evidence (see the section below).

The grading rubric · Five evidence grades

ExperimentalA controlled comparison was run.
OperationalReal production metrics, self-reported.
TestimonialOne practitioner: "works for me".
VendorThe seller's own claim.
MechanisticNo data, but a sound causal reason.

The limits of this method

A fan-out search finds what is written down and public; it under-weights what teams keep to themselves. Verification corroborates against sources, but cannot run the experiments the field is missing. Treat this as the most honest available map of the evidence – not as proof in its own right.

05 · What the evidence supports

The evidence base is thin – and dominated by testimony

Apply the rubric across the whole corpus and the distribution is humbling. In everything we reviewed, we found only one controlled experiment: a preprint by Marquardt and Brule (arXiv 2507.04889). They found that LoRA fine-tuning on roughly 100 synthetic, tone-labelled samples outperforms system prompting, and the gap widens as context grows. That is the entire experimental layer of this field – and it supports the least-used technique.

The strongest evidence short of that is Latitude's, because they measured variation rather than asserting quality – running prompts repeatedly at low temperature and watching the spread. That is what earns "Show it" its near-experimental grade. The strongest operational evidence sits in the feedback practices, where teams report pass rates and edit metrics from live production. Almost everything else – including most of what is repeated with great confidence at marketing conferences – is testimony or vendor claim.

This is not a reason to dismiss the field. Mechanistic and testimonial evidence is still worth acting on; it is just not proof, and it should not be sold as proof. The practical consequence is simple: trust the two well-evidenced techniques most, treat the rest as sensible defaults, and verify them on your own copy.

Two records kept as debunked

A widely circulated "40–60% less editing" figure that no source corroborates, and a popular voice-file scaffold wrongly attributed to a well-known writing tool. They stay because confident round numbers are common in this field, and treating them as marketing until proven is the correct default.

06 · The demonstration

We ran the framework on our own copy – and scored it blind

Part two of the research turns the framework on ourselves. It is a transparent worked demonstration, not a controlled study – two runs per condition where Latitude used ten to fifteen, and an LLM judge rather than a panel of humans – but every prompt, output, score and failure is recorded and reproducible.

Method, in full

01

Source material. Four articles from our own website, spanning analytical and story-led pieces.

02

The task. Write the opening from a neutral fact-brief – the raw material stripped of any voice.

03

Four conditions. A adjectives; B three pairs; C a spec; D the stack (spec + pairs).

04

A genuine holdout. Voice materials built only from the other three articles – no model saw its target.

05

Controlled runs. Every generation a fresh instance, two runs per condition, to expose variability.

06

Blind judging. All eight outputs scored against the published opening, then re-scored by a second model family.

Condition Voice match (mean /10) AI-smell tics Stability (/5)
A · adjectives 7.00 23 3.25
B · + pairs 6.75 19 3.75
C · spec 5.81 27 3.75
D · the stack 6.06 26 4.50

How to read this table

For a brand, the number that matters most is stability – the same brief should give you the same voice every time, not a great line one run and an off-brand one the next. Read that column first: the stack (D) leads it by a clear margin. The voice-match averages need that context: adjectives (A) score highest on average, but across a wide 5–9 spread, so the average hides how unreliable they are. Consistency is what a brand can build on. The five findings below explain the rest.

06 · What we found

Five readings the numbers force

1 The stack delivered consistency – which is what it promises.

Condition D was the most stable by a clear margin and produced the two best single outputs: on our two analytical pieces, the blind judge scored the stack 9/10. The model had never seen those articles. It was shown what our voice does, and reproduced it.

Verdict For the thing a brand actually needs – the same voice every time – the stack wins.

2 Adjectives were volatile, not useless.

Condition A averaged the highest voice score – and the widest spread, scoring everywhere from 5 to 9. Adjective prompts sometimes produce a good draft, but cannot produce the same draft twice – and you don't choose which run you get.

Verdict A high average you can't reproduce is no use to a brand. Volatility is the failure.

3 Examples suppressed AI smell better than rules.

B produced the fewest generic-AI tics and the tightest quality floor – every B output scored 6–8, no failures. The spec's banned list, by contrast, did not stop the tics: C had the most.

Verdict Show the model, don't just list rules at it.

4 The spec overfitted – our most useful failure.

Built from three mostly analytical articles, the spec encoded our "reframe the settled question" opening as a universal rule. On a story-led piece it ignored the story and imposed the template, scoring 4/10 against an adjective-only run that scored 8.

Verdict One spec can't cover everything – build from your full range, or scope it to one type.

5 Two judges, two verdicts.

A second judge from a different company agreed the stack produced the best work but reversed the ranking of the weaker conditions. Two careful judges, the same texts, different answers.

Verdict There is no clean scoreboard yet – treat every confident number, ours included, with care.


What it adds to the framework

Build the spec from a corpus as varied as what you publish, or scope it to one content type. Never ship a spec without examples – the pairs hold the floor. And a strong, factual, neutral brief did the quiet heavy lifting in every condition.

07 · Implementation stacks by scale

Resource it by volume and the fidelity you need

For leaders, the decision is what to resource. Do not buy the promise of a perfect, hands-off voice engine – the evidence says it does not exist. Fund the system and the work of improving it instead, and protect the human at the point of taste. The right build depends on your volume.

Individual / Boutique

C + B + D

A structured spec with a strong contrastive set and a banned list, persisted so it loads every session. Add sentence-seeding. Skip fine-tuning and judges; the volume doesn't justify them.

Small team / Agency

+ G, centralised

Everything above, centralised – one shared spec, one example library – plus an LLM-as-judge scoring drafts against the shared rubric. The failure mode at this scale is many writers drifting apart. Calibrate the judge first.

Enterprise / High volume

+ F, improving

Add fine-tuning, keep G as the governance gate, and keep feeding human corrections back into both. Fine-tuning lowers variance across thousands of outputs. Keep C as the human-readable source of truth, or the system becomes un-auditable.

The build order, for everyone

  1. 01Spec first (C) – use an eight-section structure, with a banned list in it from day one.
  2. 02Persist it (D). Make it load before drafting, every time.
  3. 03Capture corrections until the same ones keep recurring – the 10–15 rule – before refining anything.
  4. 04Add contrast pairs (B) where the corrections show the model keeps missing. Source the spec from the full range of what you publish.

08 · The system has to keep learning

Every approach is static. Voice work compounds only if corrections become rules.

Pass rate

70–80%

The share of content approved without significant change – the reported working range, and the metric that should replace edit time.

The 10–15 rule

10–15

Log every manual edit; when the same correction recurs across ten to fifteen pieces, promote it to a permanent rule. Zero-edit pieces become anchor examples. In short: when the model gets something wrong, show it the edit and let it learn.

Beneath the metrics sits a division of labour every successful practitioner keeps and none has tested removing: the system produces structure and first drafts; the human owns judgement – the final 30%, the ear, the call on what is actually good. But notice where the raw material for improvement comes from: the human's corrections are the training data for the spec. Pushback only compounds if you capture it.

It is worth stating the real promise plainly, as a view rather than a finding. The point of these systems is not that they automate the writing. It is that they take on the lower-order work – assembling, drafting, structuring a first pass – and free the human for the higher-order work: judging, deciding, choosing what is actually good. There is a mirror in this: to make an AI hit your voice you have to say what your voice is, which forces a sharper question than you would otherwise ask. This is the pattern leaders should take from a marketing problem into every other corner of the business.

Hold it as a point of view

Our own grading marks this division of labour as asserted, never measured – so take it as a point of view, not a result. Treat it as a governance principle rather than a proven optimum, because that is what it is.

09 · The verdict

Show it, or bake it in.

If you want an AI to sound like your tone of voice guide, the evidence supports two main techniques. Contrastive examples are the highest-return technique available to everyone; fine-tuning is the proven technique for those with the volume to justify it.

The structured spec is the workhorse that organises the examples; persistence keeps the whole thing alive; and capturing corrections – turning them into rules – is the difference between a system that improves and one that keeps starting from scratch. Whichever you choose, the job underneath is the same: drag the model off the average, shrink the decision space, and put a human at the point where taste is applied.

So the one thing to take away: stop describing how you sound, and start showing the model the difference. Do this, not that. Then keep teaching it.

The gap that remains

One thing this field still lacks, including in our own demonstration: a direct, repeatable measure of voice consistency itself. Everyone measures proxies – speed, edit time, scores against a reference. If taste is tacit and high-dimensional, there may be no single number to find. Until then, every confident claim in this space, ours included, deserves the grading we gave it.

What's next – and how to work with us

We are continuing this research: widening the demonstration, testing grounding and fine-tuning on real client voices, and pushing at the measurement problem itself. That is not a reason to wait. The move available to every leader today is the same one we made: run the test on your own copy, keep the records, and keep teaching the system. If you would like the slides, the full research, or a conversation about building this into how your organisation works, talk to us.

Build this into how your organisation works

Get the slides, the full research, or a working session on making AI sound like you – not the internet.

10 · References

Sources, graded as we found them

Primary evidence

  1. 01Marquardt & Brule, Fine-tuning beats prompting for tone consistency, arXiv 2507.04889 – a rare controlled experiment on tone of voice.
  2. 02César Miguelañez / Latitude – contrastive GOOD/BAD example blocks, validated by repeated low-temperature runs.
  3. 03Katie Parrott / Every – the 8-section AI style guide, persisted in a Claude Project.
  4. 04Chris Silvestri / Copyhackers – the five-part principle: every voice rule carries a correct and an incorrect example.
  5. 05Kate Moran / Nielsen Norman Group – four tone dimensions on 1–5 scales (built for analysis; prompt use is a practitioner adaptation).
  6. 06Will Francis – the ~56-term banned list and structural rules.

Operational practice

  1. 07SocialPilot – edits logged into rules after 10–15 posts; zero-edit posts become anchors.
  2. 08Jolissa Skow / Search Engine Land – brand voice document; a ladder of prompts → RAG → PEFT → fine-tuning; pass-rate and deviation metrics.
  3. 09David Johnson-Igra / Scribes – performance-weighted knowledge graph of an executive's own words, queried at draft time.
  4. 10Ran Isenberg – four-file voice system; the orchestration file forces "read voice files first".
  5. 11The Wall Street Journal (Ed Hyatt) – shared prompts routed through the existing house-style desk.

Vendor – graded as such

Writer.com Voice; Jasper Brand IQ; Frontitude Audit.

Published by Brilliant Noise.

A research whitepaper · June 2026 · Public · Evidence-graded.

© Brilliant Noise 2026. All rights reserved.