Voice cloning

Speech Technologyalso: voice replicationalso: synthetic voice cloningalso: audio deepfake

In one sentence

Voice cloning builds a synthetic copy of one named person's voice out of a recording of them talking, and that copy can then be made to say things the person never said.

3 seconds, 85 percent match McAfee, The Artificial Imposter, 2023Last reviewed 30 July 2026

Not to be confused with Text to speech, or Voice biometrics and anti-spoofing.

Definition

Voice cloning takes a recording of someone talking and builds a synthetic copy of how they sound.

The copy will then say whatever it is told to, including things the real person never said.

There are two broad approaches in use today, and the line between them is drawn by how much audio each one demands.

Zero-shot, or instant, cloning

  • A few seconds of reference audio is enough. From it the model derives a speaker embedding, which is a compact numeric fingerprint of vocal identity, and synthesis is then conditioned on that fingerprint.
  • Results range from rough to startlingly good. Needing so little data is exactly what turns this into a fraud vector.

Fine-tuned, or professional, cloning

  • Wants clean studio audio measured in minutes or hours, plus a training run over it.
  • What comes back is more faithful, holds together better over long content, and reproduces the prosody that makes a particular speaker recognizable.
  • This is the approach behind legitimate commercial voice licensing.

The distinction that carries the most weight, and it is not a technical one

  • Consented cloning. Whoever owns the voice said yes, on paper if the process is being run properly, with the scope and the duration written down. Audiobooks, localization, accessibility and brand voices make up a real and growing market.
  • Non-consented cloning. Everything else, from commercial use nobody authorized all the way through to impersonation fraud.
  • Neither side uses different technology. Governance is the whole of the difference.

Safeguards in current practice

  • Verification at enrollment: the speaker reads out a phrase the system picks, live on the spot, instead of handing over a file.
  • Watermarks embedded in whatever gets generated. Meta AudioSeal is open source. Google SynthID-Audio is built to outlast metadata stripping. C2PA Content Credentials attach a signed provenance manifest, which stripping does defeat.
  • Because those two break in different ways, current best practice runs a watermark and a manifest together.
  • Blocklisting of public figures, plus detection when someone tries to clone a protected voice.

Common misconception

That a trained ear will catch a cloned voice. Usually it will not, and trusting your own ear is the weak point rather than the defense. Which is why the things that actually help are procedural: ring back on a number you already had, confirm through a second channel.

Why it matters commercially

Both of these hold simultaneously. Cloning is a genuinely useful capability with a real market behind it, and it is also where most of the category's reputational and regulatory exposure originates. Sell it and you need a consent process that survives scrutiny. Buy a voice agent of any kind and you inherit a public that has formed its view of synthetic speech from fraud stories you had nothing to do with.

In voice specifically

Text carries no equivalent exposure. A forged email is judged on the address it arrived from, while a forged voice is judged on whether it sounds right, and a listener has spent a lifetime being trained to accept that as proof of identity.

Where AsqVox fits

Cloning is not among the Orb's live capabilities. This entry exists because the word shapes how the whole category is perceived, and because anyone researching voice agents will meet fraud coverage sooner or later and deserve the context around it.

Visual

Same technology, two entirely different things

Same technology, two entirely different thingsOne voice sampleA syntheticvoice that cansay anythingThrough the consent gatelegitimateAudiobook narrationLanguage localizationAccessibility and voice restorationWritten consent, with scope and durationLive verification at enrollmentWatermarkingAround the consent gatefraudFamily impersonation scamExecutive payment fraudVoice authentication bypassDeepfake fraud attempts up more than 1,300 percent in 2024, PindropProvenance toolingwhat a later claim of authorship rests onMeta AudioSeal, open source watermarkGoogle SynthID-Audio, survives metadata strippingC2PA Content Credentials, signed manifest, does not survive strippingPair a watermark with a manifest, they fail differently

The model does not know whose voice it is. The consent record does.

Three seconds of audio has been reported to produce an 85 percent voice match. Both routes run the same model on the same sample and arrive at the same output, which is why the gate is the only thing on the diagram that separates them.

Statistics

Every figure carries its source and year. Vendor numbers are labelled as vendor numbers, and where no reliable figure exists this page says so rather than borrowing one.

McAfee's "The Artificial Imposter" study, which surveyed 7,054 people in seven countries, put the threshold at three seconds of audio for a clone matching the original voice 85 percent. It also found 70 percent of those surveyed doubted they could separate a clone from the genuine article.

3 seconds, 85 percent matchvendor claim

McAfee, The Artificial Imposter, 2023 - One vendor-commissioned study is the whole provenance here. The three-second number gets repeated everywhere with the attribution stripped off, and it should never travel without it.

Working from a pool of over 1.2 billion calls it had analyzed, Pindrop counted a rise of more than 1,300 percent in deepfake fraud attempts through 2024, moving from roughly one a month to seven a day.

Up more than 1,300 percentvendor claim

Pindrop, 2025 Voice Intelligence and Security Report, 2025 - What this counts is traffic through Pindrop's own systems, so treat it as vendor telemetry and not as a census of the industry. Worth remembering too that it grew from a very small base.

Deloitte projects that fraud losses in the United States enabled by generative AI could hit USD 40 billion in 2027, against USD 12.3 billion in 2023, which works out at 32 percent compound annual growth.

USD 12.3bn to 40bnanalyst forecast

Deloitte Center for Financial Services, 2024 - A forecast from a credible analyst house, and a forecast is not a measurement. Always cite it with its date and say plainly that it is a projection.

The transparency duties in Article 50 of the EU AI Act, deepfake disclosure among them, take effect on 2 August 2026.

2 August 2026industry range

EU AI Act, Article 50, 2026 - A date set in statute rather than anything measured. Point readers at the regulatory text itself instead of paraphrasing it, and re-check it before you build a plan on it.

Rules under India's DPDP Act 2023 were notified in November 2025, phasing in to full compliance by 13 May 2027. Consent sits at the center of the regime, penalties are fixed at up to Rs 250 crore instead of being pegged to turnover, and biometric or voice data gets no category of its own.

Full compliance by 13 May 2027industry range

India Digital Personal Data Protection Act 2023, Rules notified November 2025, 2025 - That missing biometric and voice category is a real hole where voice is concerned, and it is worth putting on the table early in any India deployment.

There is no clean independent yardstick for how well synthetic speech detection or anti-spoofing actually works. The figures in circulation come from the detection vendors themselves, with no NIST-style head-to-head evaluation publishing results anyone can compare.

-no reliable figure

A conspicuous hole given what is at stake. It also means every vendor accuracy number you are shown is self-reported, with no peer result to hold it against.

Examples

In practice

One platform accepts enrollment only when the speaker reads a phrase live, refuses uploaded files outright, screens every sample against a blocklist of public figures, watermarks all generated output and ships a C2PA manifest with each file. Months later a customer contests who produced a particular clip. Re-encoding has not destroyed the watermark, and the manifest still reports the parameters used to generate it. Neither of those decides the argument, though. The consent record does.

The everyday version

Cloning is copying how a specific person sounds. Done right, it is one narrator voicing the same audiobook in eight languages, paid and willing. Done wrong, it is the call from a voice indistinguishable from your finance director, wanting a payment pushed through today. Identical technology on both sides. All that separates them is whether permission was ever given.

Usage

Who says it

  • Text-to-speech vendors, carefully, and usually in the same breath as their consent and safety process.
  • Fraud and security teams file it under threats, and the words they reach for are audio deepfake or synthetic voice fraud.
  • Legal and compliance, in the context of likeness rights, what a performer agreed to, and what has to be disclosed.

Where it turns up

  • Alongside clauses on verification and consent, watermarks, provenance standards, blocklists, likeness rights and the terms governing commercial use.
  • Put the same word in front of a bank, or anyone running voice authentication, and it becomes an attack surface to defend rather than a capability to buy. Same vocabulary, different conversation, and often a different buyer in the room.

Common misuse

  • Stretching the term to cover any synthetic voice at all. An off-the-shelf synthetic voice is a copy of nobody.
  • Advertising instant cloning while the consent mechanism sits somewhere in the small print, which invites precisely the scrutiny this category can least afford.
  • Offering detectability as reassurance. The evidence runs the other way, and overclaiming here is expensive in credibility.

Questions people ask

How much audio does voice cloning need?

Zero-shot cloning works from a few seconds of reference audio. McAfee's Artificial Imposter study, run across 7,054 respondents in seven countries, put three seconds at an 85 percent voice match. The fine-tuned or professional route asks for clean studio recordings measured in minutes or hours plus a training run, and pays that back in fidelity and in holding together over long content.

Is voice cloning legal?

It turns entirely on consent and on jurisdiction. With the speaker's written agreement in hand, and scope and duration set out in it, you are in a legitimate market: audiobooks, localization, accessibility, brand voices. Without that agreement you are somewhere on a line running from commercial use nobody authorized to outright impersonation fraud. Neither side uses different technology, so the consent record is the entire distinction.

Can you tell a cloned voice by listening to it?

Usually not. Seventy percent of McAfee's respondents doubted they could separate a clone from the genuine article, and trusting your own ear is the weak point rather than the defense. What works is procedure: ring back on a number you already had, confirm through a second channel, and read urgency as a warning rather than a reason to hurry.

How do you prove a piece of audio was synthetic?

Two tools, used together. A watermark from Meta AudioSeal or Google SynthID-Audio lives inside the audio itself, and SynthID-Audio is built to outlast metadata stripping. C2PA Content Credentials add a signed provenance manifest, which stripping does defeat. They are paired precisely because they break in different ways. Bear in mind that no NIST-style independent evaluation of detection accuracy has been published, so vendor accuracy claims have no peer result to be checked against.

Share this definition

Last reviewed 30 July 2026. Written and reviewed by Dhruv Dholakia, founder of AsqVox.