A voice can be cloned from a very short clip, so the defences that work do not depend on hearing the fake: a family code word, a callback, and a pause.
- Do not trust the voice. Researchers have copied voices from short recordings, and the US Federal Trade Commission has warned families not to trust a voice on its own.
- The scam follows a script. A loved one in distress, an authority who takes over, a demand for secrecy, a deadline, and a payment that is hard to reverse.
- Faces are not proof either. At Arup, a video call full of synthetic colleagues dissolved an employee’s suspicion and led to large transfers.
- Agree a family code word. Share it out loud rather than in a phone note or chat, and ask for it in any emergency call.
- Hang up and call back. Use the number already saved in your phone, and treat urgency and requests for gift cards, crypto or wire transfers as the warning.

Much of this site teaches you to recognise synthetic things. This page assumes you will not recognise it — because the people who lose money to voice clones are not careless, they are correct about everything except one voice. So the defenses here do not rely on your ears. They rely on a word, a callback, and a pause.
How little it takes
In 2023, Microsoft researchers demonstrated a model that could reproduce a person's voice — pitch, cadence, speaking style — from roughly three seconds of recorded audio. Not a studio session. Three seconds: a birthday clip, a voicemail greeting, a few words in someone else's livestream.Source: Microsoft Research, VALL-E — “Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers”, read at source 10 Sep 2026: it can “synthesize high-quality personalized speech with only a 3-second enrolled recording of an unseen speaker as an acoustic prompt”. January 2023. Microsoft's own description: high-quality personalised speech from “only a 3-second enrolled recording of an unseen speaker”. Trained on 60,000 hours of speech from over 7,000 speakers; code not released. A 2024 successor claimed human parity.
That is the entire technical story, and it is why "just listen carefully" stopped being advice. The US Federal Trade Commission put it plainly in a consumer alert as early as March 2023: scammers are using AI to sharpen family-emergency schemes, and you should not trust the voice.Source: US Federal Trade Commission, Consumer Alert, “Scammers use AI to enhance their family emergency schemes”, March 2023. A dated alert, cited here for when the warning was first issued rather than as current guidance. Checked 16 Sep 2026.

Stop treating a familiar voice as proof of who is calling: a short public clip can be enough raw material to copy it.
The call, beat by beat
Nearly every version runs the same five beats. Learning the shape is more useful than learning the symptoms, because the shape does not change when the technology improves.
Someone you love, in distress. An accident, an arrest, a hospital, a hostage. The voice is right, because the voice was cloned from something they posted.
A second speaker takes over: a lawyer, an officer, a doctor, a manager. Authority arrives to do the asking, so the "loved one" never has to answer questions.
Do not tell anyone. It is embarrassing, it is under investigation, it will make things worse. Secrecy exists for one reason: it prevents the one phone call that ends the scam.
It must happen now, within the hour, before the court closes. Urgency is not a detail of the story. Urgency is the attack — it exists to prevent verification.
Wire transfer, gift cards, cryptocurrency, or a courier who comes to the door for cash. Every one of them is chosen because it is hard to reverse.
Learn the script rather than listening for glitches: secrecy, a deadline and a payment that is hard to reverse are warnings whatever the voice sounds like.
It scaled to boardrooms
Early in 2024, a finance employee in the Hong Kong office of the engineering firm Arup joined a video call with people who looked and sounded like the company's chief financial officer and other staff. All of them were deepfake re-creations. Fifteen transactions followed, totalling about US$25.6 million (HK$200 million). Hong Kong police described the scam in February 2024 without naming the company; in May 2024 Arup confirmed that fake voices and images were used and that none of its internal systems were compromised.Source: CNN Business, 16 May 2024, reporting the February 2024 police account and Arup’s May confirmation; read at source 17 Sep 2026: “He subsequently agreed to send a total of 200 million Hong Kong dollars — about $25.6 million.” · “The amount was sent across 15 transactions, Hong Kong public broadcaster RTHK reported, citing police.” · Arup: “we can confirm that fake voices and images were used” and “none of our internal systems were compromised”.
The employee did what training says to do — he was suspicious of the initial email. The video call is what dissolved the suspicion. That is the lesson worth carrying home: a face and a voice are no longer verification, at any scale, in any setting.
The pattern now has a national price tag. The FBI's 2025 Internet Crime Report gives artificial intelligence a section of its own — the 2024 report had none — and counts 22,364 AI-related complaints and $893,346,472 in losses, including more than $30M in business email compromise involving AI and almost $13M in AI-involved employment scams. People aged 60 and over reported around $7.7 billion in losses of all kinds, more than any other age group.Source: FBI Internet Crime Complaint Center, 2025 IC3 Annual Report, read at source 17 Sep 2026: “In 2025, IC3 received more than 22,000 complaints reporting AI-related information. Adjusted losses of these complaints exceed $893 million.” (section count: 22,364 complaints, $893,346,472) · “In 2025, businesses reported losses over $30 million to BEC scams involving AI.” · “In 2025, victims reported losses of almost $13 million to AI-involved employment type scams.” · “60+: 201,266 complaints, $7.7 billion in losses.”
At work as at home, a face and a voice on a call are not verification: confirm any request for money through a channel you already know is genuine.
Four defenses that still work
None of these require you to detect anything. That is the point.
One word or short phrase, known only to your household. Anyone calling in an emergency must be able to say it. Choose something unguessable and unposted — not a pet's name, not a street you have written about. Say it out loud to each family member; never store it in a phone note, a group chat, or anywhere a breach could reach. Then tell the eldest members first, in two sentences: if anyone calls sounding like us in trouble, ask for the word, then hang up and call us back.
Any unexpected request for money or secrecy ends the same way: you hang up and call the person back on the number already saved in your phone. Never a number the caller gives you. This works in both directions — a boss, a bank, a government office. A legitimate caller has no reason to object to being called back.
Stop treating panic as evidence that something is real. The script is engineered to keep you from pausing, because a five-minute pause is fatal to it. Real emergencies survive verification. If someone is working hard to stop you from checking, that effort is the finding.
Wire transfer, gift cards, crypto, or a courier for cash — the channel is chosen for irreversibility, not convenience. No hospital, court, or police department takes bail in gift cards. When the request reaches this beat, the beat itself is the answer.
A fifth, quieter defense: shrink the sample. Public video and audio is the raw material. Locking down who can see family posts will not make cloning impossible, but it makes your household a slower target than the next one.
- Do not confirm names or relationships. Answering "which grandchild?" hands the script its next line.
- Do not send anything. Do not stay on the line.
- Hang up. Call the person directly on their saved number. If they do not answer, try a text, another relative, or wherever they are supposed to be.
- If money has already moved, call your bank immediately — speed is the only lever left — and report it to the FTC at reportfraud.ftc.gov (US) or your national fraud reporting line.
Set up the code word and the callback rule before you need them, and explain them to the eldest members of the family first.
If the fake is an intimate image rather than a voice, US law now gives you a removal route. Under the TAKE IT DOWN Act, platforms must offer a way to ask for non-consensual intimate images to be taken down, and must remove them within 48 hours of a valid request; the Federal Trade Commission has enforced this since 19 May 2026 and takes complaints about platforms that fail to act at TakeItDown.ftc.gov.Federal Trade Commission, FTC Begins Enforcing the TAKE IT DOWN Act, 19 May 2026, read at source 22 Sep 2026: covered platforms must give people a way to request removal and “remove those intimate images, and known identical copies, within 48 hours of a valid request.” It adds that the FTC “has launched TakeItDown.ftc.gov, a website allowing victims and survivors to submit complaints about platforms that have failed to act on valid requests”.
Voice synthesis is not the villain here. The same technology restores speech to people who have lost it, narrates books, and crosses languages. What changed is that impersonation became cheap, and the institutions we trust to verify identity have not caught up. Until they do, the defense is a household habit, not a gadget. Set the code word tonight; it takes fifteen minutes and it outlives every model release.