top of page

AI Voice Cloning and Deepfake Scams: What Your Employees Need to Spot

Writer: echoudhury77
echoudhury77
7 minutes ago
4 min read

In May 2024, an employee at engineering firm Arup joined what looked like a routine video call with the company's CFO and several colleagues. Everyone on screen looked and sounded right. Over the course of the call, the employee was instructed to make 15 transfers totaling $25 million. Every person on that call except the employee was an AI-generated deepfake — synthetic video and cloned voices built from publicly available footage of real executives.


That single call is now one of the most cited case studies in corporate security, and it illustrates a shift every business needs to understand: the newest attack vector doesn't target your firewall. It targets your employees' trust in a familiar voice.


The Scale of the Problem

This isn't a rare, exotic threat anymore. A few data points make that clear:

  • The FBI's 2025 data attributes $893 million in losses across more than 22,000 complaints to AI-enabled fraud, with impersonation and investment scams making up the bulk of it.

  • AI-enabled scams surged an estimated 1,210% during 2025, with voice-based attacks on businesses specifically climbing roughly 1,300%.

  • Perhaps most concerning: research on voice-clone calls found that 77% of people who receive a convincing AI-cloned voice call end up losing money — a success rate that would make any other attack vector a top priority overnight.

  • Fewer than 5% of voice-clone victims file an official report, meaning the true scale is almost certainly larger than these numbers suggest.


The barrier to entry for attackers has also collapsed. Cloning a convincing version of someone's voice now takes as little as three seconds of clear audio — easily pulled from a company webinar, an earnings call, a podcast appearance, or a video posted to LinkedIn. Every executive or manager with a public-facing voice is a potential source.


Why This Targets People, Not Systems

Traditional phishing relies on a bad link or a suspicious attachment — something a spam filter or an alert employee can catch. Voice and video deepfakes are different: they exploit the deepest trust signal humans have, the sound of a familiar voice giving an instruction that sounds completely in character. There's no malicious file to scan, no link to hover over. The entire attack happens in a conversation.

That's exactly why this threat belongs in security awareness training, not just in your technical security stack. No firewall stops a phone call.


What Employees Should Actually Listen and Watch For

Real-time detection is less about pixel-perfect analysis and more about noticing when a conversation deviates from how a normal one behaves:


On voice calls:

  • Flat or metallic tone. Cloned voices often lack the natural variation of real speech — the way a real voice speeds up slightly under stress, drops in register when hesitating, or includes small breath sounds between sentences.

  • Unnatural pacing. Pauses that don't line up with normal thought patterns, or a rhythm that feels just slightly "off," even if the words themselves sound right.

  • Background inconsistency. Ambient noise that shifts abruptly or doesn't match what you'd expect from where the caller claims to be.

  • Vocabulary that doesn't fit. Word choices or phrasing that a colleague wouldn't normally use.


On video calls:

  • Lip-sync drift. Deepfake video often shows mouth movements that lag or lead the audio by roughly 100–300 milliseconds — subtle, but noticeable if you're watching for it.

  • Unnatural blinking, lighting, or facial movement that doesn't track naturally with head motion.


In either case, the biggest tell is behavioral, not technical:

  • Unusual urgency — "this needs to happen right now."

  • A request to bypass normal approval steps, especially around money movement or credentials.

  • An authority figure asking for something out of character for how they normally operate.


Build Verification Into the Process, Not Into Memory

Employees shouldn't be expected to detect a well-produced deepfake by ear alone — some are genuinely convincing. The real defense is a verification step that doesn't depend on the call itself:


  1. Callback authentication. Never act on a financial request from a single call or video, no matter how convincing. Hang up and call back on a previously known, verified number — not one provided during the suspicious call.

  2. Pre-shared code words. Establish a private verification phrase with executives and finance staff that isn't used anywhere public. A cloned voice can mimic tone and cadence, but it can't produce a phrase it's never heard.

  3. Dual approval for money movement. Any wire transfer or payment change should require a second, independent confirmation through a different channel — never approved on the basis of one urgent call.

  4. A "pause" culture, not a "comply" culture. Employees need explicit permission from leadership to slow down and verify, even when the person on the call outranks them. The Arup employee later said they followed the instructions specifically because everyone on the call looked and sounded legitimate — the org chart worked against them, not for them.


Where This Belongs: Ongoing Security Awareness Training

Deepfake and voice-clone awareness shouldn't be a one-time memo — it needs to live inside your regular security awareness training alongside phishing simulations, just like any other social engineering vector. Employees who handle payments, vendor changes, or executive requests should specifically practice recognizing urgency-based pressure and know the verification steps cold, before they're tested by a real attempt.


As the cost of producing a convincing fake keeps dropping, this is quickly becoming one of the highest-return investments a business can make in its security program — not because it requires new hardware, but because it requires a habit.

Firestorm Cyber builds deepfake and social-engineering awareness into its security training programs and backs it with 24x7 live support, so if something does feel off, your team has someone to call immediately rather than guessing under pressure.


Sources: FBI Internet Crime Complaint Center (2025 data via EyeSift); Adaptive Security, deepfake detection research; CNN, Arup deepfake scam coverage (May 2024).

 
 
 

Comments


bottom of page