How Voice Deepfakes Power Modern Fraud
In January 2024 a finance worker at the engineering firm Arup paid out $25 million after a deepfake video call impersonated the company’s CFO. The employee sat in a video meeting with what looked and sounded like senior colleagues, approved 15 transfers and only learned the truth after checking with headquarters. Every person on that call except the victim was synthetic.
That single case answers the core question fast: voice deepfakes power fraud because they defeat the one check people trust most, the sound of a familiar human voice. A cloned voice turns a routine authorization into a payout. Attackers no longer need to guess a password when they can simply become the boss on the phone.
Why voice became the new attack surface
Voice is the new attack surface because it sits outside most security controls. Firewalls, multi-factor prompts and email filters never inspect the audio of a phone call, so a convincing clone walks straight past them. A three-second sample pulled from a webinar, an earnings call or a voicemail is often enough for modern cloning models to reproduce a person’s tone, accent and cadence.
The volume proves the shift is real. Deepfake fraud attempts in contact centers rose more than 1,300% in 2024, from roughly one a month to seven a day, according to the Pindrop 2025 Voice Intelligence and Security Report. When an attack method gets 40 times more common in a single year, it has moved from novelty to standard operating procedure for fraud crews.
How voice cloning enables CEO fraud and BEC
Voice cloning supercharges CEO fraud and business email compromise (BEC) by adding a verification step the attacker controls. The classic BEC email asks an employee to wire money to a new account. Staff are now trained to be suspicious of that email, so the fraudster adds a follow-up phone call in the executive’s cloned voice to confirm the request. The call is the trap, not the safeguard.
The pattern is not new. A UK energy company lost about €220,000 to a deepfaked CEO voice call in 2019, when a manager believed he was speaking to his German parent-company chief and transferred the funds to a supposed Hungarian supplier. That case is six years old. The tools have gotten cheaper, faster and far more accurate since then, while the playbook has stayed the same: urgency, authority and a voice you recognize.
The Arup loss shows how far the technique scales. A single manipulated video meeting moved $25 million. The mechanics were identical to the 2019 energy-firm scam, only the production values and the payout had grown.
The anatomy of a voice deepfake scam
A voice deepfake scam follows a repeatable four-step structure. First, reconnaissance: the attacker studies the target company, its executives and its payment process, often using LinkedIn and public recordings. Second, voice capture: a short clip of the executive speaking is scraped from a podcast, conference talk or media interview.
Third, the pressure call: the cloned voice reaches an employee in finance or accounts payable with an urgent, confidential request, usually a wire transfer or a gift-card purchase tied to a fake acquisition or tax matter. Fourth, the payout and disappearance: money moves to a mule account and is layered through several banks within hours. The 2019 energy-firm funds were routed onward almost immediately, which is why recovery is rare. Speed is the whole design.
Because each step relies on public information and cheap tooling, the barrier to entry keeps falling. Security teams tracking these campaigns increasingly rely on specialized deepfake tools to flag synthetic audio before a transfer clears, rather than depending on an employee to notice something is off mid-call.
What actually stops these attacks
Process beats intuition. The most effective defense against a cloned voice is a verification rule that does not depend on recognizing the voice at all. A mandatory callback to a known internal number, a pre-agreed code word for high-value transfers and a second human approver remove the single point of failure that voice deepfakes exploit.
Payment controls matter just as much. If the Arup employee had faced a hard rule requiring out-of-band confirmation before releasing $25 million, the synthetic call would have hit a wall no amount of realism could clear. Dual authorization on any transfer above a set threshold, a 24-hour hold on new payee accounts and a policy that no executive ever authorizes a wire by phone alone would have stopped both the 2024 and 2019 cases. The technology is impressive. The countermeasure is boring, and that is why it works.
Frequently asked questions
How much audio does an attacker need to clone a voice?
Modern cloning models can produce a usable imitation from a few seconds of clear speech. Executives who speak on podcasts, webinars, earnings calls or in media interviews hand attackers plenty of source material. There is no practical way to keep a public-facing leader’s voice fully private, which is why defenses focus on verifying requests rather than protecting the voice itself.
Are voice deepfakes really being used against companies, or is this hype?
They are in active use. A UK energy firm lost about €220,000 to a cloned CEO voice in 2019 and the engineering firm Arup lost $25 million to a deepfake video call in January 2024. Contact-center deepfake attempts rose more than 1,300% in 2024. These are documented losses, not theoretical scenarios.
Can employees tell a deepfake voice from a real one?
Increasingly, no. Cloning quality has improved to the point where trained staff on live calls cannot reliably distinguish synthetic from real speech, especially under time pressure. Relying on human ears is a losing strategy. Verification steps and payment controls that work regardless of how convincing the voice sounds are the dependable defense.
What is the difference between BEC and a voice deepfake attack?
Business email compromise (BEC) uses fraudulent emails to request payments or data. A voice deepfake attack adds a cloned phone or video call to confirm that request, defeating employees who learned to distrust the email but still trust a familiar voice. The two are often combined in a single campaign for a stronger effect.