A Ferrari executive stopped a multimillion-dollar AI voice scam last year with one simple question.
In July 2024, an unnamed executive at the Italian automaker received WhatsApp messages and then a phone call that appeared to come from Ferrari CEO Benedetto Vigna, 58. The caller’s voice and accent matched Vigna’s perfectly because scammers had cloned them. When the impersonator requested a transfer, the executive did not comply immediately. Instead, he asked the caller to name a book Vigna had recommended just days earlier. The scammer hung up.
Valeria Contreras, marketing manager at Zepo Intelligence, said that a voice can now be cloned from as little as three seconds of audio.
That brevity means a single TikTok clip, Instagram Story, or voicemail message can supply enough material for a convincing fake. Ordinary families and mid-level employees are increasingly in scope because publicly available social media content provides scammers with abundant source audio. The targets are no longer limited to celebrities or senior executives.
Voice deepfakes rose 680% year-over-year, according to figures supplied by Zepo. Deepfake fraud attempts increased 2,137% over three years, climbing from 0.1% to 6.5% of all fraud attempts. In 2024, those incidents jumped from roughly one per month to seven per day. Global fraud losses linked to generative AI are projected to rise from $12.3 billion in 2024 to $40 billion by 2027.

The technology is already being deployed in multiple schemes. Scammers can impersonate executives during live voice or video calls to authorize transfers, imitate relatives during supposed emergencies, or begin conversations through WhatsApp or SMS before following up with a cloned-voice call. The method is flexible and increasingly accessible.
Human perception offers little protection. A 2024 analysis covering 56 studies and 86,000 participants found people were only around 55% accurate at identifying deepfakes by ear or eye. Commercial detection tools can reach 96% accuracy in laboratory conditions, but that performance drops to between 50% and 65% in real-world conditions.
Subtle audio artifacts or unnatural pacing can sometimes surface in cloned voices, but experts warn these glitches are not reliable enough to make simply listening carefully a safe strategy. The strongest defense is independent verification, a step that matters because the technology is already capable of real-time interaction.
For consumers, the core advice is straightforward: if a loved one suddenly calls asking for money, hang up and call them back using a known number. If that is not possible, ask something specific and recent that would not be available on social media. Never let urgency remove the verification step.
The Ferrari incident illustrates why experts say people should stop relying on whether a voice sounds right when money or sensitive information is involved. A question only the real person could answer could be the difference between a convincing call and a scam being exposed.

