For years, the tell was the voice. A cadence that landed wrong. A surname pronounced by someone who had clearly only ever read it. Nobody trained staff to catch that. They just did.
That signal is gone.
Cloning a voice from public audio is a commodity task now. An earnings call, a conference panel recording, a podcast appearance, and an attacker has your CFO.
Pair that with a spoofed number and the call passes every check a human being can run while the line is still open.
For helpdesk and finance teams, this doesn’t weaken the informal verification layer. It removes it. Work those teams used to do on instinct now has to be done by procedure, or it doesn’t get done at all.
Worth being precise about what caller ID actually is, because the misunderstanding is the thing that makes the attack land.
On a SIP call, the calling number sits in the From header, and often in P-Asserted-Identity as well. Those fields are populated by the originating party.
A carrier that accepts traffic from a customer without tight validation passes along whatever it is handed. The display name is worse.
In most paths it is free text with no verification at all, and in the rest it is resolved from a database keyed on the same unverified number.
So treat the number on the screen as attacker-controlled input. It is a claim, not an assertion. STIR/SHAKEN exists to close this gap and it does real work, but I will come back to what it cannot tell you.
The attacker calls the service desk as an employee who is traveling. Phone is dead, laptop is locked out, client meeting starts in ten minutes.
The pressure is calibrated carefully: urgent enough that the agent skips a step, mundane enough that nothing feels wrong. The voice matches the employee.
The number matches the employee’s mobile in the directory. The agent resets the credential or enrolls a new MFA device, and the account is gone.
This one is rarely a cold call. The attacker has usually been sitting in a mailbox first and knows there is a live invoice and a payment run date.
The call arrives from the supplier’s number, in the voice of whoever accounts payable normally deals with, carrying a change of bank details that needs to land before Friday.
Written confirmation follows from a lookalike domain. Two channels appear to corroborate each other. Both of them are the attacker.
The Arup case from January 2024 is the one everyone cites, and it earns the reference. An employee in the firm’s Hong Kong office joined what looked like a routine multi-participant video call with the UK-based CFO and colleagues he recognized.
He had already been suspicious of the initial written message. The call is what removed the suspicion.
Fifteen transfers followed, totaling roughly HK$200 million. Hong Kong police said the fakes were assembled from footage that was already public.
Notice the shape of it: doubt raised in writing, doubt resolved by a synthetic face and voice. That is the entire attack in one sentence.
Aimed at people with signing authority, and increasingly at their family members. The call presents from the bank’s published fraud line.
There is a suspicious transfer, and the caller will walk the target through securing the account, which means moving money into a safe account that is not.
What used to give this away was a rushed script read by someone who did not understand the product.
Now it is fluent, unhurried, and gets the terminology right every time.
Almost every policy document says some version of it. Confirm the caller’s identity before disclosing information or making changes. Confirm it with what, exactly?
A person on a live call has two signals available: the number and the voice. Both are now forgeable at low cost. Telling staff to verify using the two artifacts the attacker controls is not a control. It is an instruction to guess, plus a paper trail for blaming the agent when the guess is wrong.
Anything that keeps verification inside the call has already failed. The only thing that works is leaving it.
Out-of-band callback carries the weight here. The agent ends the call and dials back on a number pulled from the internal directory or the supplier contract record. Never a number read out during the call.
Never the one in the signature block of the email that arrived alongside it. If the caller pushes back on the callback, that pushback is the finding.
Shared secrets set at onboarding cover the gap when a callback is too slow. A phrase agreed in person, stored against the identity record, never sent by email.
Crude, cheap, and completely indifferent to how good the voice clone is.
A first-pass check against community-reported number data such as WhoseNo will catch recycled ranges, though it will not catch a freshly spoofed number, which is why it sits before callback verification rather than replacing it.
Then the part most organizations skip. Staff need written permission to hang up on an executive. If an agent expects to be disciplined for delaying a payment the CFO demanded on the phone, the agent will not delay it.
Put the no-penalty rule in policy, say it out loud in a meeting, and have someone senior repeat it. Without that, the technical controls are decoration.
After an incident, the useful fields turn out to be the ones nobody retains by default. At minimum: Call-ID, the full From and P-Asserted-Identity headers rather than a normalized number, any Diversion or History-Info headers, the ingress trunk and source IP, the Identity header or verstat parameter if it survived to your edge, negotiated codec, and UTC timestamps on both legs.
On attestation, do not read more into it than it carries. Full attestation means the originating provider vouched for its customer’s right to use that number. It says nothing about who is actually speaking.
And on international inbound the signature is frequently absent or downgraded at the gateway, so an entire category of high-risk calls reaches you carrying no meaningful grade at all. Log it, correlate on it, do not gate on it.
The only control that survives contact with a cloned voice is the one that does not depend on the call. Hang up. Dial a number you already had. Everything after that is detail.
Hackers are actively probing AI systems, turning exposed gateways and agent tools into routes for…
Hackers are making some phishing pages harder to track by changing the code delivered to…
A cyber incident reportedly forced a British power plant to halt operations for about four…
Russian hackers have used a new backdoor called HOOKEDGE to target defense manufacturers, government bodies,…
TITAN ransomware is pairing file encryption with an ambitious claim: artificial intelligence that can sort…
A fake student resume is being used to place a remote-access tool on researchers’ Windows…