AI/Vishing: The Voice Is Not the Control
written by: Tiago Assumpcao
TLP Clear
Attackers are calling employees at financial firms and talking their way into identity workflows. The controls most crypto organizations rely on were never designed to survive a caller who sounds right.
In early August, a coordinated wave of attempted cyberattacks reached some of Wall Street's highest value firms. Bloomberg reported that Citadel, Point72 and Two Sigma were among the targets; Reuters reported that major financial services firms and money managers were approached through phone calls in which criminals tried to trick employees into granting access or handing over sensitive information.
Only one named firm described the outcome. Two Sigma told Bloomberg: "Our security team responded quickly to an attempted vishing campaign targeting Two Sigma and other investment managers, and we have no indication of any impact to our data or our systems." Bloomberg reported that Point72 informed investors it had been attacked, and that the firm's initial review found no client information was stolen. Point72 and Citadel declined press comment.
Bloomberg's description of the technique is worth quoting exactly, because it is not what the coverage headlines said: "Hackers can also listen into a phone call and mimic the voice, tone and phrasings of the speakers to create fake calls." No vendor research published since has confirmed synthetic audio in this specific activity — and the best documented campaign against the same sector in the same weeks involved human callers. That distinction matters, and we come back to it.
The tactic was voice phishing. The larger issue was trust.
That is where this becomes relevant for crypto, and we already know what a bad trust decision costs us. When Coinbase disclosed in May 2025that attackers had bribed support insiders to copy customer data, the stolen records included phone numbers — and Coinbase said the aim was "to gather a customer list they could contact while pretending to be Coinbase." Chainalysis estimatesimpersonation scams grew more than 1,400 percent in 2025.
The exposed functions are the ones that make identity and access decisions on a call or in chat: customer support and account recovery; IT help desk and identity administration; treasury and withdrawal approvals; OTC counterparty confirmation; executive operations and after hours exceptions; fraud, compliance and incident response.
On paper, these are workflows. In practice, many still rely on someone believing they recognize the person asking.
What an inbound call can cause
Before any control, run the diagnostic. The test is simple: what can an inbound call cause?
Can it reset MFA? Enroll a new device? Change a recovery email? Release a withdrawal hold? Expand API permissions? Reach customer information? Move an approval outside normal process?
If the answer is yes, the caller's realism is load bearing in that workflow — and it should not be.
The data says this is where attackers now spend their time. Mandiant's M-Trends 2026, built on more than 500,000 hours of frontline investigations, found voice phishing is the second most common initial infection vector at 11 percent of intrusions, while email phishing accounts for 6 percent. CrowdStrike's 2026 Threat Hunting Report found vishing intrusions in the first half of 2026 doubled against the second half of 2025.
The industry spent fifteen years hardening email. The most productive human attack vector is now the telephone, and most organizations have no control that covers it.
How a phone call becomes durable access
Mandiant reported in January that vishing intrusions by the cluster UNC6661 fed extortion actor UNC6240, which brands its demands as ShinyHunters. The same research tracked a second cluster using similar tradecraft, UNC6671, whose extortion emails are unbranded and which uses separate negotiation infrastructure.
The UNC6661 sequence is worth slowing down, because each step looks like a normal identity workflow until it is too late:
A caller posing as internal IT tells the employee that MFA settings are being updated.
The employee is directed to a victim branded login page and enters SSO credentials and MFA codes.
The attacker uses that access to register an attacker controlled MFA device.
Mandiant also reported that in at least one case, attackers used a Google Workspace add-on — ToogleBox Recall — to permanently delete the notification email alerting the user that a new MFA device had been enrolled. That step matters, because it shows the goal was not only access, but access that could persist quietly. Notably, the same activity used compromised accounts to send follow on phishing to contacts at cryptocurrency-focused companies.
That is how a phone call becomes durable access. The compromise does not need to start with a software vulnerability. Mandiant was explicit: "This activity is not the result of a security vulnerability in vendors' products or infrastructure." The weak point was the human process around identity, MFA enrollment, and SaaS access.
A voice gets the target into the workflow. The workflow gives the attacker persistence.
And it happens fast. CrowdStrike reports that the adversary SNARKY SPIDER "moved from account takeover to data theft in under five minutes." Dark Reading's coverage of the full report is more specific about a second adversary: "In four minutes flat, Cordial Spider had successfully vished a victim, authenticated to their network, and registered a new allowed MFA device." That four minute figure does not appear in CrowdStrike's own published summary, only in coverage of the full report. The same coverage records what happened next, and it is the more useful half: the new device listing raised alarms, and the victim expelled the attacker twelve minutes later.
Four minute compromise is survivable. It is survivable because something was watching the enrollment event.
The remediation became the pretext
On 6 August, the Google Threat Intelligence Group published follow up research on UNC6671 that should change how the above is read.
Three things have evolved. Callers now reach employees on personal mobile numbers, circumventing corporate telephony controls. They spoof the organization's legitimate help desk number. And the pretext is no longer a vague MFA update — GTIG describes an urgent help desk mandate to enable FIDO2 passkeys.
Read that again. Phishing resistant MFA is the single most recommended control against this exact attack, and attackers have turned its rollout into the lure. Victims are sent to lookalike sites where their own company name sits as a subdomain on a generic passkey root — the pattern [company].createssopasskey[.]com, alongside roots including addssopasskey[.]com, passkeyhelpdesk[.]com and passkeyenroll[.]com. Adversary in the middle infrastructure intercepts credentials and MFA tokens live, during the call. The attacker then enrolls its own MFA device, initiates unauthorized password resets for non-SSO enterprise applications, and systematically deletes the password reset confirmations, security modification alerts and company wide security notices that would have exposed it.
GTIG also documented which sectors were targeted when: manufacturing, real estate, healthcare and insurance across April and May; large technology, transportation and hospitality in June; and by July, financial services, private equity, law firms and ratings agencies. Root domain registration ran at roughly one every 2.2 days in April and May and one every 1.6 days across June and July, with seven stood up in the 72 hours of 20 to 22 July.
The extortion is profitable. GTIG reviewed 18 Bitcoin addresses tied to UNC6671's BlackFile extortion brand, receiving 141.65 BTC — about $10.69 million — between 7 January and 12 May 2026. Initial demands ran $1 million to $3 million with 50 to 75 percent reductions offered in negotiation; in over 53 percent of tracked cases the final payment averaged about $750,000.
Crypto ISAC assessment: we judge with low to moderate confidence that digital asset firms are a plausible next target set for this activity. They share the SaaS and identity provider topology this actor exploits, they hold immediately liquid value, and many answer to the same financial regulators. This is an inference from an observed targeting pattern, not observed activity against our sector.
One thing GTIG's reporting does not say: it makes no claim that AI generated voice was used. It describes human callers with spoofed caller ID. Reuters' follow up on 6 August, covering the same campaign against private equity firms and market infrastructure, described the hackers' use of "low-tech tactics such as phone calls," and quoted GTIG's Austin Larsen: "Sophisticated is not the right word. It is just really effective."
Two attacks are being conflated
Two different things are being called AI vishing, and they need separating.
The first is bespoke and human. A small crew, a spoofed help desk number, a personal mobile, a good script. It is expensive per target and it already pays: in over half of UNC6671's tracked cases the final payment averaged about $750,000, which funds a great many phone calls. It is running in financial services right now, and it does not need AI. A separate cluster, Scattered LAPSUS$ Hunters, is advertising $500 to $1,000 upfront per call for callers — and per Dataminr, recruiting women specifically, because a female voice falls outside the caller profile help desk staff are trained to flag. Attackers are deliberately engineering human voices, because human voices still work.
The second is volume. The same conversation, run against every employee on a leaked org chart at once. Heiding et al. find human operated vishing unprofitable at U.S. wages, while AI powered vishing "appears to be economically viable for several models."
The first attack proves recognition is not a control. The second is what happens when defeating recognition stops costing anything.
Neither is defeated by better listening. Both are defeated by the same processes.
The quiet trust layer is breaking
Security teams usually talk about vishing as social engineering, which is accurate but incomplete. The more interesting problem is that humans have always used recognition as a shortcut for trust. A boss's cadence, a colleague's tone, an IT support person's normal way of speaking. Those cues were never strong authentication, but they carried weight.
The FBI has warned directly about this, and it is the clearest government evidence that synthetic audio is already in operational use. In a May 2025 IC3 PSA, the Bureau said actors had used text messages and AI generated voice messages to impersonate senior U.S. officials "in an effort to establish rapport before gaining access to personal accounts." By July 2026, it was warning that scammers were impersonating the IC3 itself: "Scammers generate videos for real time video chats with alleged company executives, law enforcement, or other authority figures," and returning to prior fraud victims with fake recovery offers. Anyone handling recovery requests should read that twice.
There is now a measurement of how badly recognition performs. In a study accepted for publication in Expert Systems with Applications, researchers at Harvard Kennedy School and Harvard SEAS, with co-authors at Meta, tested six voice models — including Meta's own — against human baselines with 4,100 U.S. participants. Listeners correctly identified an AI voice 70.3 percent of the time. In scam scenarios, one model reached statistical parity with human voices on sentiment, persuasiveness, trustworthiness and human likeness.
Seven in ten is a majority. It is also a 29.7 percent error rate, measured in a laboratory where participants knew they were being asked to spot fakes. No control failing three times in ten, with no failure signal, would be accepted anywhere else in a security architecture. We would not accept it from a firewall.
Two findings matter more than the headline number.
Prior exposure to AI conferred no protection. Neither general AI familiarity nor voice assistant usage improved detection accuracy. Whatever awareness alone buys here, it is not detection.
Persuasiveness, not human likeness, predicted compliance. Human likeness did not independently predict whether people complied. Attackers do not need a perfect voice. They need a good script.
So hardening perception fails twice over: staff catch the fake seven times in ten, and catching it is not what determines whether they comply.
The obvious rebuttal is to detect synthetic audio automatically. Not available yet either. The organisers of the AT-ADD detection challenge report that under real world degradation — noise, reverberation, replay, compression — their strongest baseline detector reaches 76.73 percent macro-F1, and a conventional spectrogram model just 47.41 percent. Those are challenge baselines rather than telephony specific results, but synthetic audio detection is nowhere near a deployable control on a live call.
Crypto is already living the workflow half of this
Our sector does not need to imagine how a voice becomes a loss.
Support tooling. Coinbase's May 2025 disclosure describes cash offers to a small group of insiders to copy data from customer support tools — names, addresses, email and phone numbers — explicitly so attackers could contact customers while pretending to be Coinbase. Coinbase refused a $20 million extortion demand. Per the Maine Attorney General filing, 69,461 people were affected; Coinbase's Q2 and Q3 2025 shareholder letters put related costs at roughly $355 million combined. In December 2025, Brian Armstrong said a former Coinbase customer service agent had been arrested in India.
The support team itself. In April 2026 Kraken disclosed an extortion attempt built on a threat actor's recruitment of client support staff, first surfaced by a tip in February 2025, involving two instances of improper access to limited customer data. Around 2,000 client accounts, or 0.02 percent of users, were affected. CSO Nick Percoco was direct: "our systems were never breached; funds were never at risk; we will not pay these criminals; we will not ever negotiate with bad actors."
Self service permissions. In August 2025 Binance documented callers "using spoofed numbers or voice-over-IP technology" to impersonate Binance support, walking victims step by step through expanding their own API permissions. In Binance's words: "By manipulating victims into expanding API permissions, such as enabling withdrawal functions, attackers gain near-total control." This is our closest analogue to attacker MFA enrollment — and our assessment is that it is harder to catch, because the change originates from the legitimate user on their own device.
Synthetic familiarity on video.February 2026 Mandiant reporting describes UNC1069 — suspected with high confidence to have a North Korea nexus, and observed against centralized exchanges, venture capital, payments, brokerage, staking and wallet infrastructure — contacting a victim from a compromised Telegram account belonging to an executive at a crypto company, moving them to a spoofed Zoom domain, and presenting what the victim reported as a deepfake video of a CEO from another cryptocurrency company. Mandiant was careful: it "was unable to recover forensic evidence to independently verify the use of AI models in this specific instance." Kaspersky has separately documented a comparable campaign against Web3 executives, which it attributes to BlueNoroff — a threat actor Mandiant describes as overlapping with UNC1069 — where "these videos were, in fact, real recordings secretly taken from other victims." So this may be the same operator, and the faces were stolen rather than generated. Neither vendor draws a conclusion about identity controls. Ours is this: whether the face on the call was synthesized or lifted from an earlier victim, the control that failed was recognition.
The scale, and the raw material. Chainalysis estimates $17 billion was stolen in crypto scams and fraud in 2025, with impersonation scams growing more than 1,400 percent over 2024 and average payment severity up more than 600 percent. The FBI's 2025 Internet Crime Report logged $11.37 billion in cryptocurrency related losses, up 22 percent. And the input keeps being refreshed: on 13 August, Trezor disclosed a breach at its shipping and logistics provider ShipMonk exposing 13,689 customers, 11,742 of them with name, email, phone number and shipping address together. Trezor's warning was blunt — the data lets scammers "make fake phone calls" and "potentially impersonate banks, crypto exchanges, or even Trezor." In January, reporting on investigator ZachXBT's on-chain work described a single victim losing a reported $282 million in Bitcoin and Litecoin after an attacker impersonating Trezor "Value Wallet" support obtained their wallet backup. Published asset breakdowns of that total do not fully reconcile, and Trezor's own systems were not involved.
Now set the two halves side by side. UNC1069's deepfake drove malware delivery, not a transaction. The Coinbase, Trezor and Binance cases turned on human callers working customers, not synthetic voices working staff. To our knowledge, as of publication, no confirmed public case exists of AI generated voice driving a crypto firm's own help desk to reset MFA, enroll a device, or release funds.
The workflow playbook is fully mature in our sector. The AI voice capability is mature elsewhere and already productized: Abnormal AI documented in April 2026 a service called ATHR selling for $4,000 plus 10 percent of profits, running an automated ten section call script, shipping with prebuilt credential panels for Coinbase, Binance, Gemini and Crypto.com alongside Google, Microsoft, Yahoo and AOL. Half its target brands are digital asset platforms.
Those two halves have not been confirmed converging on a crypto help desk. That is a gap to prepare inside, not a reason for comfort.
Vibe based verification is dead
For years, organizations quietly relied on "this feels right" as part of verification. A familiar voice, a known name, a convincing backstory, and the right tone could move a request forward.
That era is ending, because vibe can now be manufactured. A realistic voice can be generated. Caller ID can be spoofed — including the help desk's own number. A lookalike site can be built in minutes. Internal language can be gathered from past breaches, LinkedIn, Slack screenshots, support scripts, or prior interactions. And as UNC6671 shows, even a security roadmap becomes usable pretext.
Regulators have started saying this out loud. The New York Department of Financial Services issued an industry letter in February 2026 on an ongoing targeted vishing campaign, advising DFS regulated entities — a population that includes licensed virtual currency businesses under 23 NYCRR Part 500 — to verify identity by means beyond caller ID, review MFA enrollment controls and help desk access permissions, and monitor for anomalous authentication activity. The Financial Services Sector Coordinating Council's AI and Identity/Authentication Workstream, supported by the American Bankers Association and the Better Identity Coalition, published mitigations guidance in 2026 whose second named tactic is, verbatim, "Social Engineering Using Deepfake Voice Against a Financial Services Call Center or Help Desk."
What crypto teams can do now
Use known good callbacks for sensitive requests. Help desk, treasury, OTC, support and executive workflows should verify high risk requests through numbers registered in advance or previously confirmed channels — never the number provided during the inbound call. Because UNC6671 spoofs the legitimate help desk number, "the number looked right" is no longer evidence of anything.
Move privileged users to phishing resistant MFA, and defend the migration itself. FIDO2 security keys and passkeys break adversary in the middle credential replay, because the authenticator signs a challenge bound to the legitimate origin and will not respond to a lookalike domain. They do not protect the enrollment path, the account recovery path, or any phishable fallback factor left enabled — which is exactly where UNC6671 operates. Announce enrollment campaigns through verified internal channels first, tell staff explicitly that IT will never call to walk them through passkey setup, and expect attackers to time their calls to your rollout.
Ban credential or MFA code entry prompted by inbound calls. Employees should never enter credentials, MFA codes, or recovery information because someone called and instructed them to. Internal IT calls should be treated as untrusted until verified through a known process.
Harden the help desk as an identity system, not a service desk. Mandiant's frontline guidance for this activity recommends out of band approval from a known manager before password resets, live video verification with physical government ID held next to the face for high risk resets, and — during an active campaign — temporarily disabling self service password reset and pausing MFA and device registration. Consider issuing temporary access codes rather than performing resets.
Instrument identity changes, not just logins. See the indicators below. Alert on MFA enrollment preceded by authentication failures, removal of an existing MFA factor, new device registration, recovery channel changes, deletion of security notification emails, and unusual SaaS access following any support interaction.
Verify with the unrehearsable. When Kraken identified a DPRK operative in its own hiring pipeline, the tell was live: the candidate "occasionally switched between voices, indicating that they were being coached through the interview in real time," and then "couldn't convincingly answer real-time questions about their city of residence or country of citizenship." Shared verification phrases work on the same principle — they do not care how real the caller sounds. Rotate them when roles, exposure or personnel change.
Build vishing into incident response playbooks, and rehearse the exception path. Cover vished MFA resets, attacker enrolled devices, help desk and support impersonation, executive impersonation, and urgent transaction pressure. Response steps should include identity provider review, SaaS session revocation, support ticket review, and follow on fraud checks. Tabletop the moments attackers actually exploit: someone locked out, someone traveling, someone senior appearing to ask for help before market open.
Indicators and observable patterns
From GTIG's UNC6671 reporting and Mandiant's January research. Defanged; validate against your own environment before deploying.
Phishing infrastructure. Victim company name as a subdomain of a generic authentication root — [company].createssopasskey[.]com, [company].addssopasskey[.]com — plus roots including passkeyhelpdesk[.]com and passkeyenroll[.]com. GTIG's indicator table lists more than 70 such domains, created between 4 April and 3 August 2026. Watch for newly registered domains pairing your company name with passkey, sso, mfa, okta, access, support or internal.
Identity provider. GTIG recommends querying Okta and Microsoft Entra ID audit logs for MFA registration events (system.multifactor.factor.setup) immediately preceded by authentication failures (user.authentication.auth_via_mfa) or abandoned push challenges. Also alert on removal of an existing factor immediately before enrollment of a new one, and on multiple MFA enrollments from a single source IP across different users.
Automated exfiltration. GTIG observed scripting-library user agents including python-requests and WindowsPowerShell on high volume document access and download in Microsoft 365 and SharePoint. Scripted user agents on document access are not people browsing.
Notification suppression. Deletion of password reset confirmations, security modification alerts and company wide security notices from a mailbox shortly after a support interaction. In the January activity this was achieved with the Google Workspace add-on ToogleBox Recall — review and restrict OAuth app authorizations accordingly.
Closing
Crypto teams do not need to win a listening contest against synthetic voices. A three in ten error rate is not a control, and detection tooling will not rescue it over a phone line.
What is winnable is process.
UNC6671 is calling personal phones with your help desk's number on the screen, telling people it is time to enroll a passkey.
The voice is no longer the control. The process has to be.
What we are watching
Whether the targeting pattern observed against financial services extends to digital asset firms, and whether passkey migration pretexts appear against crypto help desks.
The first confirmed case of synthetic voice driving a crypto identity workflow — MFA reset, device enrollment, or withdrawal release.
Whether AI vishing services move from credential panels aimed at retail exchange customers to pretexts aimed at exchange staff.
Sources: Reuters (5 and 6 Aug 2026); Bloomberg (5 Aug 2026) — Bloomberg and Reuters both block automated retrieval, so the Two Sigma statement and the voice-mimicry sentence were verified verbatim against three independent renderings of the Bloomberg wire — Fortune (6 Aug 2026), Claims Journal (6 Aug 2026) and InvestmentNews (6 Aug 2026); Google Threat Intelligence Group / Mandiant — UNC6671 (6 Aug 2026), UNC1069 (9 Feb 2026), ShinyHunters SaaS data theft and accompanying frontline guidance (30 Jan 2026), M-Trends 2026 (23 Mar 2026); CrowdStrike 2026 Threat Hunting Report (3 Aug 2026) and reporting thereon; FBI IC3 PSA I-051525 (15 May 2025), PSA I-072026-PSA (20 Jul 2026), and 2025 Internet Crime Report; Heiding et al., Evaluating AI Models' Capability to Automate Voice Phishing Attacks (arXiv:2607.09970, accepted, Expert Systems with Applications); Xie et al., AT-ADD Challenge Evaluation Plan (arXiv:2604.08184, Apr 2026); Chainalysis 2026 Crypto Crime Report (13 Jan 2026); Coinbase (15 May 2025), Maine AG breach notification, Coinbase Q2/Q3 2025 shareholder letters, Armstrong statement (Dec 2025); Kraken (1 May 2025; Apr 2026 via BleepingComputer); Binance (22 Aug 2025); Trezor via BleepingComputer (13 Aug 2026); Kaspersky Securelist (28 Oct 2025); NYDFS industry letter (6 Feb 2026); FSSCC AI and Identity/Authentication Workstream (2026); Abnormal AI (16 Apr 2026); Dataminr via The Hacker News (25 Feb 2026); ZachXBT via Brave New Coin (10 Jan 2026).
Tiago Assumpcao is the Technical Director of Threat Intelligence Engineering at Crypto ISAC. With over 25 years in cybersecurity, his career demonstrates a clear trajectory: from pioneering exploit defenses in the early 2000s to shaping the security posture of today’s crypto ecosystem. He has contributed to foundational memory protection mechanisms, served as an editor at Phrack, and led security initiatives at BlackBerry, IOActive, and Coinspect. His work spans from pioneering exploit defenses to securing digital asset ecosystems against systemic risks and nation-state threats. He recently participated as a panelist at Upbit D Conference and DeFi Security Summit.
The Crypto ISAC is a member-driven, not-for-profit organization that works together to curb malicious actors, address vulnerabilities, share intelligence, and move security forward to protect the crypto ecosystem. We are founded by leading crypto organizations and designed for cryptosecurity experts to address the security and trust challenges that face crypto today and shape the crypto ecosystem of tomorrow.