
Ten Minutes, $600,000: How a Face and a Voice Hijacked Trust Itself
TL;DR
Face- and voice-swap fraud didn't break a technology — it broke a trust contract never designed to withstand synthetic attacks. Labeling rules are catching up, but scammers were never aiming at the label. They're aiming at your verification habits.
At 11:40 a.m., WeChat Rang
At 11:40 a.m. on April 20, 2023, a businessman surnamed Guo — the legal representative of a tech company in Fuzhou — was going about his day when WeChat buzzed with a video call request. It was from a “friend.”
Guo picked up. On the screen was a familiar face, speaking in a familiar voice. The caller said a friend of his was bidding on a project out of town and urgently needed 4.3 million yuan (roughly $600,000) in deposit money. Could Guo wire it through his company’s account? The funds would be transferred back shortly.
This wasn’t the crude, faceless request of a typical scam call. Guo recognized the person on the video. He recognized the voice. Following the bank account number the caller provided, he wired the money in two transfers. The caller even sent, as a courtesy, a screenshot of a bank transfer receipt to prove the money was already on its way from his end.
By 11:49 a.m., the transfer was done. Less than ten minutes had passed.
Afterward, Guo casually messaged his “friend” on WeChat to confirm. The reply that came back was a single question mark: “?”
That was the moment he understood: the face and the voice on that video call had both been fake (Xinhua, 2023; Fuzhou News Network, 2023). Police in Fuzhou and Baotou, working through the night with the banks, managed to freeze and recover 3.3684 million yuan; what happened to the remaining roughly 940,000 yuan is not disclosed in public records (Xinhua, 2023).
The case was later flattened into a sensational headline that spread across the Chinese internet: “$600,000 stolen in ten minutes.” But treating it as just another “AI gone wild” curiosity misses what actually makes it worth remembering. This wasn’t an AI malfunction — it was a direct hit on the one verification habit almost nobody questioned: that seeing a face and hearing a voice on a video call meant it was really them. This case was the first time the Chinese public, at scale, watched that proof get punched through.

From Deceiving One Person to Deceiving an Entire Boardroom
If Guo’s ten minutes were just an isolated stroke of bad luck, this story would end here. But placed on the timeline of the past three years, it looks less like an endpoint and more like an opening move.
In January 2024, an employee at the Hong Kong branch of a multinational engineering firm received a phishing email purporting to be from the company’s UK-based chief financial officer, then was pulled into what looked like a multi-person video conference. Hong Kong’s government later confirmed, in a written reply to the Legislative Council, what that “meeting” actually was: a pre-recorded video fabricated from downloaded public footage and the impersonated officer’s voice. Everyone on screen besides the victim was fake, and there was no real interaction between him and the fraudsters — after the fake “CFO” issued the transfer instructions, the meeting was cut short on some pretext, and the rest of the instructions came through instant messaging (Hong Kong Government press release, 2024). Over the following week, the employee wired money in 15 separate transactions totaling roughly HK$200 million — more than $25 million (Hong Kong Government press release, 2024; CNN, 2024; The Paper, 2024). According to 21jingji.com, the material used to synthesize the executives’ faces and voices included the company’s own YouTube videos — though that specific detail comes from media reporting, not from the government’s official statement (21jingji.com, 2024).
This wasn’t a coincidental escalation — it was the same attack technique climbing the trust chain: from deceiving one person to deceiving a whole room of people simultaneously “seeing” each other. The numbers tell the slope of that curve more precisely than any single case can.
China’s Ministry of Public Security reported in 2023 that a targeted crackdown had cracked 79 “AI face-swap” fraud cases and arrested 515 suspects (Xinhua, 2023). Chen Hongxiang, head of the Third Criminal Tribunal at the Supreme People’s Court, said in an interview that fraud using AI “face-swapping” and “voice-changing” is “more deceptive,” that “related cases are increasingly common,” and that courts would “crack down severely” (21st Century Business Herald, 2024). By early 2026, in a separate interview, he went further: malicious abuse of AI face-swap and voice-cloning technology in telecom and internet fraud has become “more concealed and more deceptive,” and he called for accelerated research and timely regulatory rules (Sina Finance, 2026) — meaning that three years on, regulators are still chasing the pace of the technology. The problem hasn’t been solved, only described more precisely.
This means the people at risk aren’t just “people who understand AI” or “elderly people prone to being fooled.” A 2025 Beijing Youth Daily headline put it bluntly: “AI fraud victims are not just the elderly.” Victims have included corporate legal representatives, company employees, and university students — anyone who answers video calls and trusts a familiar voice falls within reach of this kind of scam (The Beijing News, 2025).
When “Seeing” and “Hearing” Used to Be Enough
Before diving into how the technology works, it’s worth correcting a detail that easily gets distorted, and returning to what the victims themselves actually went through — what they did wrong, or whether they did anything wrong at all.
The most widely circulated version of the Guo case claims that “AI face-swapping only needs one photo.” That isn’t technically accurate. A 2024 Xinhua anti-fraud report explicitly notes that real-time video face-swapping is not a matter of “one photo and you’re done” — it requires training a model on video and audio material shot from multiple angles, in significant volume (Xinhua, 2024). In other words, the fact that scammers could convincingly impersonate Guo’s “friend” implies that this friend — whether he realized it or not — had left behind enough public footage: WeChat Moments videos, livestream recordings, short-video appearances, any of which could become raw material.
This point is even starker in the Hong Kong CFO case: the scammers used the company’s own official YouTube videos as source material (21jingji.com, 2024). In other words, the victims’ exposure came precisely from something they’d never thought to guard against — the fact that publicly posted footage of their own faces is itself a digital asset that can be weaponized.
Guo’s own response wasn’t actually careless. He didn’t wire money on the strength of a phone call alone — he insisted on video verification, which used to be a reasonable, even cautious, precaution. What broke through his defenses wasn’t a lapse in vigilance; it was the fact that the verification method itself had already stopped working, something he — like nearly everyone else at the time — simply didn’t know. Related analysis from institutions including China University of Political Science and Law points to a shared feature across these cases: victims almost always performed some version of a “quick look, quick listen” confirmation, and it was exactly that threshold — once considered reliable enough — that the technology bypassed (see the Zhihu series “Shuke Fa Ping,” 2024).
The Hong Kong employee’s experience follows the same pattern. He didn’t skip the process — he “sat in” on what looked like a formal multi-person video conference, with several “colleagues” and “superiors” appearing to take turns laying out the instructions. But per Hong Kong’s official account, that meeting had been recorded in advance; nobody on screen was real, and he never had a genuine exchange with the other side. In the old logic, more people in the room, taking turns speaking, should have meant a stronger defense against fabrication. In this case, even that instinct — “more people present means more credible” — was turned against him by a video that couldn’t be questioned or interrupted, because it wasn’t happening live at all.
Cracking Open the Master Key
The two cases above show that what broke through the victims’ defenses wasn’t a lack of caution. To understand exactly how this line of defense gets bypassed, we need to unpack two terms first.
Deepfake refers to the use of deep learning models to replace or overlay one person’s face, expressions, and mouth movements onto another piece of footage, making it look as though that person is speaking or acting in real time. AI voice cloning refers to training a model on a sample of a target’s voice, then synthesizing new speech that mimics their tone and cadence — feeding it arbitrary text, or driving it live, so that “this voice” says whatever the scammer wants it to say.
A 2024 People’s Daily report summarized this category of fraud into two techniques — “AI voice cloning” and “AI face-swapping” — the former simulating a familiar or authoritative voice to earn trust, the latter fabricating a video likeness to convince the victim they’re talking to a real person. The two are frequently combined to complete a single chain: build trust, apply pressure, extract a transfer (People’s Daily, 2024).
The operational process breaks down into roughly three steps.
Step one: gathering material. Scammers collect enough video and audio samples of the target through social media, public speeches, livestream recordings, or even a preliminary “test call.”
Step two: synthesis and real-time driving. A face-swap model replaces the scammer’s own face with the target’s in real time, or a voice-cloning model makes the target’s “voice” say whatever text the scammer inputs, live. This step requires no understanding, on the model’s part, of the relationship between the victim and the person being impersonated — the model’s only job is to look and sound convincing; it isn’t in the business of the con itself.
Step three: embedding the fabrication into social engineering. The fabricated identity gets inserted into a pre-existing trust scenario — a friend needing to route funds through your account, a boss ordering an urgent transfer, a CFO assigning a task — exploiting people’s habitual deference to existing relationships and authority, not any “intelligence” on the part of the AI.
This is also why reducing these cases to “AI turned evil” is a misreading. The AI model here functions like a well-made master key; what actually determines which door gets opened and what gets stolen is the person holding the key and the script of deception they’ve designed. Xinhua’s anti-fraud guidance also offers a technical correction: real-time face-swapping has real requirements around the volume of source material, network conditions, and how distinctive the target’s features are — not every photo can instantly generate a convincing real-time video. It also notes that video calls may still carry tells — abnormal blink rates, wandering eye focus, mismatched expressions and lip movement, disjointed speech logic — though these tells are becoming harder to catch with the naked eye as the technology advances (Xinhua, 2024; People’s Daily, 2024).
What Actually Broke Is a Trust Contract Never Built to Withstand Synthetic Attacks
Having unpacked the technical process, one more fundamental question remains: why does this “master key” happen to fit the lock called “seeing a face, hearing a voice”? Here is a judgment, not just a restatement of fact: the root cause of these fraud cases isn’t “AI running out of control.” It’s that a verification method society has relied on for decades — recognizing a voice, recognizing a face — was never designed to withstand synthetic attacks. It worked in the past only because forging it was expensive enough to be impractical.
Before deepfake technology matured, impersonating someone’s voice and likeness required specialized equipment, significant time, and a fairly high technical bar — out of reach for the ordinary scammer. That made “confirm identity via video call” an extremely cost-effective verification method, one that filtered out most scams at essentially zero cost. Deepfakes crashed that cost of forgery. You no longer need specialized equipment or a team — off-the-shelf, open-source tools can get you started. The amount and type of source material needed, and how convincing the result is, varies a lot by tool and by how high a bar you’re trying to clear — there’s no single number of samples that applies across the board. What can be said with confidence is that the cost and skill required have fallen materially; they haven’t disappeared, but they’ve fallen far enough that the math has flipped. Once forging something costs less than verifying it, the old verification method stops being “reliable” and slides toward becoming a vulnerability — while most people’s and institutions’ security habits are still stuck in the old world.
This is also why framing this story around “how scary AI is” misses the point. The question that actually deserves scrutiny is: why, even today, do so many corporate and personal transfer-confirmation processes still rely on a single communication channel that can be synthesized? Guo’s method of verification was “glance at a video.” The Hong Kong employee’s was “sit in on a meeting.” Both, fundamentally, staked identity verification on one inherently fragile information channel, with no independent second channel to cross-check.
The Rules Have Arrived — But They’re No Silver Bullet
Once the root problem is laid bare, what has regulation actually offered in response? Three years after the case broke, the response has moved from “warning the public” to “institutional mandate.”
China’s Provisions on the Administration of Deep Synthesis of Internet Information Services, which took effect in 2023, was the first regulation to require deep-synthesis service providers to build management systems covering user registration, algorithm mechanism review, content publication review, data security, personal information protection, and anti-telecom-fraud measures — and to require conspicuous labeling for synthesized faces and voices that could cause public confusion or mistaken identity (Cyberspace Administration of China, 2022).
By March 2025, the Cyberspace Administration of China, together with the Ministry of Industry and Information Technology, the Ministry of Public Security, and the National Radio and Television Administration, jointly issued the Measures for Labeling AI-Generated and Synthesized Content, paired with the mandatory national standard GB 45438-2025, Cybersecurity Technology — Labeling Methods for AI-Generated and Synthesized Content. Both took effect together on September 1, 2025. The rules require that all AI-generated text, images, audio, video, and virtual scenes carry both an explicit label (a visible badge or watermark users can see directly) and an implicit label (a marker embedded in file metadata, imperceptible but traceable). That same day, several major Chinese social platforms rolled out “AI-generated” badges, and the rules explicitly prohibit anyone from deleting, altering, forging, or concealing such labels (Xinhua, 2025; Cyberspace Administration of China, 2025).
But the labeling system is no silver bullet, and that needs to be said plainly rather than glossed over. First, the Labeling Measures themselves leave an opening — under certain conditions of user responsibility and log retention, some scenarios are exempted from explicit labeling, leaving room for the rules to be gamed. Second, labels can be re-encoded across platforms, maliciously stripped, or outright forged, and whether enforcement’s technical tracing capability can keep pace remains an open question. Third, and most critically: labeling targets the content itself, while scammers target the act of verification. A fabricated video tagged “AI-generated” is meaningless if the victim never thinks to check for the label in the first place. The system is moving forward, but it solves the problem of “can synthetic content be identified” — not the problem of “will people habitually rely on a single channel to verify identity.”
Chen Hongxiang, head of the Third Criminal Tribunal at the Supreme People’s Court, acknowledged in an early-2026 interview that the malicious abuse of AI face-swap and voice-cloning technology in telecom fraud has become “more concealed and more deceptive,” and that judicial authorities need to “accelerate research into the new situations and problems brought by deepfake technology” (Sina Finance, 2026). Translated, that sentence means: three years on, this cat-and-mouse game has no clear winner yet. The rules are catching up, but the technology keeps running.
Stop Treating “Seeing” and “Hearing” as the Finish Line
The system is still catching up — but individuals and companies don’t have to wait for it. Compressed into one takeaway: stop treating “I saw his face” or “I heard his voice” as the end point of identity verification.
At the practical level, here are a few recommendations drawn from Xinhua’s and People’s Daily’s anti-fraud guidance, along with the lessons exposed by these cases themselves (Xinhua, 2024; People’s Daily, 2024):
- For any request involving a large transfer that carries urgency (“right now,” “someone’s waiting”) — whether it arrives by video, voice, or text — hang up first, then re-verify through a separate, independent channel: call the person’s long-saved phone number, or confirm in person.
- Don’t treat “the face on this video” or “the voice on this call” as sufficient verification on its own, especially for instructions involving money.
- Internal corporate transfer-approval processes should require confirmation from at least two independent information sources, not deference to a single communication channel — even when that channel is “a video conference I watched with my own eyes.”
- Technical tells are still worth watching for, but shouldn’t be relied on: abnormal blink rates, slightly mismatched expressions and lip movement, jumps in speech logic — these are useful as supporting signals, but they’ll keep getting harder to catch by eye as the technology improves, and they can’t be the sole basis for trust.
- On a personal level, reducing unnecessary exposure of public video and audio material — especially long, continuous conversations with a clear frontal view — objectively raises the bar for gathering source material. It’s not a complete solution, just a higher barrier.
What began with a ten-minute video call in Fuzhou ultimately points not to “how smart AI has become,” but to a plainer reminder: as forging a face or a voice keeps getting cheaper and easier, any trust built solely on “seeing” and “hearing” needs to be redesigned.
Sources
- AI换脸新骗局!福州一老板10分钟被骗走430万元 - 福州新闻网
- 法制日报 相关报道
- 公安机关侦破“AI换脸”相关案件79起 抓获犯罪嫌疑人515名 - 新华网
- 新华网防骗提示报道(2024年2月)
- 人民日报相关报道(2024年4月)
- LCQ9: Combating frauds involving deepfake - Hong Kong Government press release, June 26, 2024
- 香港CFO深伪诈骗案 - CNN
- 香港警方:跨国公司被“AI换脸”成CFO的骗子欺骗,损失高达2亿港币 - 澎湃新闻
- 震惊!“变脸”冒充CFO,骗走两个亿!香港最大AI诈骗案细节曝光 - 21经济网
- 10分钟被骗430万,被AI诈骗的不止老年人 - 新京报
- AI换脸织就的诈骗迷网 - 最高人民检察院
- 专访最高法刑三庭庭长陈鸿翔:从严打击利用AI“换脸”“变声”的新型网络诈骗(2024年3月) - 中国法院网
- 最高法:恶意滥用AI换脸、拟声技术 电诈手法更隐蔽更具迷惑性(2026) - 新浪财经
- 《互联网信息服务深度合成管理规定》 - 国家网信办
- 关于印发《人工智能生成合成内容标识办法》的通知 - 国家网信办
- 9月1日起 AI生成合成内容必须添加标识 - 新华网