Skip to content

A

Ten Minutes, $600,000: How a Face and a Voice Hijacked Trust Itself

After the Headlines

#Technology#AI Safety#Judgment

TL;DR

Face- and voice-swap fraud didn't break a technology — it broke a trust contract never designed to withstand synthetic attacks. Labeling rules are catching up, but scammers were never aiming at the label. They're aiming at your verification habits.

At 11:40 a.m., WeChat Rang

At 11:40 a.m. on April 20, 2023, a businessman surnamed Guo — the legal representative of a tech company in Fuzhou — was going about his day when WeChat buzzed with a video call request. It was from a “friend.”

Guo picked up. On the screen was a familiar face, speaking in a familiar voice. The caller said a friend of his was bidding on a project out of town and urgently needed 4.3 million yuan (roughly $600,000) in deposit money. Could Guo wire it through his company’s account? The funds would be transferred back shortly.

This wasn’t the crude, faceless request of a typical scam call. Guo recognized the person on the video. He recognized the voice. Following the bank account number the caller provided, he wired the money in two transfers. The caller even sent, as a courtesy, a screenshot of a bank transfer receipt to prove the money was already on its way from his end.

By 11:49 a.m., the transfer was done. Less than ten minutes had passed.

Afterward, Guo casually messaged his “friend” on WeChat to confirm. The reply that came back was a single question mark: “?”

That was the moment he understood: the face and the voice on that video call had both been fake (Xinhua, 2023; Fuzhou News Network, 2023). Police in Fuzhou and Baotou, working through the night with the banks, managed to freeze and recover 3.3684 million yuan; what happened to the remaining roughly 940,000 yuan is not disclosed in public records (Xinhua, 2023).

The case was later flattened into a sensational headline that spread across the Chinese internet: “$600,000 stolen in ten minutes.” But treating it as just another “AI gone wild” curiosity misses what actually makes it worth remembering. This wasn’t an AI malfunction — it was a direct hit on the one verification habit almost nobody questioned: that seeing a face and hearing a voice on a video call meant it was really them. This case was the first time the Chinese public, at scale, watched that proof get punched through.

Two screens showing a face breaking apart

From Deceiving One Person to Deceiving an Entire Boardroom

If Guo’s ten minutes were just an isolated stroke of bad luck, this story would end here. But placed on the timeline of the past three years, it looks less like an endpoint and more like an opening move.

In January 2024, an employee at the Hong Kong branch of a multinational engineering firm received a phishing email purporting to be from the company’s UK-based chief financial officer, then was pulled into what looked like a multi-person video conference. Hong Kong’s government later confirmed, in a written reply to the Legislative Council, what that “meeting” actually was: a pre-recorded video fabricated from downloaded public footage and the impersonated officer’s voice. Everyone on screen besides the victim was fake, and there was no real interaction between him and the fraudsters — after the fake “CFO” issued the transfer instructions, the meeting was cut short on some pretext, and the rest of the instructions came through instant messaging (Hong Kong Government press release, 2024). Over the following week, the employee wired money in 15 separate transactions totaling roughly HK$200 million — more than $25 million (Hong Kong Government press release, 2024; CNN, 2024; The Paper, 2024). According to 21jingji.com, the material used to synthesize the executives’ faces and voices included the company’s own YouTube videos — though that specific detail comes from media reporting, not from the government’s official statement (21jingji.com, 2024).

This wasn’t a coincidental escalation — it was the same attack technique climbing the trust chain: from deceiving one person to deceiving a whole room of people simultaneously “seeing” each other. The numbers tell the slope of that curve more precisely than any single case can.

China’s Ministry of Public Security reported in 2023 that a targeted crackdown had cracked 79 “AI face-swap” fraud cases and arrested 515 suspects (Xinhua, 2023). Chen Hongxiang, head of the Third Criminal Tribunal at the Supreme People’s Court, said in an interview that fraud using AI “face-swapping” and “voice-changing” is “more deceptive,” that “related cases are increasingly common,” and that courts would “crack down severely” (21st Century Business Herald, 2024). By early 2026, in a separate interview, he went further: malicious abuse of AI face-swap and voice-cloning technology in telecom and internet fraud has become “more concealed and more deceptive,” and he called for accelerated research and timely regulatory rules (Sina Finance, 2026) — meaning that three years on, regulators are still chasing the pace of the technology. The problem hasn’t been solved, only described more precisely.

This means the people at risk aren’t just “people who understand AI” or “elderly people prone to being fooled.” A 2025 Beijing Youth Daily headline put it bluntly: “AI fraud victims are not just the elderly.” Victims have included corporate legal representatives, company employees, and university students — anyone who answers video calls and trusts a familiar voice falls within reach of this kind of scam (The Beijing News, 2025).

When “Seeing” and “Hearing” Used to Be Enough

Before diving into how the technology works, it’s worth correcting a detail that easily gets distorted, and returning to what the victims themselves actually went through — what they did wrong, or whether they did anything wrong at all.

The most widely circulated version of the Guo case claims that “AI face-swapping only needs one photo.” That isn’t technically accurate. A 2024 Xinhua anti-fraud report explicitly notes that real-time video face-swapping is not a matter of “one photo and you’re done” — it requires training a model on video and audio material shot from multiple angles, in significant volume (Xinhua, 2024). In other words, the fact that scammers could convincingly impersonate Guo’s “friend” implies that this friend — whether he realized it or not — had left behind enough public footage: WeChat Moments videos, livestream recordings, short-video appearances, any of which could become raw material.

This point is even starker in the Hong Kong CFO case: the scammers used the company’s own official YouTube videos as source material (21jingji.com, 2024). In other words, the victims’ exposure came precisely from something they’d never thought to guard against — the fact that publicly posted footage of their own faces is itself a digital asset that can be weaponized.

Guo’s own response wasn’t actually careless. He didn’t wire money on the strength of a phone call alone — he insisted on video verification, which used to be a reasonable, even cautious, precaution. What broke through his defenses wasn’t a lapse in vigilance; it was the fact that the verification method itself had already stopped working, something he — like nearly everyone else at the time — simply didn’t know. Related analysis from institutions including China University of Political Science and Law points to a shared feature across these cases: victims almost always performed some version of a “quick look, quick listen” confirmation, and it was exactly that threshold — once considered reliable enough — that the technology bypassed (see the Zhihu series “Shuke Fa Ping,” 2024).

The Hong Kong employee’s experience follows the same pattern. He didn’t skip the process — he “sat in” on what looked like a formal multi-person video conference, with several “colleagues” and “superiors” appearing to take turns laying out the instructions. But per Hong Kong’s official account, that meeting had been recorded in advance; nobody on screen was real, and he never had a genuine exchange with the other side. In the old logic, more people in the room, taking turns speaking, should have meant a stronger defense against fabrication. In this case, even that instinct — “more people present means more credible” — was turned against him by a video that couldn’t be questioned or interrupted, because it wasn’t happening live at all.

Cracking Open the Master Key

The two cases above show that what broke through the victims’ defenses wasn’t a lack of caution. To understand exactly how this line of defense gets bypassed, we need to unpack two terms first.

Deepfake refers to the use of deep learning models to replace or overlay one person’s face, expressions, and mouth movements onto another piece of footage, making it look as though that person is speaking or acting in real time. AI voice cloning refers to training a model on a sample of a target’s voice, then synthesizing new speech that mimics their tone and cadence — feeding it arbitrary text, or driving it live, so that “this voice” says whatever the scammer wants it to say.

A 2024 People’s Daily report summarized this category of fraud into two techniques — “AI voice cloning” and “AI face-swapping” — the former simulating a familiar or authoritative voice to earn trust, the latter fabricating a video likeness to convince the victim they’re talking to a real person. The two are frequently combined to complete a single chain: build trust, apply pressure, extract a transfer (People’s Daily, 2024).

The operational process breaks down into roughly three steps.

Step one: gathering material. Scammers collect enough video and audio samples of the target through social media, public speeches, livestream recordings, or even a preliminary “test call.”

Step two: synthesis and real-time driving. A face-swap model replaces the scammer’s own face with the target’s in real time, or a voice-cloning model makes the target’s “voice” say whatever text the scammer inputs, live. This step requires no understanding, on the model’s part, of the relationship between the victim and the person being impersonated — the model’s only job is to look and sound convincing; it isn’t in the business of the con itself.

Step three: embedding the fabrication into social engineering. The fabricated identity gets inserted into a pre-existing trust scenario — a friend needing to route funds through your account, a boss ordering an urgent transfer, a CFO assigning a task — exploiting people’s habitual deference to existing relationships and authority, not any “intelligence” on the part of the AI.

This is also why reducing these cases to “AI turned evil” is a misreading. The AI model here functions like a well-made master key; what actually determines which door gets opened and what gets stolen is the person holding the key and the script of deception they’ve designed. Xinhua’s anti-fraud guidance also offers a technical correction: real-time face-swapping has real requirements around the volume of source material, network conditions, and how distinctive the target’s features are — not every photo can instantly generate a convincing real-time video. It also notes that video calls may still carry tells — abnormal blink rates, wandering eye focus, mismatched expressions and lip movement, disjointed speech logic — though these tells are becoming harder to catch with the naked eye as the technology advances (Xinhua, 2024; People’s Daily, 2024).

The Scam Chain

What Actually Broke Is a Trust Contract Never Built to Withstand Synthetic Attacks

Having unpacked the technical process, one more fundamental question remains: why does this “master key” happen to fit the lock called “seeing a face, hearing a voice”? Here is a judgment, not just a restatement of fact: the root cause of these fraud cases isn’t “AI running out of control.” It’s that a verification method society has relied on for decades — recognizing a voice, recognizing a face — was never designed to withstand synthetic attacks. It worked in the past only because forging it was expensive enough to be impractical.

Before deepfake technology matured, impersonating someone’s voice and likeness required specialized equipment, significant time, and a fairly high technical bar — out of reach for the ordinary scammer. That made “confirm identity via video call” an extremely cost-effective verification method, one that filtered out most scams at essentially zero cost. Deepfakes crashed that cost of forgery. You no longer need specialized equipment or a team — off-the-shelf, open-source tools can get you started. The amount and type of source material needed, and how convincing the result is, varies a lot by tool and by how high a bar you’re trying to clear — there’s no single number of samples that applies across the board. What can be said with confidence is that the cost and skill required have fallen materially; they haven’t disappeared, but they’ve fallen far enough that the math has flipped. Once forging something costs less than verifying it, the old verification method stops being “reliable” and slides toward becoming a vulnerability — while most people’s and institutions’ security habits are still stuck in the old world.

This is also why framing this story around “how scary AI is” misses the point. The question that actually deserves scrutiny is: why, even today, do so many corporate and personal transfer-confirmation processes still rely on a single communication channel that can be synthesized? Guo’s method of verification was “glance at a video.” The Hong Kong employee’s was “sit in on a meeting.” Both, fundamentally, staked identity verification on one inherently fragile information channel, with no independent second channel to cross-check.

The Rules Have Arrived — But They’re No Silver Bullet

Once the root problem is laid bare, what has regulation actually offered in response? Three years after the case broke, the response has moved from “warning the public” to “institutional mandate.”

China’s Provisions on the Administration of Deep Synthesis of Internet Information Services, which took effect in 2023, was the first regulation to require deep-synthesis service providers to build management systems covering user registration, algorithm mechanism review, content publication review, data security, personal information protection, and anti-telecom-fraud measures — and to require conspicuous labeling for synthesized faces and voices that could cause public confusion or mistaken identity (Cyberspace Administration of China, 2022).

By March 2025, the Cyberspace Administration of China, together with the Ministry of Industry and Information Technology, the Ministry of Public Security, and the National Radio and Television Administration, jointly issued the Measures for Labeling AI-Generated and Synthesized Content, paired with the mandatory national standard GB 45438-2025, Cybersecurity Technology — Labeling Methods for AI-Generated and Synthesized Content. Both took effect together on September 1, 2025. The rules require that all AI-generated text, images, audio, video, and virtual scenes carry both an explicit label (a visible badge or watermark users can see directly) and an implicit label (a marker embedded in file metadata, imperceptible but traceable). That same day, several major Chinese social platforms rolled out “AI-generated” badges, and the rules explicitly prohibit anyone from deleting, altering, forging, or concealing such labels (Xinhua, 2025; Cyberspace Administration of China, 2025).

But the labeling system is no silver bullet, and that needs to be said plainly rather than glossed over. First, the Labeling Measures themselves leave an opening — under certain conditions of user responsibility and log retention, some scenarios are exempted from explicit labeling, leaving room for the rules to be gamed. Second, labels can be re-encoded across platforms, maliciously stripped, or outright forged, and whether enforcement’s technical tracing capability can keep pace remains an open question. Third, and most critically: labeling targets the content itself, while scammers target the act of verification. A fabricated video tagged “AI-generated” is meaningless if the victim never thinks to check for the label in the first place. The system is moving forward, but it solves the problem of “can synthetic content be identified” — not the problem of “will people habitually rely on a single channel to verify identity.”

Chen Hongxiang, head of the Third Criminal Tribunal at the Supreme People’s Court, acknowledged in an early-2026 interview that the malicious abuse of AI face-swap and voice-cloning technology in telecom fraud has become “more concealed and more deceptive,” and that judicial authorities need to “accelerate research into the new situations and problems brought by deepfake technology” (Sina Finance, 2026). Translated, that sentence means: three years on, this cat-and-mouse game has no clear winner yet. The rules are catching up, but the technology keeps running.

Stop Treating “Seeing” and “Hearing” as the Finish Line

The system is still catching up — but individuals and companies don’t have to wait for it. Compressed into one takeaway: stop treating “I saw his face” or “I heard his voice” as the end point of identity verification.

At the practical level, here are a few recommendations drawn from Xinhua’s and People’s Daily’s anti-fraud guidance, along with the lessons exposed by these cases themselves (Xinhua, 2024; People’s Daily, 2024):

  • For any request involving a large transfer that carries urgency (“right now,” “someone’s waiting”) — whether it arrives by video, voice, or text — hang up first, then re-verify through a separate, independent channel: call the person’s long-saved phone number, or confirm in person.
  • Don’t treat “the face on this video” or “the voice on this call” as sufficient verification on its own, especially for instructions involving money.
  • Internal corporate transfer-approval processes should require confirmation from at least two independent information sources, not deference to a single communication channel — even when that channel is “a video conference I watched with my own eyes.”
  • Technical tells are still worth watching for, but shouldn’t be relied on: abnormal blink rates, slightly mismatched expressions and lip movement, jumps in speech logic — these are useful as supporting signals, but they’ll keep getting harder to catch by eye as the technology improves, and they can’t be the sole basis for trust.
  • On a personal level, reducing unnecessary exposure of public video and audio material — especially long, continuous conversations with a clear frontal view — objectively raises the bar for gathering source material. It’s not a complete solution, just a higher barrier.

What began with a ten-minute video call in Fuzhou ultimately points not to “how smart AI has become,” but to a plainer reminder: as forging a face or a voice keeps getting cheaper and easier, any trust built solely on “seeing” and “hearing” needs to be redesigned.


Sources