Skip to content
Mindcraft Impuls

AI agents:
When the voice of the authority is not human at all

Approx. 6 minutes read

Experience this topic as an interactive Cyber Snack:
just click and learn it all in 5 minutes.

Cyber Snack: start AI agents interactively

The phone rings, and the caller claims to be from a cybersecurity agency. The voice sounds calm, competent and urgent: your computer may be affected by well-known malware, and you should install a diagnostic tool right away. What used to be an elaborately prepared scam call can now be handled by an AI agent - autonomously, convincingly and hundreds of times at once.

Deepfake voices are not a new topic. What is new is that AI no longer just provides a voice, but runs the entire conversation: it responds to questions, notices hesitation, reassures, pushes and adapts its arguments. This fundamentally changes social engineering - at work and just as much on private phones.

The August Cyber Snack makes exactly that tangible: it shows how an attack using AI agents can unfold, why it scales so well, and which simple rules help employees and families protect themselves.

Real or AI? On the phone, hardly distinguishable

In the Cyber Snack, you hear four voices and have to decide which one is AI-generated. The answer: all four. That is exactly the point. Current speech models produce voices with natural intonation, small pauses and fitting emotion. On the phone, with the slight distortion of the line, the difference is even harder to notice.

We have already described how effective fake voices and faces can be in our insight Web meetings - including the Arup case, in which an employee in Hong Kong transferred around 25 million US dollars after a video conference with fake colleagues. Back then, a lot of human preparation went into the attack. With AI agents, exactly that effort disappears.

How such an attack can unfold

Illustration of a phone call with an official badge and a download prompt on a laptop
An official-sounding call, an urgent request, a download - that is all the attack needs.

The background is real: in September 2025, Mandiant and the Google Threat Intelligence Group reported on BRICKSTORM, a backdoor used by suspected China-nexus attackers to spy on law firms, SaaS providers and technology companies, among others. It mainly hid on network appliances and VMware servers where classic endpoint security barely reaches, and on average remained undetected for more than a year. In December 2025, the US agency CISA, the NSA and the Canadian Cyber Centre followed up with a joint analysis.

Headlines like these are the perfect hook for fraudsters. The Cyber Snack plays this through in a realistic scenario: an employee in a New York office starts the day as usual. Then the phone rings. A supposed CISA officer is on the line, calm and credible: "We have indications that your computer may be affected by BRICKSTORM. Please stay on the line, we need to stop the malware from spreading."

The employee knows the news reports. The call sounds like help, not like an attack - and that is exactly how the manipulation begins. He is asked to download a "diagnostic tool". He does. What he has installed is not help from an agency, but malware. The twist: on the other end of the line there was no human, but an AI agent.

The key question is no longer: 'Does this voice sound real?' It is: 'Am I following the process - no matter how real it sounds?'

What is genuinely new about AI agents

Illustration of an AI agent calling many smartphones at the same time
One AI agent can hold many conversations at the same time. The target is whoever picks up.

In the past, deepfake calls were mostly prepared audio clips or voices controlled live by a human. That took effort. This is why such attacks were mainly aimed at selected individuals: people with budget responsibility, access to confidential data or influence over payments.

An AI agent works differently. The attacker only sets a goal, for example: "Get the person to download a diagnostic tool." The AI finds its own way there. It notices skepticism, explains patiently, applies pressure at the right moment, flatters or threatens - depending on what works with the person on the other end. That is agentic AI: it pursues a goal and acts autonomously to reach it.

The decisive difference is scale. In the past, a social engineer made every call personally, one after the other. Today, an AI agent can hold hundreds of conversations in parallel. The target is then no longer a specific person, but anyone who happens to pick up, is stressed and clicks. Social engineering becomes faster, cheaper and scalable.

This is not science fiction. Researchers at the University of Illinois showed as early as October 2024 that voice-enabled AI agents can autonomously carry out common phone scams - at an average cost of less than one US dollar per attempt. In April 2026, Abnormal Security described ATHR, a platform sold on criminal forums on which an AI voice agent conducts vishing calls automatically and a single operator manages many conversations at once. And in May 2025, the FBI warned of a campaign in which criminals used AI-generated voice messages to impersonate senior US government officials.

The real CISA also warns about fake CISA calls

The scenario in the Cyber Snack is fictional, but the scam behind it is not. As early as June 2024, the US agency CISA publicly warned of fraudsters impersonating its staff on the phone. Its advice: real CISA employees never ask for wire transfers, cash, cryptocurrency or gift cards - and never ask you to keep the conversation secret.

Not just companies: families and children are targets too

You may be thinking: this affects companies, but not me as a private person. Unfortunately, it does. AI agents can impersonate authorities, schools, sports clubs or familiar people - on private phones too. And they do not only call adults.

Imagine your child receives a call, supposedly from the police: "Your parents have had an accident. Is anyone at home? How can we reach your family?" Or the agent poses as a teacher, coach or family friend and asks for the address, daily routine, names, holiday plans or whether the child is home alone. Children and teenagers are particularly vulnerable when a voice sounds adult, official and urgent.

The more real and urgent a call sounds, the harder it is to stay calm. That is exactly what fraudsters exploit. An AI agent never gets impatient. It repeats, reassures, threatens or flatters for as long as it takes until something works.

Five rules for suspicious calls

The good news: the best protection is not technical, but human and informed. Anyone who knows the following rules does not need to recognize whether a voice is real. They just need to follow the process.

1. Authorities do not make urgent surprise calls

Real authorities use official channels and known contact routes. An unexpected, urgent call demanding immediate action is a warning sign.

2. No remote access, no software, no credentials

Legitimate organizations never ask for remote access, software installations or confidential credentials over the phone. Do not click, do not install, do not grant access.

3. Do not discuss - hang up

An AI agent has an answer to every objection. Do not get drawn into a discussion. End the call politely but firmly.

4. Call back via an official number

Look up the number yourself, for example on the official website or the intranet. Never dial the number the caller gives you - it can be just as fake as their story.

5. Report it, even if nothing happened

At work, every such call should be reported to IT security - even if you did not click anything. That way, colleagues can be warned in time.

Family code word and role play: how to protect your children

Illustration of a family under a protective shield with a key symbol that blocks a suspicious call
An agreed family code word protects where the voice alone can no longer be trusted.

Families have a simple but very effective tool: a family code word. It is a secret word or short phrase known only to the family and selected trusted people - for example "Blue Panda" or "Operation Kiwi". Anyone who calls on behalf of the family and urgently wants something must be able to say the code. If they cannot, the call ends.

It is important that children know how to react when someone asks for their address, school or passwords, or wants to know whether they are alone:

  • Say stop and do not panic.
  • Ask for the code word.
  • End the call.
  • Call the parents or a known trusted person directly.
  • Never act immediately - not even when the voice sounds official.

The best way to practise this is through short role plays. One parent plays the fraudster, and the child practises: stay calm, ask for the code word, hang up and call the parents themselves. It is not about perfection, but about pausing briefly instead of answering automatically.

What awareness teams should take from this

For CISOs and awareness managers, the most important lesson is this: the question "Is this voice real?" can no longer be answered reliably on the phone. Training that relies on spotting deepfakes therefore falls short. Clear, simple procedures that work regardless of how convincing a caller sounds are more effective.

This includes: employees need an easy-to-find official phone number for verification, permission to turn away even supposedly senior callers, and a low-threshold reporting channel - explicitly also for calls where nothing happened. Every report helps to detect an ongoing wave early and to warn colleagues.

Our insight Ransomware shows how quickly a single installed “tool” can turn into a serious security incident. And the core question from Online services and AI tools applies here too: what information am I giving out - and to whom, actually?

Conclusion

AI agents are not dangerous in themselves. In the hands of fraudsters, however, they make social engineering faster, more convincing and more scalable. One voice, one authority, one urgent download - that is all it takes if nobody pauses.

Do not blindly trust the voice. Trust the process: do not discuss, do not click, hang up, call back via an official number and report it at work.

Sources

Google Threat Intelligence Group / Mandiant, September 24, 2025: "Another BRICKSTORM: Stealthy Backdoor Enabling Espionage into Tech and Legal Sectors"; CISA, NSA and Canadian Centre for Cyber Security, December 4, 2025: Malware Analysis Report "BRICKSTORM Backdoor" (AR25-338A).

CISA, June 12, 2024: alert on scammers impersonating CISA employees by phone.

FBI / IC3, May 15, 2025: Public Service Announcement "Senior US Officials Impersonated in Malicious Messaging Campaign" (I-051525-PSA).

Fang, Bowman, Kang et al. (University of Illinois Urbana-Champaign), October 2024: "Voice-Enabled AI Agents can Perform Common Scams", arXiv:2410.15650; Abnormal Security, April 16, 2026: "AI Meets Voice Phishing: How ATHR Automates the Full TOAD Attack Chain".

Topic cluster