The threat posed by technology-based fraud attempts on the internet and via telecommunications networks has increased significantly as digitalization progresses. Surveys and crime statistics show that more than 70 percent of citizens in Germany have already been targeted by online criminals, with a particular increase in manipulated contact attempts being recorded during the busy summer months.
Fraudsters are increasingly using generative artificial intelligence (AI) to clone the voices of trusted individuals or bank employees with deceptive realism, or to simulate faces in video calls using deepfakes. However, security analysts point out that despite their technological advancements, these systems exhibit specific vulnerabilities in live operation that consumers can specifically test.
A key technological weakness of many current AI systems is the processing time required for real-time speech and image processing. When a conversation unfolds unpredictably, subtle delays often occur in the natural flow of speech. Experts therefore recommend confronting suspicious callers with rapid, consecutive questions that disrupt the rhythm. While a human counterpart adapts immediately, using typical filler words, automated systems usually react with a time lag or treat each question in isolation. Similar shortcomings are evident with manipulated videos: If the other party is spontaneously asked to turn the camera or interact with objects in the room, the algorithms often fail, resulting in unnatural image orientations or asynchronous background movements.
Furthermore, artificially generated media can be exposed through close observation of anatomical and linguistic details. AI models still exhibit deficiencies in the precise synchronization of lip movements and spoken sounds, particularly with so-called plosive sounds like "P," "B," or "M," for which the human mouth must be completely closed. Blinking behavior also provides important clues, as it is often rigid, mathematically uniform, or completely detached from the context of what is being said in artificial faces. Another effective test is the deliberate use of irony, sarcasm, or absurdly formulated humor. Since generative language models primarily analyze texts based on probabilities, they usually interpret ambiguous statements literally and remain rigidly bound to their predetermined conversational script.
However, critical market observers emphasize that consumer responsibility is reaching its limits, as the quality of synthetic media is rapidly improving. Detecting manipulation in the private sphere is further complicated by the psychological pressure exerted by fraudsters, for example, through shock calls. IT security experts therefore demand that protection against AI fraud not be solely reliant on the vigilance of potential victims. Rather, financial institutions and telecommunications providers have a responsibility to implement more robust technical defense mechanisms and automated verification procedures to block manipulated data streams before they reach the end user.