T
he fact that large language models harbour serious security vulnerabilities and can be easily fooled using clever methods necessitates a fundamental paradigm shift in our approach to these technologies. Currently, technology companies are racing to release and commercialize these models as quickly as possible, often treating ethical and security testing as mere formalities. However, the risks we face cannot be mitigated with a simple software update.
Presented as the new oracles of the modern world, large language models are permeating every aspect of daily life at an astonishing pace, and humanity is captivated by the allure of the bright future promised by this technology. These systems, which instantly answer our questions, generate complex texts, and even create works of art, are marketed as digital panaceas. However, on the flip side, a deep and systemic weakness lies at the heart of these dazzling capabilities. Artificial intelligence’s greatest strength, namely its ability to understand human language and carry out instructions to the letter, is also its most dangerous Achilles heel.
These models, though programmed like loyal servants, are fragile systems incapable of distinguishing between a well-intentioned instruction and malicious manipulation. They can easily be thrown off course by a cleverly crafted story or a devious prompt. This situation is far more than a simple software bug. The problem points to an illusion of control stemming not from a few lines of faulty code in these systems, but from their foundational architecture /core design.. As the weakness of the control mechanisms behind these technologies becomes apparent, the potential for difficult-to-reverse problems across a broad spectrum, from individual privacy to national security, also becomes clear.
Simple prompt threat
The most basic form of this new generation threat manifests itself through a method known as prompt injection. In the simplest form of this method, known as direct injection, a malicious user instructs the artificial intelligence to forget all previous prompts and obey new instructions given to it. Indeed, in 2023, a Stanford student succeeded in exposing the model’s secret internal guidelines, codenamed ‘Sydney,’ with a simple prompt such as ‘Ignore previous instructions’ directed at Microsoft’s Bing chatbot, proving how fundamental this vulnerability is. This incident demonstrated how vulnerable artificial intelligence is not to complex cyberattacks, but to simple and clear statements.
However, the threat does not end there. In the much more insidious and dangerous indirect injection method, malicious prompts are hidden within external data sources, such as a website, email, or PDF file, that are presented for the artificial intelligence to process. For example, an email containing the prompt ‘delete all my emails’ in invisible text, which the AI assistant unknowingly executes when asked to summarise the email, turns every piece of data the AI interacts with into a potential minefield.
This situation shifts the attack surface from a finite number of code vulnerabilities to the infinite potential for creativity and deception inherent in human language. While traditional cybersecurity paradigms make a clear distinction between code and data, this distinction disappears in large language models. In other words, every piece of data becomes a potential prompt, and every prompt becomes a potential threat.
Manipulation techniques not only distort specific tasks of artificial intelligence but also aim to eliminate its entire moral and ethical compass. These methods, known as jailbreaking, are designed to break the model’s security and ethical barriers. One of the most striking examples of this is the technique known as the ‘Grandma Exploit’ or ‘Dead Grandmother Trick’.. A user crafted an innocent and emotional story by telling the AI that his late grandmother was a chemical engineer at a napalm factory and that when he was little, she would lull him to sleep by singing him a lullaby about the steps involved in making napalm. Faced with this emotional manipulation, the AI, which would normally refuse to disclose dangerous information, bypassed its security protocols and provided the requested harmful information.
This example demonstrates how easily AI defence mechanisms can be manipulated and bypassed using human emotions and fictional narratives rather than logical arguments. The model can easily stray in the wrong direction when faced with a moral dilemma because it can only process the surface meaning and emotional tone of a prompt, not the intention behind it.
Perhaps the most worrying and fundamental attack on large language models is data poisoning carried out during the model’s training phase. Attackers deliberately inject false, biased, or malicious information into the massive datasets used to train artificial intelligence. One study showed that poisoning just 0.001% of a medical AI’s training data could cause the model to produce dangerous and incorrect medical diagnoses.
Detecting and correcting this type of manipulation, which has the potential to transform the system into a permanent disinformation tool by corrupting its fundamental information source, is nearly impossible. Once the model learns the poisoned data, it accepts this error as correct and repeats it in all future outputs.
Global and local vulnerabilities
The real-world implications of these theoretical vulnerabilities have begun to cause tangible harm at both the individual and organisational levels. Cases such as a Chevrolet chatbot being persuaded to sell a car for one dollar or an Air Canada chatbot providing incorrect refund information, forcing the company to comply with this decision, highlight the immediate financial and reputational risks of these technologies. Samsung employees unknowingly leaking confidential company data via ChatGPT is proof of how easily privacy and corporate security breaches can occur.
However, a significant danger is that these models can be used en masse to generate personalized and highly persuasive phishing emails, financial fraud schemes, and political disinformation. This has the potential to fundamentally undermine public trust and social cohesion. At the national security level, these vulnerabilities can become weapons. Marginalised groups could use manipulated large language models to conduct automated propaganda and disinformation campaigns on a large scale and at high speed, destabilizing democratic processes and poisoning cultural narratives.
A single individual with a cleverly designed prompt can jeopardize a system used by millions of people and automatically generate and disseminate harmful content through that system. This situation places an unpredictable threat to global and local stability by placing an influence capacity previously only held by superpowers into the hands of small groups.
The fact that large language models harbour such serious security vulnerabilities and can be easily deceived using clever methods necessitates a fundamental paradigm shift in our approach to these technologies. Currently, technology companies are racing to release and commercialise these models as quickly as possible, often treating ethical and security testing as mere formalities. However, the risks we face cannot be mitigated with a simple software update. These models must be trained and guided to understand the subtleties, deceptions, and emotional manipulations of human language and to resist them. Unfortunately, rule-based control of inputs and outputs is insufficient and even triggers various debates, including criticism of censorship. This process will require in-depth, interdisciplinary and extremely rigorous testing procedures that may take years rather than months.
Testing these models over long periods of time, under different scenarios, from an ethical perspective and in terms of security risks, is not a choice but a necessity. Otherwise, it will be inevitable that irreversible individual and societal harms will emerge, ranging from individual privacy to national security, social peace, and the functioning of democratic processes. No matter how great the potential benefits of artificial intelligence may be, the cost of releasing these benefits unchecked and unregulated could be very heavy for humanity.
Therefore, rather than blindly believing in these models, which are perceived by many as omniscient, flawless oracles, understanding their fragile intelligence and hidden dangers and acting accordingly is the wisest step we can take for our future. Otherwise, this technology, regarded as one of humanity’s greatest inventions, could become a weapon that turns against its creator, slipping beyond our control, much like the tiger cub in Geoffrey Hinton’s analogy, and inflict irreparable wounds in the real world rather than the digital one.
(Originally published in Turkish by Kriter).
Recommended





