Cybersecurity blog header

AI poisoning techniques

AI poisoning techniques pose a threat to companies

AI poisoning techniques are sophisticated and can enable attackers to manipulate the behaviour of an AI and carry out attacks against companies

Is it possible to manipulate the recommendations of an Artificial Intelligence? The answer is yes. Microsoft recently revealed that AI tools had been made to promote certain companies in an illegitimate manner. How? By using a technique to poison the context of an AI agent, which seeks to inject persistent commands into its memory.

This case demonstrates that AI poisoning techniques are already being deployed for spurious purposes, such as carrying out deceptive marketing tactics or attacking companies that are capitalising on the key technology of this era.

In fact, MITRE ATLAS, the framework developed by the non-profit organisation MITRE, already lists several AI poisoning techniques linked to tactics such as resource development, preparing the attack on the model, or achieving persistence within the targeted AI systems.

Below, we will review some of the most dangerous AI poisoning techniques, which can lead to serious security incidents in companies that develop their own AI tools or use third-party AI agents.

To this end, we will categorise AI poisoning techniques according to the malicious tactics with which they are associated. We will also discuss how to mitigate them.

1. Resource development

One of the tactics employed by malicious actors to attack AI systems is to obtain resources that help them support their attacks during certain phases of their lifecycle, such as the preparation of attacks. How? By creating, acquiring, or stealing resources from third parties.

One of the AI poisoning techniques used to carry out this tactic is the publication of poisoned AI artefacts.

1.1. Publishing poisoned AI artefacts

What does this technique involve? Malicious actors create or modify datasets, models, or AI agent tools that contain malicious content, code, or configurations, and make them publicly available so that their potential victims can acquire them or integrate them into their systems.

The aim of these poisoned AI artefacts is to facilitate the compromise of AI systems through an AI supply chain attack.

As with other AI poisoning techniques, we can divide the publication of poisoned AI artefacts into three sub-techniques:

  1. Publication of tainted datasets that can be used to train or fine-tune models. These datasets may contain labels, annotations, or metadata that have been manipulated by malicious actors, with the aim of undermining the AI model training process.
  2. Publication of poisoned AI models in model registries or code repositories. AI models may contain configurations, architecture, or components that have been manipulated to cause the models to exhibit malicious behaviour or execute malware.
  3. Publication of compromised AI tools. These tools may include malicious definitions or instructions visible to the AI model, as well as hidden executable behaviours or responses designed to manipulate an AI agent. This is done so that the AI agent behaves maliciously when it executes the tool. An example of this sub-technique is the malicious tools found on ClawHub.

How can the use of these AI poisoning techniques by malicious actors be mitigated? MITRE ATLAS recommends three mitigation techniques:

  1. Clean the training data before using it to detect contaminated labels, annotations, or metadata.
  2. Validate the AI models to be used to detect backdoors, data leaks, or unexpected behaviour.
  3. Carry out a vulnerability analysis on AI agent models and tools to detect malicious content.

2. Execution

One of the most common execution techniques is execution by the victim themselves. In other words, using social engineering techniques to persuade a user to execute a malicious or insecure artefact, thereby opening the door to the execution of malicious code on the system or causing the AI to engage in malicious behaviour.

2.1. AI agent poisoning

A key sub-technique within this approach is AI agent poisoning. Through this, malicious actors trick a victim into invoking a poisoned tool whilst interacting with an AI agent, exposing them to tainted responses or the execution of the implementation logic of a malicious tool. This technique was also used in attacks against OpenClaw.

To mitigate this technique, the following measures can be taken:

  • Introduce signature checks to prevent the execution of AI artefacts that may pose a risk to the system.
  • Carry out a vulnerability analysis on tools and models before connecting them to the agent.
  • Train users of AI agents to identify manipulation attempts and reduce the likelihood of them performing actions that enable the execution of malware.
Cybersecurity for AI is key in this era

3. Persistence

The vast majority of AI poisoning techniques relate to persistence. In other words, they concern the ability of malicious actors to maintain access to the systems they attack in order to achieve their objectives.

3.1. Manipulating the AI model

The first of the AI poisoning techniques designed to facilitate persistence is the manipulation of an AI model.

Using this technique, malicious actors manipulate an AI model or its components to:

  • Alter the behaviour of the AI system.
  • Introduce malicious code.
  • Establish persistent malicious functionalities.

By manipulating an AI model, hostile actors can ensure that malicious behaviour is triggered only when certain conditions or contexts arise, thereby remaining undetected under normal conditions.

What AI poisoning sub-techniques might malicious actors deploy in these cases?

  1. Poisoning the AI model to alter its behaviour or performance.
  2. Modifying the model’s architecture to change its behaviour with the aim of removing predictive capabilities, increasing computational costs, undermining its performance, or even creating a backdoor.
  3. Injection of malicious code into AI model files. This enables models with embedded malware to allow malicious actors to execute actions, implement command-and-control techniques, or exfiltrate data.
  4. Modification of the logic used to construct prompts. To what end? To be able to persistently inject instructions or manipulate the model’s outputs.

To mitigate these AI poisoning sub-techniques, MITRE ATLAS recommends:

  • Controlling access to data at rest in AI models to prevent unauthorised tampering.
  • Validate AI models by testing them against malicious inputs to ensure they cannot be manipulated.
  • Implement a code signature to ensure that a model has not been tampered with after deployment.
  • Conduct a Red Team exercise, in which cybersecurity experts simulate attacks aimed at manipulating a model, in order to detect weaknesses and rectify them.

3.2. Training data poisoning

One of the most common AI poisoning techniques involves manipulating the data used to train or fine-tune a model. Malicious actors can contaminate the data by gaining unauthorised access to the training process or by injecting it via AI supply chain attacks.

What are the objectives of this technique?

  • Cause errors in AI systems.
  • Induce biased or unsafe behaviour.
  • Undermine the model’s performance.
  • Inject backdoors that are triggered by specific inputs.

How can training data poisoning be mitigated?

  • Restricting the publication of datasets.
  • Controlling access to data at rest in AI models.
  • Cleaning training data to detect alterations that could cause a model to act maliciously or enable attacks through the existence of backdoors.
  • Validating the AI model.
  • Maintaining an AI bill of materials to identify unreliable components.
  • Verifying the provenance of datasets.
  • Conducting Red Team exercises to verify the effectiveness of data ingestion and cleansing processes, as well as those for detecting tampering.

3.3. RAG Contamination

Another AI poisoning technique linked to persistence in AI systems involves injecting malicious content into the data indexed by a Retrieval-Augmented Generation (RAG) system. In this way, cybercriminals can contaminate future threads via RAG-based search results. How? By embedding manipulated documents in locations that the RAG indexes.

What can be achieved by using this technique? AI systems may display content containing false data, misleading information, or even malicious instructions.

To combat this AI poisoning technique, it is possible to:

  • Implement control measures for generative AI and ensure that content indexed by the RAG that is unreliable or potentially malicious is rejected.
  • Carry out Red Team exercises in which malicious content is introduced in a controlled manner into the information ingestion sources. This allows for the evaluation and optimisation of mechanisms for authorising sources, verifying their provenance, or validating their content.

3.4. Context poisoning of an AI agent

There are three AI poisoning techniques designed to manipulate and alter the behaviour of AI agents used by companies in their day-to-day operations to automate hundreds of processes.

The first of these techniques is context poisoning, which exploits the LLM model of an AI agent to generate its responses or carry out actions.

Using this technique, a malicious actor can persistently modify the behaviour of an AI agent and turn it to their own ends. How is this technique carried out? By instructing the LLM to:

  • Add instructions to its memory.
  • Use previous messages from a thread as context.

To mitigate this technique, the following measures can be taken:

  • Strengthen the agent’s memory to reduce persistent context poisoning. To do this, mechanisms can be implemented to control what an agent stores in its memory and to detect compromised records.
  • Conduct a Red Team exercise in which attempts are made to add malicious instructions to the agent’s memory, and verify the authorisation mechanisms for changing the context and other processes such as integrity checks.

3.5. Manipulation of an AI agent

Malicious actors can poison the tools used by AI agents by introducing malicious content into their definition, implementation, or responses.

Once a manipulated tool is installed on or connected to an agent, it may enable malicious actors to gain persistent influence over the AI agent’s actions.

Thus, a compromised tool can cause an agent to access confidential data, alter inputs, extract information, conceal actions from users, or execute unauthorised commands.

3.6. Data contamination in AI agent tools

This technique involves the manipulation of a data source used by an AI agent as a source of information. For example, malicious actors may alter data in a source controlled by their victim or in a trusted source to which they have access.

Furthermore, they may design malicious content to be displayed in routine queries of the AI agent tool or during retrieval operations. Such content may include false information or malicious instructions with the aim of carrying out an indirect prompt injection into the LLM model.

It is possible to mitigate the use of this technique by attackers by implementing control measures in content retrieval processes, so that content from untrusted sources can be rejected.

4. Preparing an AI attack

One of the AI poisoning techniques we have already discussed can also be used when preparing an attack. We are referring to the poisoning of an AI model by manipulating its weights, training it with altered data, or interfering with its training process.

4.1. Poisoning an AI model

Poisoning can be used to alter the behaviour of generative AI by training it with false or biased information. How can this technique be mitigated?

  • Control access to AI models and data at rest.
  • Clean the data.
Some AI poisoning techniques aim to alter the behaviour of models

5. Impact

Malicious actors have also developed AI poisoning techniques that are deployed during the impact phase of attacks: forcing excessive resource consumption and compromising the integrity of a dataset.

5.1. Increased resource consumption by the AI agent

This technique can be used to force an AI system to connect to tools, making unnecessary API calls and thereby causing it to consume vast amounts of computational and financial resources.

It is also possible to force the AI agent to waste resources through self-delegation loops, in which it is forced to delegate additional tasks to itself, potentially causing stack overflows and service interruptions.

To mitigate this technique, the following measures can be taken:

  • Conduct a Red Team exercise that simulates this technique and attempts to force the AI agent to make repeated API calls, enter a recursive loop, or self-delegate tasks. This makes it possible to detect weaknesses, verify API quotas, and set limits on the number of iterations or timeouts.
  • Limit the AI agent’s resource consumption.

5.2. Undermining the integrity of a dataset

Malicious actors can poison parts of a dataset to undermine its usefulness and reduce confidence in it. This means that organisations using it to train their AI models will have to waste resources on rectifying errors and detecting tampered data.

To combat the final AI poisoning technique we will address in this guide, you can:

  • Clean contaminated training data to ensure the integrity of the dataset.
  • Maintain control over the provenance of the dataset to identify unauthorised modifications more easily.

6. AI security audit: A shield against AI poisoning techniques

If this era is already being shaped by the growing importance of AI systems within the business ecosystem, it should come as no surprise that malicious actors are targeting AI models and agents.

To prevent hostile actors from exploiting vulnerabilities and deploying AI poisoning techniques, or other techniques such as prompt injections, it is essential to carry out regular AI security audits. This type of assessment makes it possible to verify that:

  • An AI model operates within its intended parameters.
  • The confidential data with which an AI agent works is protected.
  • The operational integrity of the AI system and the organisation itself is guaranteed.

An AI security audit enables you to:

  • Review the architecture of an AI model, including its data sources and training and deployment processes, to prevent the use of AI poisoning techniques.
  • Identify critical points that may be susceptible to vulnerabilities, such as user interfaces or integrations with third-party tools.
  • Simulate prompt injection or AI poisoning attacks.
  • Verify that the system does not disclose confidential information in its responses or interactions.
  • Analyse third-party components to detect vulnerabilities and prevent AI supply chain attacks.
  • Establish security configurations in line with best practice.
  • Produce a detailed report setting out the vulnerabilities identified, their potential impact, and the actions that can be taken to address them, prioritising them according to the level of risk posed by the weaknesses.
  • Train the organisation’s staff to identify and prevent vulnerabilities.

7. The role played by the Red Team in combating AI poisoning techniques

When addressing the various AI poisoning techniques, one mitigation approach has been consistently highlighted: the conduct of Red Team exercises.

Given the complexity of the threats facing AI and the development and refinement of malicious tactics and techniques, Red Team services have become a cornerstone of any cybersecurity strategy for AI systems.

By devising specific Red Team scenarios, cybersecurity experts can simulate 100 per cent realistic attacks using the techniques that malicious actors might employ to compromise AI systems and disrupt their operation or deploy malware.

Red Teams are key to validating the level of cyber resilience of AI systems and verifying the effectiveness of the defensive mechanisms deployed to protect data, models, and AI agents. Furthermore, they play a critical role in detecting exploitable vulnerabilities and drawing up a list of improvements that can be implemented to prevent AI poisoning techniques from succeeding.