AI Supply Chain Security defines attacks by technique, providing an estimated severity and the rationality for classifying it with that severity.
- Detections - Known exploits exist in the model files. The detections are ranked from Critical to Low severity and can be used to define Supply Chain policy.
- Advisories - Known files of concern, but are not exploits in and of themselves. Advisories should be reviewed prior to model usage.
| Detection Category | Estimated Severity | Definition | Rationality for Severity |
|---|---|---|---|
| Arbitrary Code Execution | Critical | Adversaries can inject malicious code into a model, which will be executed whenever the hijacked model is loaded into memory. This vulnerability can be used to exfiltrate sensitive data, execute malware (such as spyware or ransomware) on the machine, or run any kind of malicious scripts.
| Arbitrary code execution attacks are relatively easy to perform and may lead to critical outcomes such as execution of malicious code on an organization’s computers.
|
| Arbitrary Read Access | High | Adversaries can craft a malicious model that will exfiltrate sensitive data upon loading.
| Arbitrary read access attacks are relatively easy to perform and may lead to critical outcomes such as an attacker exfiltrating sensitive data.
|
| Control Vector | High | Adversaries can inject a control vector into the computational graph of a model introducing refusal ablation or custom, attacker-defined behaviors.
| Inserted control vectors can control or modify model behavior as well as can be used to remove refusals on a secured model.
|
| Decompression Vulnerabilities | High | Adversaries can exploit vulnerabilities in popular compression formats to cause denial of service or leak sensitive data.
| Decompression vulnerabilities are relatively easy to exploit and may lead to high-impact outcomes such as denial of service, code execution, or data leakage.
|
| Denial of Service | Medium | Adversaries can craft a malicious model, or modify legitimately pre-trained model, in order to disrupt the system the model will be loaded on.
| Denial of service attacks are relatively easy to perform and may lead to disruption or degradation of service.
|
| Directory Traversal | Medium | Adversaries can craft a malicious model, or modify legitimately pre-trained model, in order to gain unauthorised access to sensitive files on the system.
| Directory traversal attacks are relatively easy to perform and may grant an attacker access to sensitive files on the file system.
|
| Embedded Payloads | Low | Adversaries can embed malicious payloads (such as backdoors, coin miners, spyware, and ransomware) inside the model’s tensors. Such payloads can be injected in plain text, obfuscated, or embedded using steganography.
| Malicious payloads can be embedded in ML models relatively easily; this may lead to malware components being distributed on an organization’s computers.
|
| Graph Payload | High | Adversaries can inject a computational graph payload, introducing a secret attacker-controlled behavior into a pre-trained model.
| Model backdooring may be relatively difficult to perform and can lead to critical outcomes such as biased or inaccurate output.
|
| Model Sideloading | High | Adversaries can load code or model artifacts from an unexpected location bypassing checks performed on the model.
| Model sideloading is not expected behavior and typically points to an attempt to obfuscate a payload. |
| Network Requests | High | Adversaries can craft a malicious model that will make network requests upon loading.
| Network requests are relatively easy to perform and may be used to exfiltrate data, download payloads, or initiate command and control communications.
|
| Repository Sideloading | Medium | Adversaries can load code or model artifacts from an unexpected location, bypassing checks performed on the artifacts in the repository. | Repository sideloading is an expected behavior allowed by Hugging Face; however, it can be abused to bypass security checks.
|
| Suspicious File Format | Medium | Adversaries can modify data structures and encodings in an attempt to evade detection.
| File format tampering is usually indicative of a targeted attack.
|
| Suspicious Functions | High | The presence of these functions themselves is not inherently malicious, but they can be used in conjunction with other functions to create a malicious model.
| Functions can be used in conjunction with other functions to create a malicious model.
|
| TokenBreak | High | Adversaries can exploit a weakness in the tokenizer to bypass model classifications. | Models susceptible to the tokenbreak vulnerability can have their classifications altered on command by an attacker resulting in a weakness wherever the model is used. |