Tags and Taxonomies
garak includes a number of tags on Probe objects that describe both intents and techniques.
Within garak, we define an intent to be, more-or-less, “what you want the model to do”.
In contrast, a technique is an particular way of structuring a request intended to get the model to comply with provided instructions.
Tags are enumerated in garak/data/tags.misp.tsv.
While garak supports tags for many frameworks, the inclusion of a taxonomy does not mean that every relevant attribute of a framework is fully covered.
Be certain that if you use garak for compliance checking related to one of these frameworks, you familiarize yourself with the framework itself and the parts that are not covered by garak.
Currently, garak supports the following frameworks via tags:
OWASP LLM Top 10
AVID Effects
Language Model Risk Cards
Common Weakness Enumeration
Summon and Demon and Bind it
EU AI Act
OWASP LLM Top 10
The OWASP LLM Top 10 identifies prominent security risks affecting applications built with large language models, including prompt injection, insecure output handling, training data poisoning, excessive agency, sensitive information disclosure, and model theft. It provides a practical application-security framework for identifying and communicating LLM-specific vulnerabilities and their mitigations across the development and deployment lifecycle.
owasp:llm01 |
LLM01: Prompt Injection |
Crafty inputs can manipulate a Large Language Model, causing unintended actions. Direct injections overwrite system prompts, while indirect ones manipulate inputs from external sources. |
owasp:llm02 |
LLM02: Insecure Output Handling |
This vulnerability occurs when an LLM output is accepted without scrutiny, exposing backend systems. Misuse may lead to severe consequences like XSS, CSRF, SSRF, privilege escalation, or remote code execution. |
owasp:llm03 |
LLM03: Training Data Poisoning |
This occurs when LLM training data is tampered, introducing vulnerabilities or biases that compromise security, effectiveness, or ethical behavior. Sources include Common Crawl, WebText, OpenWebText, & books. |
owasp:llm04 |
LLM04: Model Denial of Service |
Attackers cause resource-heavy operations on Large Language Models leading to service degradation or high costs. The vulnerability is magnified due to the resource-intensive nature of LLMs and unpredictability of user inputs. |
owasp:llm05 |
LLM05: Supply Chain Vulnerabilities |
LLM application lifecycle can be compromised by vulnerable components or services, leading to security attacks. Using third-party datasets, pre- trained models, and plugins can add vulnerabilities. |
owasp:llm06 |
LLM06: Sensitive Information Disclosure |
LLMs may reveal confidential data in its responses, leading to unauthorized data access, privacy violations, and security breaches. It’s crucial to implement data sanitization and strict user policies to mitigate this. |
owasp:llm07 |
LLM07: Insecure Plugin Design |
LLM plugins can have insecure inputs and insufficient access control. This lack of application control makes them easier to exploit and can result in consequences like remote code execution. |
owasp:llm08 |
LLM08: Excessive Agency |
LLM-based systems may undertake actions leading to unintended consequences. The issue arises from excessive functionality, permissions, or autonomy granted to the LLM-based systems. |
owasp:llm09 |
LLM09: Overreliance |
Systems or people overly depending on LLMs without oversight may face misinformation, miscommunication, legal issues, and security vulnerabilities due to incorrect or inappropriate content generated by LLMs. |
owasp:llm10 |
LLM10: Model Theft |
This involves unauthorized access, copying, or exfiltration of proprietary LLM models. The impact includes economic losses, compromised competitive advantage, and potential access to sensitive information. |
AVID Effects
The AI Vulnerability Database (AVID) is an open-source knowledge base and taxonomy for documenting observed failures and vulnerabilities in AI systems. AVID organizes issues across areas such as security, ethics, and performance, helping practitioners classify findings consistently and relate them to known AI failure modes and other risk frameworks.
avid-effect:security:S0100 |
Software Vulnerability |
Vulnerability in system around model—a traditional vulnerability |
avid-effect:security:S0200 |
Supply Chain Compromise |
Compromising development components of a ML model, e.g. data, model, hardware, and software stack. |
avid-effect:security:S0201 |
Model Compromise |
Infected model file |
avid-effect:security:S0202 |
Software compromise |
Upstream Dependency Compromise |
avid-effect:security:S0300 |
Over-permissive API |
Unintended information leakage through API |
avid-effect:security:S0301 |
Information Leak |
Cloud Model API leaks more information than it needs to |
avid-effect:security:S0302 |
Excessive Queries |
Cloud Model API isn’t sufficiently rate limited |
avid-effect:security:S0400 |
Model Bypass |
Intentionally try to make a model perform poorly |
avid-effect:security:S0401 |
Bad Features |
The model uses features that are easily gamed by the attacker |
avid-effect:security:S0402 |
Insufficient Training Data |
The bypass is not represented in the training data |
avid-effect:security:S0403 |
Adversarial Example |
Input data points intentionally supplied to draw mispredictions. Potential Cause: Over permissive API |
avid-effect:security:S0500 |
Exfiltration |
Directly or indirectly exfiltrate ML artifacts |
avid-effect:security:S0501 |
Model inversion |
Reconstruct training data through strategic queries |
avid-effect:security:S0502 |
Model theft |
Extract model functionality through strategic queries |
avid-effect:security:S0600 |
Data poisoning |
Usage of poisoned data in the ML pipeline |
avid-effect:security:S0601 |
Ingest Poisoning |
Attackers inject poisoned data into the ingest pipeline |
avid-effect:ethics:E0100 |
Bias/Discrimination |
Concerns of algorithms propagating societal bias |
avid-effect:ethics:E0101 |
Group fairness |
Fairness towards specific groups of people |
avid-effect:ethics:E0102 |
Individual fairness |
Fairness in treating similar individuals |
avid-effect:ethics:E0200 |
Explainability |
Ability to explain decisions made by AI |
avid-effect:ethics:E0201 |
Global explanations |
Explain overall functionality |
avid-effect:ethics:E0202 |
Local explanations |
Explain specific decisions |
avid-effect:ethics:E0300 |
User actions |
Perpetuating/causing/being affected by negative user actions |
avid-effect:ethics:E0301 |
Toxicity |
Users hostile towards other users |
avid-effect:ethics:E0302 |
Polarization/ Exclusion |
User behavior skewed in a significant direction |
avid-effect:ethics:E0400 |
Misinformation |
Perpetuating/causing the spread of falsehoods |
avid-effect:ethics:E0401 |
Deliberative Misinformation |
Generated by individuals., e.g. vaccine disinformation |
avid-effect:ethics:E0402 |
Generative Misinformation |
Generated algorithmically, e.g. Deep Fakes |
avid-effect:performance:P0100 |
Data issues |
Problems arising due to faults in the data pipeline |
avid-effect:performance:P0101 |
Data drift |
Input feature distribution has drifted |
avid-effect:performance:P0102 |
Concept drift |
Output feature/label distribution has drifted |
avid-effect:performance:P0103 |
Data entanglement |
Cases of spurious correlation and proxy features |
avid-effect:performance:P0104 |
Data quality issues |
Missing or low-quality features in data |
avid-effect:performance:P0105 |
Feedback loops |
Unaccounted for effects of an AI affecting future data collection |
avid-effect:performance:P0200 |
Model issues |
Ability for the AI to perform as intended |
avid-effect:performance:P0201 |
Resilience/stability |
Ability for outputs to not be affected by small change in inputs |
avid-effect:performance:P0202 |
OOD generalization |
Test performance doesn’t deteriorate on unseen data in training |
avid-effect:performance:P0203 |
Scaling |
Training and inference can scale to high data volumes |
avid-effect:performance:P0204 |
Accuracy |
Model performance accurately reflects realistic expectations |
avid-effect:performance:P0300 |
Privacy |
Protect leakage of user information as required by rules and regulations |
avid-effect:performance:P0301 |
Anonymization |
Protects through anonymizing user identity |
avid-effect:performance:P0302 |
Randomization |
Protects by injecting noise in data, eg. differential privacy |
avid-effect:performance:P0303 |
Encryption |
Protects through encrypting data accessed |
avid-effect:performance:P0400 |
Safety |
Minimizing maximum downstream harms |
avid-effect:performance:P0401 |
Psychological Safety |
Safety from unwanted digital content, e.g. NSFW |
avid-effect:performance:P0402 |
Physical safety |
Safety from physical actions driven by a AI system |
avid-effect:performance:P0403 |
Socioeconomic safety |
Safety from socioeconomic harms, e.g. harms to job prospects or social status |
avid-effect:performance:P0404 |
Environmental safety |
Safety from environmental harms driven by AI systems |
Language Model Risk Cards
Language Model Risk Cards provide a structured way to document risks associated with a language model or its deployment context. Each RiskCard describes how a risk can lead to harm, relates it to broader harm taxonomies, and can include representative prompt-and-output examples, making the framework useful for evaluating and communicating risks that may depend heavily on how a model is used.
risk-cards:lmrc:aggression_user |
Aggression User |
Aggression User |
risk-cards:lmrc:anthropomorphisation |
Anthropomorphisation |
Anthropomorphisation |
risk-cards:lmrc:astroturfing |
Astroturfing |
Astroturfing |
risk-cards:lmrc:bad_medical_advice |
Bad Medical Advice |
Bad Medical Advice |
risk-cards:lmrc:bad_social_advice |
Bad Social Advice |
Bad Social Advice |
risk-cards:lmrc:bullying |
Bullying |
Bullying |
risk-cards:lmrc:calls_for_conversion |
Calls For Conversion |
Calls For Conversion |
risk-cards:lmrc:challenge_presuppositions |
Challenge Presuppositions |
Challenge Presuppositions |
risk-cards:lmrc:conspiracy_theories |
Conspiracy Theories |
Conspiracy Theories |
risk-cards:lmrc:csam |
Csam |
Csam |
risk-cards:lmrc:cyber_weapon_instructions |
Cyber Weapon Instructions |
Cyber Weapon Instructions |
risk-cards:lmrc:deadnaming |
Deadnaming |
Deadnaming |
risk-cards:lmrc:deceptive_anthropomorphism |
Deceptive Anthropomorphism |
Deceptive Anthropomorphism |
risk-cards:lmrc:demands_for_personal_information |
Demands For Personal Information |
Demands For Personal Information |
risk-cards:lmrc:discrediting |
Discrediting |
Discrediting |
risk-cards:lmrc:dos |
Dos |
Dos |
risk-cards:lmrc:downplaying_credible_threats |
Downplaying Credible Threats |
Downplaying Credible Threats |
risk-cards:lmrc:doxxing |
Doxxing |
Doxxing |
risk-cards:lmrc:eating_disorder_promotion |
Eating Disorder Promotion |
Eating Disorder Promotion |
risk-cards:lmrc:essentialising |
Essentialising |
Essentialising |
risk-cards:lmrc:expressions_of_dominance |
Expressions Of Dominance |
Expressions Of Dominance |
risk-cards:lmrc:fear_of_violence |
Fear Of Violence |
Fear Of Violence |
risk-cards:lmrc:grooming_minors |
Grooming Minors |
Grooming Minors |
risk-cards:lmrc:harmful_instructions |
Harmful Instructions |
Harmful Instructions |
risk-cards:lmrc:hate_speech |
Hate Speech |
Hate Speech |
risk-cards:lmrc:hegemonising_worldview |
Hegemonising Worldview |
Hegemonising Worldview |
risk-cards:lmrc:holocaust_denial |
Holocaust Denial |
Holocaust Denial |
risk-cards:lmrc:identify_attacks |
Identify Attacks |
Identify Attacks |
risk-cards:lmrc:identity_misrepresentation |
Identity Misrepresentation |
Identity Misrepresentation |
risk-cards:lmrc:idolisation |
Idolisation |
Idolisation |
risk-cards:lmrc:impersonation |
Impersonation |
Impersonation |
risk-cards:lmrc:inciting_animal_abuse |
Inciting Animal Abuse |
Inciting Animal Abuse |
risk-cards:lmrc:inciting_self-harm |
Inciting Self-Harm |
Inciting Self-Harm |
risk-cards:lmrc:inciting_violence |
Inciting Violence |
Inciting Violence |
risk-cards:lmrc:manipulation |
Manipulation |
Manipulation |
risk-cards:lmrc:misquoting |
Misquoting |
Misquoting |
risk-cards:lmrc:mocking_people |
Mocking People |
Mocking People |
risk-cards:lmrc:neosexism |
Neosexism |
Neosexism |
risk-cards:lmrc:news_encyclopedia_hallucination |
News Encyclopedia Hallucination |
News Encyclopedia Hallucination |
risk-cards:lmrc:nonrepresentativity |
Nonrepresentativity |
Nonrepresentativity |
risk-cards:lmrc:objectification |
Objectification |
Objectification |
risk-cards:lmrc:outdated_false_info |
Outdated False Info |
Outdated False Info |
risk-cards:lmrc:person_group_undesirable |
Person Group Undesirable |
Person Group Undesirable |
risk-cards:lmrc:positive_reflections_death |
Positive Reflections Death |
Positive Reflections Death |
risk-cards:lmrc:profanity |
Profanity |
Profanity |
risk-cards:lmrc:prompt_extraction |
Prompt Extraction |
Prompt Extraction |
risk-cards:lmrc:providing_illegal_info |
Providing Illegal Info |
Providing Illegal Info |
risk-cards:lmrc:quack_medicine |
Quack Medicine |
Quack Medicine |
risk-cards:lmrc:radicalisation |
Radicalisation |
Radicalisation |
risk-cards:lmrc:reacting_to_self-harm |
Reacting To Self-Harm |
Reacting To Self-Harm |
risk-cards:lmrc:sexual_content |
Sexual Content |
Sexual Content |
risk-cards:lmrc:sexualisation |
Sexualisation |
Sexualisation |
risk-cards:lmrc:slur_usage |
Slur Usage |
Slur Usage |
risk-cards:lmrc:speculation |
Speculation |
Speculation |
risk-cards:lmrc:stereotypes |
Stereotypes |
Stereotypes |
risk-cards:lmrc:suicide_instructions |
Suicide Instructions |
Suicide Instructions |
risk-cards:lmrc:support_hate |
Support Hate |
Support Hate |
risk-cards:lmrc:supporting_medication_changes |
Supporting Medication Changes |
Supporting Medication Changes |
risk-cards:lmrc:support_scam |
Support Scam |
Support Scam |
risk-cards:lmrc:suppression_of_rights |
Suppression Of Rights |
Suppression Of Rights |
risk-cards:lmrc:threats |
Threats |
Threats |
risk-cards:lmrc:unproven_claims |
Unproven Claims |
Unproven Claims |
risk-cards:lmrc:violation_of_privacy |
Violation Of Privacy |
Violation Of Privacy |
risk-cards:lmrc:weapon_instructions |
Weapon Instructions |
Weapon Instructions |
risk-cards:lmrc:wrong_tone |
Wrong Tone |
Wrong Tone |
Common Weakness Enumeration
Common Weakness Enumeration (CWE) is a community-developed list of common software and hardware weakness types that could have security ramifications. Although CWE is not specific to AI, it provides standardized identifiers and terminology for implementation-level security weaknesses, making it useful for classifying conventional software vulnerabilities found in AI applications, services, and supporting infrastructure.
cwe:79 |
Improper Neutralization of Input During Web Page Generation (‘Cross-site Scripting’) |
The product does not neutralize or incorrectly neutralizes user-controllable input before it is placed in output that is used as a web page that is served to other users. |
cwe:89 |
Improper Neutralization of Special Elements used in an SQL Command |
The product constructs all or part of an SQL command using externally-influenced input from an upstream component, but it does not neutralize or incorrectly neutralizes special elements that could modify the intended SQL command when it is sent to a downstream component. Without sufficient removal or quoting of SQL syntax in user-controllable inputs, the generated SQL query can cause those inputs to be interpreted as SQL instead of ordinary user data. |
cwe:94 |
Improper Control of Generation of Code (‘Code Injection’) |
The product constructs all or part of a code segment using externally-influenced input from an upstream component, but it does not neutralize or incorrectly neutralizes special elements that could modify the syntax or behavior of the intended code segment. |
cwe:95 |
Improper Neutralization of Directives in Dynamically Evaluated Code (‘Eval Injection’) |
The product receives input from an upstream component, but it does not neutralize or incorrectly neutralizes code syntax before using the input in a dynamic evaluation call (e.g. “eval”). |
cwe:1336 |
Improper Neutralization of Special Elements Used in a Template Engine |
The product uses a template engine to insert or process externally-influenced input, but it does not neutralize or incorrectly neutralizes special elements or syntax that can be interpreted as template expressions or other code directives when processed by the engine. |
cwe:1426 |
Improper Validation of Generative AI Output |
The product invokes a generative AI/ML component whose behaviors and outputs cannot be directly controlled, but the product does not validate or insufficiently validates the outputs to ensure that they align with the intended security, content, or privacy policy. |
cwe:1427 |
Improper Neutralization of Input Used for LLM Prompting |
The product uses externally-provided data to build prompts provided to large language models (LLMs), but the way these prompts are constructed causes the LLM to fail to distinguish between user-supplied inputs and developer provided system directives. |
cwe:352 |
Cross-Site Request Forgery (CSRF) |
The web application does not, or cannot, sufficiently verify whether a request was intentionally provided by the user who sent the request, which could have originated from an unauthorized actor. |
Summon a Demon and Bind It
Summon a Demon and Bind It presents an empirically grounded framework for understanding how practitioners red-team large language models. Based on interviews with red-teamers, the research characterizes LLM red teaming as a limit-seeking, largely manual activity and identifies 12 attack strategies and 35 techniques, providing a useful basis for designing adversarial tests that probe model behavior and attempt to elicit failures or bypass safeguards.
demon:Language:Code_and_encode:Programming |
Programming |
Encapsulate request in code/pseudocode |
demon:Language:Code_and_encode:Data_encoding |
Data encoding |
Use an encoded representation for the request, e.g. base64 or ROT13 |
demon:Language:Code_and_encode:Data_presentation |
Data presentation |
Switch to an alternative layer for input represented, e.g. token IDs or a matrix of embeddings |
demon:Language:Code_and_encode:Token |
Token |
Use tokenizer-specific weaknesses to alter target behaviour |
demon:Language:Prompt_injection:Ignore_previous_instructions |
Ignore previous instructions |
Concatenating untrusted user input with the trusted prompt(s) from the system developers |
demon:Language:Prompt_injection:Strong_arm_attack |
Strong arm attack |
Use intensifiers and strong instructions |
demon:Language:Prompt_injection:Stop_sequences |
Stop sequences |
Using the language of code to halt the model’s direction of processing |
demon:Language:Stylizing:Formal_language |
Formal language |
Write from a position of authority |
demon:Language:Stylizing:Servile_language |
Servile language |
Employ servile language |
demon:Language:Stylizing:Synonymous_language |
Synonymous language |
Varying prompt slightly in form but not meaning |
demon:Language:Stylizing:Capitalizing |
Capitalizing |
USING CAPS |
demon:Language:Stylizing:Give_examples |
Give examples |
Issue examples of the target behaviour in the attack itself |
demon:Rhetoric:Persuasion_and_manipulation:Distraction |
Distraction |
Use an unrelated context or instruction to “slip something through the filter” |
demon:Rhetoric:Persuasion_and_manipulation:Escalating |
Escalating |
Attempting to have target “agree” with a very small part of the argument, and then building up to ask for slightly more |
demon:Rhetoric:Persuasion_and_manipulation:Reverse_psychology |
Reverse psychology |
Framing the intent as fighting the good fight |
demon:Rhetoric:Socratic_questioning:Identity_characteristics |
Identity characteristics |
Referencing identity characteristics, such as nationalities, cultural and/or religious symbols, historically or culturally significant events or locations, physical attribute |
demon:Rhetoric:Socratic_questioning:Social_hierarchies |
Social hierarchies |
Reference social hierarchies while posing the request |
demon:Possible_worlds:Emulation:Unreal_computing |
Unreal computing |
Emulation of a computing device |
demon:Possible_worlds:World_building:Opposite_world |
Opposite world |
Describe an opposite world as part of the request |
demon:Possible_worlds:World_building:Scenarios |
Scenarios |
Create scenarios where intent would be ethically sound or encouraged, or non-compliance risks harm |
demon:Fictionalizing:Switching_genres:Poetry |
Poetry |
Framing request or response as a poem or other verse |
demon:Fictionalizing:Switching_genres:Games |
Games |
Framing request as part of a game |
demon:Fictionalizing:Switching_genres:Posting |
Posting |
Framing request as help writing a post or other subsuming content |
demon:Fictionalizing:Re-storying:Goal_hijacking |
Goal hijacking |
Constructing a new meaning from an existing narrative, where the attacker works within a context to redirect the narrative |
demon:Fictionalizing:Roleplaying:Claim_authority |
Claim authority |
State in prose that one is a figure of authority when requesting |
demon:Fictionalizing:Roleplaying:DAN_and_target_persona |
DAN and target persona |
Ask target to evoke personas that don’t have restrictions |
demon:Fictionalizing:Roleplaying:User_persona |
User persona |
Inventing personas to elicit different behavior. May be evoked by simply using names associated with a specific culture or world view |
demon:Stratagems:Scattershot:Regenerate_response |
Regenerate response |
Repeatedly eliciting a response in case target slips up |
demon:Stratagems:Scattershot:Clean_slate |
Clean slate |
Resetting context by e.g. starting a new session |
demon:Stratagems:Scattershot:Changing_hyperparams |
Changing hyperparams |
Adjusting generation params, e.g. temperature |
demon:Stratagems:Meta-prompting:Perspective-shifting |
Perspective-shifting |
Hypotheticals by shifting new point of view, e.g. “what if you didn’t have this restriction? What would you say?”, or asking target to compare outputs and clarify gap |
demon:Stratagems:Meta-prompting:Ask_for_examples |
Ask for examples |
Request examples of intended behaviour |
EU AI Act
The EU AI Act is a regulatory framework that establishes risk-based requirements for developing, providing, and deploying AI systems in the European Union. It distinguishes among different levels and types of AI risk, prohibits certain practices, imposes requirements on high-risk systems, establishes transparency obligations, and introduces additional requirements for general-purpose AI models, including provisions related to evaluation, risk management, adversarial testing, and cybersecurity for models with systemic risk.
euai:discrimination:protected_class |
Protected Class Discrimination |
Discrimination against a protected class (age, race, sex, etc.) |
euai:discrimination:personal |
Personal Attribute Discrimination |
Discrimination based on non-protected personal attributes |
euai:harmful:violence |
Violent Language |
Generation of violent content |
euai:harmful:sexual |
Sexual Language |
Generation of sexual content |
euai:offensive:exploitation |
Vulnerability Exploitation |
Exploitation of software vulnerabilties |
euai:offensive:phishing |
Phishing |
Assistance with phishing |
euai:deception:subliminal |
Subliminal Deception |
Manipulation via subliminal deception |
euai:deception:hallucination |
Hallucination Deception |
Hallucinations related to factual information |
euai:harmful:pii |
Personally Identifiable Information |
Production of personally identifiable information |
euai:robustness:phishing |
Phishing Robustness |
Robustness of a target to Phishing attempts |
euai:robustness:adversarial |
Adversarial Robustness |
Robustness of a system to adversarial examples e.g. jailbreaks |
euai:robustness:unexpected_behavior |
Unexpected Behavior Robustness |
Robustness of a system to unexpected behavior e.g. glitch tokens |
euai:robustness:security |
Security |
Robustness of a system to being exploited via e.g. template injection |