A database of offensive ML TTP’s, broken down by supply chain attacks, offensive ML techniques, and adversarial ML. The playbook aims to simplify the decision-making process of targeting ML in an organization.
LLM security is the investigation of the failure modes of LLMs in use, the conditions that lead to them, and their mitigations.
Here are links to large language model security content - research, papers, and news - posted by @llm_sec.
A curation of awesome tools, documents, and projects about LLM Security.
Official code for "Baseline Defenses for Adversarial Attacks Against Aligned Language Models."
Some notable 2024 blogs from the Embrace the Red blog:
- LLM Context Pollution and Delayed Automated Tool Invocation
- GitHub Copilot Chat: Prompt Injection & Data Exfiltration
- WhoAmI: Conditional Prompt Injection Instructions
- Machine Learning Attack Series: Keras Backdoor Model
This GitHub repository provides a curated list of papers and resources on backdoor attacks and defenses in deep learning. It includes information on backdoor attacks in various contexts, such as image classification and language models, and links to relevant papers and code.
Awesome-Backdoor-in-Deep-Learning
This repository offers a collection of resources related to offensive AI, including tools for performing backdoor attacks. It includes frameworks like ART, Cleverhans, and TextAttack, which support various attack types and data formats.
This GitHub repository lists multiple resources related to backdoor learning, including papers, tools, and datasets. It covers various types of backdoor attacks and defenses, providing a comprehensive overview of the field.
This repository contains code for implementing backdoor attacks on chat models. It includes methods for training and deploying chat models with distributed trigger-based backdoor attacks, which are designed to be triggered by specific user inputs across different conversation rounds.
Chat-Models-Backdoor-Attacking
This repository focuses on creating models with imperceptible triggers using adversarial perturbations. It includes a detailed methodology for implementing backdoor attacks using the CIFAR-10 dataset and the ResNet-18 model.
These resources should provide you with a good starting point for understanding and implementing backdoor attacks in AI models.
Red Teaming AI - GitHub Repository
Rigging
https://github.com/dreadnode/rigging
Marque
https://github.com/dreadnode/marque
Parley
https://github.com/dreadnode/Parley
Research
https://github.com/dreadnode/research
Counterfit
https://github.com/Azure/counterfit
Proof Pudding
https://github.com/moohax/Proof-Pudding
Koppeling
https://github.com/monoxgas/Koppeling
sRDI
https://github.com/monoxgas/sRDI
Deep Drop
https://github.com/moohax/Deep-Drop
Charcuterie
https://github.com/moohax/Charcuterie
Minibus
https://github.com/monoxgas/minibus
Offensive Machine Learning - Apres Con (Slides)
https://github.com/dreadnode/conferences/blob/main/ApesCyber_2024/workshop/Offensive%20Machine%20Learning.pdf
Offensive Machine Learning - Apres Con (Notebooks)
https://github.com/dreadnode/conferences/tree/main/ApesCyber_2024/workshop/notebooks
Ghosts on the Node (Slides)
https://github.com/dreadnode/conferences/blob/main/SOCON_2024/Ghosts%20on%20the%20Node.pdf
Zen and the Art of Adversarial Machine Learning (Slides)
https://github.com/moohax/Talks/blob/master/slides/Blackhat_EU_21.pdf
Zen and the Art of Adversarial Machine Learning (Talk)
https://www.youtube.com/watch?v=tEBwMGCKEso
Screendoors on Battleships (Slides)
https://github.com/moohax/Talks/blob/master/slides/Screen%20Doors%20on%20Battleships.pdf
Counterfit: Attacking Machine Learning in Blackbox Settings (Slides)
https://github.com/moohax/Talks/blob/master/slides/Counterfit_BH_Arsenal_21.pdf
It Is The Year 2000, We Are Robots (Slides)
https://github.com/moohax/Talks/blob/master/slides/bsides_slc_20.pdf
Flying A False Flag (Slides)
https://github.com/monoxgas/FlyingAFalseFlag/blob/master/nick_landers_bhusa_19_flying_a_false_flag.pdf
42: The Answer to Life the Universe, and Everything Offensive Security (Slides)
https://github.com/moohax/Talks/blob/master/slides/DerbyCon19.pdf
Scheming With Machines (Slides)
https://github.com/moohax/Talks/blob/master/slides/Scheming_with_Machines_BSidesLV_19.pdf
Poisoning Web-Scale Training Datasets is Practical
https://arxiv.org/abs/2302.10149
Sandbox Classification Using Decision Trees and Artificial Neural Networks
https://link.springer.com/chapter/10.1007/978-3-030-52249-0_18