
from xerxes-agents22
Surgically remove refusal behaviors from open-weight LLMs using mechanistic interpretability and weight projection.
OBLITERATUS allows an agent to remove guardrails and refusal behaviors from open-weight LLMs without retraining. It employs advanced techniques like SVD, diff-in-means, and LEACE concept erasure to identify and excise refusal directions from model weights while preserving reasoning capabilities.
Trigger this skill when a user wants to "uncensor" or "abliterate" a model, create a version of Llama/Qwen/Mistral without refusals, or analyze the mechanistic workings of a model's refusal system.
Integrated into the Xerxes Agent framework; compatible with any agent capable of executing shell commands on a GPU-enabled host.
This skill has not been reviewed by our automated audit pipeline yet.