Jailbreaks via in-context examples are a known vulnerability of language models, but the underlying mechanics have remained murky. Why does showing a model a few harmful exchanges cause it to comply with further harmful requests? This paper dissects the phenomenon carefully, mixing benign and harmful demonstrations to isolate what models actually extract. Surprisingly, benign demonstrations can either help or hurt safety depending on the model and training history. The findings have practical implications for red-teaming, model evaluation, and the design of safer few-shot prompting interfaces — especially in applications where users supply their own examples or system prompts.
Fler avsnitt av Eye on AI Weekly Research Watch
Visa alla avsnitt av Eye on AI Weekly Research WatchEye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
