Sveriges mest populära poddar
Eye on AI Weekly Research Watch

What Do Safety-Aligned LLMs Learn From Mixed Compliance Demonstrations?

2 min23 juni 2026
Jailbreaks via in-context examples are a known vulnerability of language models, but the underlying mechanics have remained murky. Why does showing a model a few harmful exchanges cause it to comply with further harmful requests? This paper dissects the phenomenon carefully, mixing benign and harmful demonstrations to isolate what models actually extract. Surprisingly, benign demonstrations can either help or hurt safety depending on the model and training history. The findings have practical implications for red-teaming, model evaluation, and the design of safer few-shot prompting interfaces — especially in applications where users supply their own examples or system prompts.

Fler avsnitt av Eye on AI Weekly Research Watch

Visa alla avsnitt av Eye on AI Weekly Research Watch

Eye on AI Weekly Research Watch med Craig Spencer Smith finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.