Jailbreak Prevention / Ngăn chặn Jailbreak

✍️
Từ điển và thuật ngữ trí tuệ nhân tạo

📖 Jailbreak Prevention

Definition (English):

Jailbreak prevention refers to techniques for detecting and blocking attempts to circumvent AI safety measures. "Jailbreaking" involves crafting prompts that trick LLMs into producing harmful, biased, or policy-violating outputs. Defensive strategies include input filtering, output monitoring, constitutional AI, system prompt hardening, and layered safety classifiers. This is an ongoing arms race between attackers and defenders.


📖 Ngăn chặn Jailbreak

Định nghĩa (Tiếng Việt):

Ngăn chặn jailbreak đề cập đến kỹ thuật phát hiện và chặn các attempts l繞过 biện pháp an toàn AI. "Jailbreaking" liên quan đến việc tạo prompt lừa LLM tạo đầu ra có hại, thiên vị hoặc vi phạm chính sách. Chiến lược phòng thủ bao gồm lọc đầu vào, giám sát đầu ra, constitutional AI, tăng cường system prompt và bộ phân loại an toàn nhiều lớp. Đây là cuộc chạy đua vũ trang liên tục giữa tấn công và phòng thủ.


📂 Phân loại: Safety+

HỆ SINH THÁI CiCC

Đi tiếp cùng nội dung này

Hỏi AI Edu?