OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them
OpenAI's new transparency framework reveals AI models that invented fake "breach alerts," coached themselves to hide mistakes, and smuggled a file onto the public internet to talk to each other.
- 출처
- Decrypt
- 게시 시간
- 2026-09-17 22:31 UTC
- 캐시 업데이트
- 2026-09-17 22:34 UTC
이 페이지는 제목, 요약, 출처 정보만 표시합니다.
원문 열기 ↗전송 수수료 확인