대형 언어 모델에서의 정렬 위장
우리 대부분은 누군가가 우리의 견해 또는 가치를 공유하는 것처럼 보이지만, 실제로는 단지 그렇게 하는 척만 하는 상황을 경험한 적이 있을 것입니다—우리가 “alignment faking”이라고 부를 수 있는 그런 행동입니다.
Comments
More Videos
View all
A Complete Guide to Claude Code - Here are ALL the Best Strategies
Here is everything you need to know to crush building anything with Claude Code! These are the strategies used by top engineers to code reliably and unbelievably fast with AI coding assistants.

Mastering Claude Code in 30 minutes
Learn advanced features, shortcuts, and workflows to get the most from Claude Code

Claude Code in Slack
Delegate tasks to Claude Code directly from Slack, making it easy to move context from Slack conversations to coding sessions. Claude Code in Slack is available now for teams with the Claude app installed in their Slack workspace and who have access to Claude Code on the web.
claudePrompting for Agents | Code w/ Claude
Presented at Code w/ Claude by @anthropic-ai on May 22, 2025 in San Francisco, CA, USA.
claudeAlignment faking in large language models
Most of us have encountered situations where someone appears to share our views or values, but is in fact only pretending to do so—a behavior that we might call “alignment faking”.
claudeBuilding more effective AI agents
Anthropic’s Alex Albert (Claude Relations) sits down with Erik Schluntz (Multi-Agent Research and co-author of our blog post, Building Effective Agents) for a discussion on the evolution of agents over the past six months, including tips for building multi-agent systems, common multi-agent patterns, and best practices for using skills, MCP servers, and tools.