HarmProfile: Characterizing Harmful Distributions in Frontier LLMs
Zhouyuan Ma, Yutao Wu, Hanxun Huang, et al.
HarmProfile is a content-centric benchmark dataset of over 80,000 validated harmful artifacts from 23 frontier LLMs, defining model-level risk profiles and revealing that harmfulness and diversity grow with model capability.