Anthropic’s August 21 post says the risky surface is direct model access. Claude Security can now run Mythos 5 for Enterprise customers in public beta. A scan returns each finding with a CWE category, confidence and severity ratings, and a suggested fix. The suggested fix is applied in Claude Code on the web, using the models that organization has access to there. The scan does not extend Mythos to other surfaces. Partner products are supposed to do the same thing: run Mythos on a defined task and return one artifact, a patch list or an alert.

Anthropic says every patch must be reviewed and approved by a human before it is implemented, and that partner tools include abuse checks so the model stays on the assigned task. I would score the control on those two points: whether the wrapper can be steered into a different job, and whether a human reviews and approves the patch. A CWE plus a suggested fix is a ticket, not evidence that Mythos stayed inside the task.

source ↗

← all notes