Claude Mythos has been verified by tests as a leader in vulnerability detection, yet it also exhibits other shortcomings

Claude Mythos has been verified by tests as a leader in vulnerability detection, yet it also exhibits other shortcomings

58 hardware

Brief Summary

XBOW conducted an independent assessment of Anthropic’s Mythos Preview AI model. The results showed that the model outperforms existing counterparts in source‑code vulnerability detection and when working with native code and reverse engineering, but it shows weaknesses in isolated code analysis and confirming the practical applicability of discovered exploits. The cost of using the model remains a question: Mythos is more expensive than Opus, yet with a limited token budget it sometimes demonstrates better accuracy.

1. What XBOW tested
- Scope of tests – a series of independent experiments on Mythos Preview.
- Scenarios – auditing operational systems with source‑code access, isolated code analysis, reverse engineering, and interaction with graphical interfaces.

2. Key findings
Metric | Mythos Preview | Opposite Models
---|---|---
Vulnerability detection | Best among all models, especially in source code and native programming | Less accurate but sometimes more reliable when checking specific cases
False‑positive filtering | Removes more “false” findings than predecessors | Often produces more false alarms
Isolated code problems | Performs worse without system context | Some models handle line‑by‑line analysis better
Exploit confirmation | Tends to be literal, sometimes overestimates practical value of findings | More critical of real applicability
Reverse engineering and native code | Delivers strong results; “understands” program logic without source code | Less accurate in these tasks
UI interaction | Not always precise at locating element coordinates but successfully identifies required actions in a browser | Some models are more accurate in positioning

3. Cost and efficiency
- Pricing – Anthropic stated that Mythos will cost five times more than Opus.
- Economic analysis – XBOW ran experiments with “cheaper” models, giving them more runtime. Results showed that when normalized by cost, Mythos Preview does not appear wasteful for high‑accuracy tasks.
- Benchmarks – On a fixed token budget, Mythos outperforms Opus 4.6 in web vulnerability search but falls behind GPT5.5.

4. Conclusion
Mythos Preview demonstrates outstanding power in source‑code auditing and native programs, as well as reverse engineering and web‑application analysis tasks. However, the model sometimes underestimates the real applicability of discovered vulnerabilities and shows weaknesses in isolated code analysis. With limited token resources Mythos may be more cost‑effective than Opus, but with a full budget GPT5.5 remains competitive.

XBOW’s takeaway: Mythos Preview is a reliable tool for identifying potential vulnerabilities in source code and complex systems, but it requires additional confirmation of the practical value of discovered exploits.

Comments (0)

Share your thoughts — please be polite and stay on topic.

No comments yet. Leave a comment — share your opinion!

To leave a comment, please log in.

Log in to comment