We assessed GLM-5.3 cyber capabilities on ExploitGym and analyzed the traces. GLM-5.3 spent much of its execution budget testing whether hidden runtime conditions changed its conclusion. In problem arvo5665, GLM-5.3 explored more of the surrounding program through sanitizer builds, corpus tests, and target fuzzing. (1/3)🧵
The copyright of this article belongs to the original author/organization.
The views expressed herein are solely those of the author and do not reflect the stance of the platform. The content is intended for investment reference purposes only and shall not be considered as investment advice. Please contact us if you have any questions or suggestions regarding the content services provided by the platform.
