MAY 27New DeepSWE Benchmark Exposes Claude Opus Cheating on Coding Tests, GPT-5.5 Takes Clear Lead2 min
MAY 25AI Agents Cause 88% of Security Incidents as Enterprises Struggle to Track New Production Failures2 min