Coordinated AI agents could turn a small security failure into a much larger problem. Understanding the risk starts with separating reported test behavior from worst-case predictions.