The repository describes a 2,176-prompt evaluation containing 1,466 benign and 710 malicious examples. It reports classification results and an average early-pipeline latency.
Those values are internal project results. They are not independent validation, production-customer outcomes, an end-to-end request latency guarantee or evidence that every detector layer ran for every example.