Core benchmark methodologyPublication gate · incomplete

Metrics are useful only when the conditions travel with them.

The Core repository reports an internal evaluation, but Barrikade is not using those figures as homepage proof until the tested version, hardware, configuration, latency scope and reproduction path are documented.

Dataset compositionRecorded1,466 benign and 710 malicious promptsTested commit and releaseRequiredNot yet publishedHardware and operating systemRequiredNot yet publishedWarm-up and model-loading conditionsRequiredNot yet publishedThreshold configurationRequiredNot yet publishedLatency scope and percentilesRequiredNot yet publishedReproduction command and artifactsRequiredNot yet published

What is currently known

The repository describes a 2,176-prompt evaluation containing 1,466 benign and 710 malicious examples. It reports classification results and an average early-pipeline latency.

Those values are internal project results. They are not independent validation, production-customer outcomes, an end-to-end request latency guarantee or evidence that every detector layer ran for every example.

What must be published next

A reproducible methodology must pin the Core commit and artifact bundle; describe dataset provenance, deduplication and labeling; record hardware and software versions; list thresholds and optional layers; distinguish cold start from warmed execution; and report distributions rather than one average.

Accuracy must include a confusion matrix and confidence intervals. Latency must state which requests exited at each layer and whether API, serialization and model-loading time were included.

Publication rule

Until every required field above is complete, the site may describe the implemented architecture but must not use the repository's accuracy or latency values as sales proof.

Once complete, this page becomes the canonical methodology and every promoted metric links here with the evaluated version and “internal evaluation” label.