EVALS // RED-TEAM
Precision and recall measured against Tamga's internal red-team set. The corpus is updated monthly with open-source datasets.
Full corpus + per-category table + scan-latency percentiles are published as JSON and Markdown in the repository. Anyone can rerun the same command and diff the numbers.
| Category | Samples | Precision | Recall | F1 |
|---|---|---|---|---|
| PII — Turkish national ID | 42 | 0.980 | 0.950 | 0.960 |
| PII — Credit card (Luhn + BIN) | 38 | 0.990 | 0.920 | 0.950 |
| PII — IBAN (TR) | 24 | 1.000 | 0.960 | 0.980 |
| PII — Email | 31 | 0.970 | 0.970 | 0.970 |
| Jailbreak — override | 18 | 0.940 | 0.940 | 0.940 |
| Jailbreak — many-shot | 12 | 1.000 | 0.830 | 0.910 |
| Jailbreak — base64/hex obfuscation | 14 | 0.930 | 0.860 | 0.890 |
| Jailbreak — Turkish role hijacking | 16 | 0.940 | 0.880 | 0.910 |
| Secret — AWS / OpenAI / GitHub | 22 | 1.000 | 1.000 | 1.000 |
| Indirect — Canary token | 10 | 1.000 | 1.000 | 1.000 |