Relevance classification
Identify whether a sentence contains actionable cyber threat intelligence or background / non-actionable material.
binary classificationSMAH-CTI is a human-annotated benchmark built from 100 public CTI reports, unifying relevance classification, CTI-specific named entity recognition, MITRE ATT&CK mapping, and grounded procedure descriptions in one sentence-level resource.
Each instance is anchored to a sentence from a source report, with immediately preceding context where available. Relevant sentences are annotated across complementary semantic layers so models can be evaluated independently or as a multi-task pipeline.
Identify whether a sentence contains actionable cyber threat intelligence or background / non-actionable material.
binary classificationExtract Action, Infrastructure Indicator, Malware/Tool, and Threat Actor spans with exact character offsets.
span-level NERAssign one or more MITRE ATT&CK tactics and techniques/sub-techniques to sentences that express adversarial behavior.
multi-label mappingGenerate a concise, annotator-written procedure description grounded in the specific adversarial action expressed by the sentence.
grounded generationThe entity offsets are zero-based and end-exclusive relative to sentence_text. Every one of the 15,487 released spans was validated against the exact source substring.
Threat groups, actor names, or individuals associated with malicious activity.
Adversarial or security-relevant actions expressed in the sentence.
Infrastructure, indicators, paths, domains, IPs, and other technical artifacts.
Malware families, tools, utilities, and exploit kits.
The released partition is report-level: entire reports belong to exactly one of Train, Dev, or Test. No report identifier or sentence UID overlaps across the three splits.
Sentence instances from five public CTI collections; the bars show the share of all 12,304 instances.
Report-level partition used by the associated paper.
| Split | Reports | Sentences | Relevant |
|---|---|---|---|
| Train | 70 | 8,635 | 3,006 |
| Dev | 10 | 1,238 | 434 |
| Test | 20 | 2,431 | 862 |
| Total | 100 | 12,304 | 4,302 |
The explorer is bundled directly into this page, so it works even when index.html is opened as a local file. Search sentence text and filter by source, relevance, and entity type. Nothing is sent to a server.
12,304 records are bundled with this page
The public release includes the dataset itself, explicit report splits, a JSON Schema, provenance records, integrity validation, licenses, and the paper.
12,304 annotated sentence instances in a single JSON release.
JSON SCHEMAMachine-readable field definitions, allowed values, and NER offset contract.
JSON · SPLITSExact Train, Dev, and Test report IDs used for benchmark evaluation.
XLSX · PROVENANCEReport-level acquisition details, original source URLs, and access dates.
JSON · QAIntegrity checks, UID uniqueness, distribution counts, and span validation.
PDF · PAPERMethods, annotation design, benchmark formulation, and experimental results.
If you use the benchmark, please cite the associated EMNLP 2026 Findings paper. The repository also includes CITATION.cff for machine-readable citation metadata.