TAGGING EVALUATION SAMPLE Purpose: inspect tagging suggestions against a human-reviewed reference set. Open tagging-evaluation.csv in a spreadsheet. All 20 records are fictional. Keep identifiers as text. Each row is one document's decision for the tag ACCESS (physical access and accommodations; excludes account passwords), with binary 1=yes and 0=no reference and suggested values. The sample is a calculation demonstration, not a vendor benchmark. Count outcomes: Reference=1, suggested=1: true positive (TP), 6. Reference=0, suggested=1: false positive (FP), 4. Reference=1, suggested=0: false negative (FN), 2. Reference=0, suggested=0: true negative (TN), 8. Total=20; reference positives=8; suggested positives=10. Precision=TP/(TP+FP)=6/10=60%. Recall=TP/(TP+FN)=6/8=75%. If either denominator is zero, report that measure as undefined with its counts. Inspect missed and inappropriate tags; overall accuracy can hide rare-tag errors. Vocabulary worksheet (one record per term): Stable code / preferred label: Alternative labels / definition: Included example / excluded counterexample: Related or broader concepts: Owner / vocabulary version / change reason: Review sample / tool version / threshold / evaluation date: To adapt: 1. Define each term, inclusion/exclusion boundaries, synonyms and examples. 2. Build an independently reviewed reference set representative of your content. 3. Record tool/version/configuration, date and threshold. 4. Keep a separate development set when tuning; assess on held-out documents. 5. Count errors per tag, including rare and consequential terms. 6. Have an accountable person review suggestions before consequential publication. 7. Repeat when vocabulary, tool or content changes. Do not combine records across different tags without retaining per-tag results. Source guide: https://yenra.com/meta-tag-creating-programs/ Prepared October 4, 2026.