Self-calibrating SOC quality assurance with AI scoring and SOP improvement
Build a SOC analyst QA workflow. Use Tines Cases as the source of truth, Tines Records for storage, Anthropic for scoring, and Microsoft Graph for the SharePoint SOP library. Each week, sample a couple of recently closed triage cases per analyst, gather everything a reviewer would need: case notes, disposition, the detection rule behind it, and the SOP that should have been followed and have a Claude agent score the analyst's work on accuracy, process and documentation, with a verdict and a confidence level. Persist the scores, produce an Excel workbook and written report attached to a summary case, and escalate anything the AI shouldn't decide alone (rework verdicts, low confidence, unapproved high-severity cases) to a manager. When a manager disagrees with a score, let them record the correction on a form; feed those corrections back into the agent's rubric so it calibrates over time, and keep an audit record of the override. Periodically mine the accumulated scoring data for patterns that point at the SOPs rather than the analysts, and have an agent propose revised SOPs backed by the specific cases that justify each change. Publish proposals as drafts alongside the live SOPs, never overwrite one and route them through manager approval. Finish with a dashboard: quality trends and breakdowns by analyst and alert type, the cases currently awaiting review with the AI's reasoning and an inline override, the pending SOP proposals with their evidence, a banner when something needs attention, cross-filtering, and an Excel export.
What this prompt builds
A quality assurance workflow for security operations centers that automatically scores every closed case using AI against defined rubrics, surfaces only cases needing human review through a dashboard, captures manager overrides to calibrate the scoring agent over time, and mines patterns across cases to propose evidence-backed SOP improvements. Designed for SOC managers who need full QA coverage without manual review bottlenecks, replacing sparse sampling with continuous automated scoring that learns from human judgment. Delivers real-time quality signals, targeted oversight, and a feedback loop that improves both the scorer and the underlying procedures analysts follow.
The problem
Security operations teams that close hundreds of cases each week struggle to maintain quality assurance at scale. Traditional QA requires manual re-review of closed cases to verify analysts followed procedure, documented reasoning, and reached correct dispositions, but human bandwidth typically limits coverage to a small sample, leaving mistakes and compliance gaps undetected for long periods. Simply automating the scoring with AI introduces a new problem: without oversight, the scorer's own mistakes or blind spots go uncorrected, and managers lack visibility into which scored cases actually need attention or a way to turn corrections into lasting improvements. Beyond individual case scoring, there is no mechanism to surface patterns that point to flawed SOPs rather than analyst error, so process quality remains static even as QA data accumulates. Teams need full-coverage QA that stays calibrated to human judgment and feeds insights back into the procedures analysts follow.
Solution and impact
This workflow scores every closed SOC case automatically against accuracy, process adherence, and documentation rubrics using Claude, replacing sparse manual sampling with full coverage and immediate consistency. A dashboard surfaces only cases requiring human review (rework verdicts, low-confidence scores, unapproved critical cases) through review banners and badges, making oversight targeted rather than exhaustive. When managers override a score, the correction is captured with reasoning and fed back into the agent's rubric over time, creating a visible calibration trail that builds trust and improves scoring accuracy. The workflow also mines accumulated findings to propose evidence-backed SOP revisions routed through approval, extending the feedback loop beyond individual cases to the underlying procedures. SOC managers gain real-time quality signals across all cases, spend review time only where human judgment is needed, and drive continuous improvement in both the AI scorer and the SOPs it measures against.