Let AI read the scope. Never let it grade the work.
Large language models are very good at one thing that verification needs badly: reading messy human language and turning it into structure. A scope of work says "migrate approximately 85 users to Business Premium and enforce MFA for all active users." A model can turn that into three requirements with counts, in seconds, and show you the sentence each one came from.
Models are very bad at another thing verification needs: being right every time about facts they cannot see. Ask a model whether EDR was deployed and it will read the ticket, notice the words "deployment completed," and say yes with confidence. That confidence is the danger.
The split
Draw a hard line. AI interprets: it proposes requirements, fills in expected values, and explains results in plain language. Software verifies: deterministic code compares expected with observed and returns a result that is the same every time for the same inputs. The two never share a function.
Why the line holds
Because deterministic code can be tested. Every control gets a test for passed, failed, warning and unknown. A model cannot be tested that way; it can only be sampled. When a client asks "how do you know?", the answer has to be "here is the rule and here is the evidence," not "the model was fairly sure."
We caught this in our own build. The model once filled a license name as "Microsoft 365 Business Premium" when the tenant reports "Business Premium." Exact comparison would have failed 85 users. The design caught it because the model is only allowed to fill blanks, and the blank is validated against the values the real system uses.