Better than what you do now, by enough to matter
The question is never "can AI do this". It is whether it does it better than what you do now, by enough to justify building and running it.
So we measure the current process first. Then we build the smallest thing that beats it, put it where the work actually happens rather than in a separate tool nobody opens, and monitor it afterwards — because accuracy at launch tells you nothing about accuracy in a year.