A native-English tutor audited the system against his own evaluations of three Korean learners. The transcription came back 98–100% identical to his. The analysis did not: he found roughly twice as many issues, and on one sample he flagged eight where the system flagged one. The misses clustered in three patterns that Korean speakers hit constantly — articles, prepositions, and plurals.
The cause was in the prompt, and it was deliberate. It instructed the model to omit anything borderline and to silently discard weakly supported findings. The system had been tuned for precision, and it had overpaid: a tool that only reports what it is certain of tells an advanced learner almost nothing.
The fix was to stop asking one pass to do two jobs. A first pass now sweeps for recall, with explicit examples of the three categories that matter. A second verifies the candidates and restores precision, narrowing overlapping findings to the smallest accurate span instead of dropping them. Precision became a stage that could be tuned without suppressing recall — and the gap between them became something we could measure rather than argue about.