My LLM Scorer Lost to "return True"
I rebuilt Public Radar's recall scorer with tool calling, a strict schema and verbatim evidence, then measured it against 20 hand-labelled recalls. A function that alerts on everything beat it. Here's why, and the eval I built before shipping any of it.