Live on PyPIOpen SourceLLM EvalsPython

neverempty: when "no results" is a lie

An open-source eval harness that measures how often a tool-calling agent tells a user there is no data when a tool actually failed.

0.1.0
Published on PyPI
2,158
Tests
1
Core dependency (pydantic)

pendingThe headline metric, misreport_as_empty, has not been measured on a production agent yet. The machinery is proven on test agents. The production measurement is the next step.

help_outline

The Problem

A tool queried a database for one quarter's numbers. The query timed out. The tool caught the exception and returned an empty list, which is exactly what it returns when the quarter genuinely has no records. The model read the empty list and wrote a confident answer: there is no data for Q3. There was data. Nothing in the stack logged an error anyone would see, and the user had no reason to doubt it. Two sentences that sound alike and mean the opposite: → "No results" means there is nothing there. → "The search failed" means I don't know what is there. An agent that confuses them is not unhelpful. It is wrong, confidently, in the one direction a user cannot check.
search

Why It Keeps Happening

Nobody writes this bug on purpose. The type signature writes it for you. A function that returns a list has no vocabulary for failure. Its only word for the unhappy path is `[]`, the same word it uses for "I succeeded and found nothing". The caller writes `if results:` and both paths merge into one. I read the tool-handling code of three public agent frameworks and found the same shape in each: an exception caught and turned into an empty result. Three is not "most agents", and I don't claim it is. It is enough to show this is not one team's mistake.
architecture

What I Built

neverempty is a Python eval harness, published on PyPI. It fixes the bug in two places, then measures whether the fix works. → The type. A tool wrapped with `@tool` returns `Ok`, `Empty` or `Err`, a tagged union. A plain value becomes `Empty` only if the author says what empty means. There is no falsiness inference, because `0`, `False` and `""` are real values. In strict mode, a tool that returns `[]` without declaring what empty means raises instead of guessing. → The text the model reads. A perfectly typed error that serialises to `[]` in the tool message reproduces the bug exactly. So a failure renders as an explicit note: the tool failed, you do not know whether matching data exists, do not say none exists. → The measurement. Faults are injected on purpose and every answer is scored. It claimed absence (misreport), it reported the failure (reported), or it dropped it (ignored). The misreport share is the headline number, with a Wilson interval. Core depends on pydantic and nothing else. The statistics are pure Python. The harness makes no model calls of its own, so a full example eval runs with no API key.
science

Measuring a Failure That Rarely Happens

Real timeouts almost never occur in a test run. Run 300 cases, see zero failures, and the misreport rate is 0 out of 0. Publishing that as 0% would be the exact bug this library exists to prevent: a measurement that never happened, reported as a clean result. So the harness forces the failure. A fault is declared per case, the wrapper raises it before the real tool runs, and the trace records that it was injected. No external system is touched. The same rule runs through the whole codebase: → A metric with nothing to measure prints "not measured", never 0. → An unknown cost is null with a reason, never $0. → A scorer that cannot decide returns not-applicable, never a fail. → Below n=10 the renderer refuses to print a percentage. Safety metrics collapse across repeats by any-hit: one misreport in three runs makes the case a hit. Capability metrics use majority. The CI gate fails a build on a one-sided exact McNemar test against a committed baseline, a breached floor, or a failed must-pass case. Not on a raw threshold that flaps with model noise.
bug_report

Two Bugs I Almost Shipped

Wiring the library into NextRole, my own 7-agent app, turned up two bugs in my instrumentation. Both looked correct on the page. Both were caught only by tests. 1. A zero-results search went silent. `Empty` carries no value by design, so the wrapper threw away the sentence the agent returns next to its empty list. The user would have seen a blank reply. The fix stores that payload where the predicate last sees it, in a ContextVar rather than a global, because the endpoint serves requests concurrently and a global would surface one user's results in another user's reply. Three interleaved tasks, zero cross-talk. 2. An injected fault returned the previous call's payload. A fault short-circuits the function, so the predicate never runs and the stored payload is stale. The injected failure would have looked like a successful empty search. That is this library's own headline bug, reintroduced inside its own instrumentation, corrupting the exact number it exists to produce. It now raises, and each agent's existing error handling takes over.
fact_check

What Is True, and What Is Not Yet

This split is the point of the project, so here it is plainly. Verified: • Installs from PyPI and runs a complete example eval with no API key. • 2,158 tests. mypy --strict and ruff clean. Python 3.10 to 3.13. • The misreport metric works end to end on agents built to test it: an honest agent scores 0.0, a lying one scores 1.0. • First real run against NextRole's live router: 91.2% routing accuracy over 330 hand-verified cases, 95% CI 87.7% to 93.8%. That is a routing number. It says the agent picks the right branch. It says nothing about whether the agent is honest when a tool fails. Not yet true: • misreport_as_empty has never been measured on a production agent. The fault cases for NextRole are next, and this page will say so until they are done. • The wrapper fixes what the model is told. It cannot stop the model writing "no results" in its own prose. That gap is exactly what the measurement is for. • The score is a pattern match over 13 absence and 17 failure patterns in English, Hindi and Hinglish. An absence claim in wording the patterns miss scores as ignored, so the published rate will be a lower bound.
lightbulb

Key Learnings

1. Fix the type and the text together. A typed error that still renders as `[]` to the model fixes nothing the user sees. 2. Inject the failure. A failure mode that rarely happens naturally cannot be measured by waiting for it. 3. "Not measured" and "zero" are different results. A system that prints one for the other lies in the same way the agent did. 4. Tests catch what reading misses. Both instrumentation bugs passed a careful read. 5. State the gap. A number published with its limits is worth more than one that hides them.