Evidence and trust
The fields that say how far a Search result or an eval grade can be trusted, what each one means when it is null, and a policy for acting on them in code.
Every Search result and every eval grade carries fields that say how much it can be trusted. They are the part of a response your code should branch on before it uses the rest. This page collects them in one place; the Search and Evals API references carry the full detail.
Three rules hold across all of them:
- Null is not zero. A field that could not be measured comes back
nullor with anunavailablestatus, never as0. A zero is a measurement. - A refusal is not a fail. When a grade could not be made honestly, the run says
not_applicableand gives a reason. Nothing was assessed, so there is no score to read into. - Quotes are retrieved, never written. Every customer quote in a response is a real utterance from your corpus. The model cites ids; the server swaps them back to the stored text and drops any id it does not recognise.
On a Search result
| Field | What it tells you | Act on it |
|---|---|---|
mode_ran | Which of the four lanes answered. Lanes read different stores at different grains. | Compare two results only when they ran on the same lane. |
corpus.store, corpus.grain | The physical store (warehouse, pgvector, theme_index) and what one row counts. | A theme is a cluster, not an utterance. Do not quote it as one person's words. |
corpus.empty_reason | Why an empty result is empty. | read_failed is an error: retry. no_match means nothing was said. weak_matches_only means rephrase. |
corpus.speaker_scope | Whose speech the store can hold: all, or external_only (customer side only). | Do not report an internal-vs-external split off an external_only store. |
corpus.coverage.basis | Whether the time-range numbers describe the whole match set, a truncated slice, or nothing (unavailable). | Read this before any date or recency claim. |
corpus.truncated | A cap cut the result. | Do not treat the rows as the population. |
compiled.sql | The query as you asked for it, before the gate added your tenant filter. | Keep it as a record of intent. |
detail.internal.status | On the plain-language lane, whether the SQL ran (ok) or not (failed). | failed during a burst can be the warehouse-read budget. Back off before rewording. |
On an eval grade
| Field | What it tells you | Act on it |
|---|---|---|
headline.submitted.checks (passed / total) | How many rubric lines your own copy cleared. The only number on the report about your writing. | Report this fraction. |
gate.passed | The send/hold bit for the copy you submitted. boolean or null. | false holds. null abstains: pass through, do not block. |
not_applicable_reason | Why a run graded nothing (empty_corpus, not_outreach, evidence_unavailable, …). overall_score is null beside it. | Drop refused runs from averages. Branch on the reason. |
Quote tier | Whose voice a quote is: account (one named company), segment (a cohort), corpus (a recurring pattern), external (the public web). | A claim may not reach past its tier. "You told us" needs account. |
corpus_only | The judge had nothing but corpus patterns to cite. | Any account-specific claim in the graded copy was reaching past its evidence. |
confidence.level | high, moderate or low, with reasons (thin evidence, an errored leg, a truncated prompt). | Hold a moderate or low score loosely. It never changes the score. |
lift_reportable | Whether the before/after difference is larger than the judge's own run-to-run spread. | Do not quote lift when this is false. |
usage on each artifact | as_provided, simulated_specimen, reusable_prompt, illustration_only. | Never send an illustration_only message. Keep the reusable_prompt. |
evidence_scope, candidate_scope | Whether the evidence and the graded rewrite were pinned from another run. | Attribute a difference between two runs only when both say pinned. |
What a grade is, and is not
A grade is a structured second opinion on writing, grounded in cited quotes. Each rubric line is a binary pass or fail, so on a five-line rubric the headline score takes exactly six values: 1, 1.8, 2.6, 3.4, 4.2 and 5. It does not predict reply rate or conversion.
One run is one draw. Grading the same text twice can move the score by about two rubric lines, which is why comparing two runs withholds a delta that falls inside that spread. To measure a small change, grade each version several times and compare medians. The current bars and spreads are on Models and limits.
A policy for acting on a result
A starting point for code that acts without a person in the loop. Tighten it for anything with real cost.
| Situation | Do |
|---|---|
Search: empty_reason is read_failed, or detail.internal.status is failed | Retry with backoff. Do not report "nothing found". |
Search: coverage.basis is unavailable, or truncated is true | Use the rows, but do not state a population, a date range or a recency. |
Search: speaker_scope is external_only | Quote it as customer voice only. |
Eval: gate.passed is true | Proceed. |
Eval: gate.passed is false | Hold. Regenerate with the improved prompt from a full run, then grade again. |
Eval: gate is null or gate.passed is null | Abstain and pass through. Surface not_applicable_reason to a person. |
Eval: confidence.level is low, or corpus_only is true | Proceed only with a person reviewing the copy. |
The patterns that use these fields end to end are on Patterns.