Docs

Prompt and Message Eval

Paste an email and the prompt behind it. Get back a grade against what your buyers actually said, the specific lines to change, and a reusable prompt your whole team can run - with a worked example of a real cold email graded end to end

Slug: prompt-and-message-eval. The eval Amdahl ships. It answers the question every rep asks and no tool has ever answered honestly: is this email any good, and what specifically would make it better?

Not "is it well written". Not "does it match our tone". Does it say something these buyers actually care about, and can we prove it? The grade is scored against verbatim quotes pulled from your own calls and emails, so the feedback is your customers' words, not a model's taste.

Needs no setup. It is the default, so a run with no eval param uses it.


Start here: a real one, graded

Here is a cold email a rep sent last week. It is not a strawman — it is the shape of most first-touch outbound.

text
Subject: Quick question

Hi Dana - saw Northwind is scaling fast. We help engineering teams
ship 10x faster with best-in-class rollout tooling. Worth a quick
15 minutes next week?

And the prompt behind it:

text
Write a short cold email to a VP of Engineering about our rollout tooling.
Keep it under 100 words and friendly.

Run both through the eval and this is the shape of what comes back. Rewrite is the default — it produces the replacement prompt and the specimen message this walkthrough shows, so the "mode" line is optional. Send "mode": "advisory" instead when your prompt is a document you are keeping; it grades the same way but returns anchored suggestions against what you already have — see advisory mode.

What the report opens with
failpass

Your email argued speed. These buyers said the problem is approvals.

Your draft scored 1.6 and did not clear the bar. The improved version scored 4.4 against the same quotes, under the same blind rubric.

fail to pass

It leads on the approval bottleneck the cohort actually named, attributes the claim to the cohort it came from, drops the unverifiable 10x, and asks for a teardown instead of a call.

One run. Re-run a few times and read the median.
Scored blind3 real quotesGraded against your whole corpus

1. Your message, graded — with the reason for every number

DimensionVerdictWhy
Relevant positioningfailNothing here is about Northwind. "Scaling fast" applies to every company in the segment; the rest is a product description.
GroundingfailNo claim maps to anything these buyers said. The retrieved quotes talk about release approvals, not speed.
Verified specificsfail"10x faster" is not supported by any evidence in the corpus, and it is the kind of number a buyer will ask you to defend.
Differentiationfail"Best-in-class rollout tooling" is a category, not a position. Three of your competitors say it.
CTA claritypassThe ask is clear and low-friction. It is the strongest part of the email.

Headline: 1.8/5 — one of five checks passed.

Every rubric line is a binary pass/fail, so the headline is not a holistic guess and not an average of five opinions. It is the pass fraction placed on a five-point axis:

code
score_15 = 1 + 4 × (passed / total)

Five lines means the dial takes exactly six values — 1, 1.8, 2.6, 3.4, 4.2, 5 — and nothing in between. 1.8 is exactly one line passed. The response ships checks_passed and checks_total next to score_15 for that reason: "1 of 5 checks" says what the number measured; "1.8" on its own reads like a rating on a scale that does not exist. Quote the fraction.

Here are the same five as the console renders them, each row judged on both sides and expandable to the reasoning behind its verdict:

The dimensions, expandable
Relevant positioningYours1/5Improved5/5 pass
Your draftNothing here is about Northwind. "Scaling fast" applies to every company in the segment; the rest is a product description.
ImprovedIt leads on approvals, which is what these buyers said the problem was.
GroundingYours1/5Improved5/5 pass
Your draftNo claim maps to anything these buyers said. The retrieved quotes talk about release approvals, not speed.
ImprovedIt uses their exact language, attributed to the cohort the evidence actually came from.
Verified specificsYours1/5Improved4/5 pass
Your draft"10x faster" is not supported by any evidence in the corpus, and it is the kind of number a buyer will ask you to defend.
ImprovedThe one number in it is a number Amdahl can point at.
DifferentiationYours2/5Improved4/5 pass
Your draft"Best-in-class rollout tooling" is a category, not a position. Three of your competitors say it.
ImprovedThe audit trail as the product is a position, not a category claim.
CTA clarityYours3/5Improved4/5 pass
Your draftThe ask is clear and low-friction. It is the strongest part of the email.
ImprovedAn artifact instead of a meeting is lower friction for a cold account.

2. The quotes it graded you against

These are verbatim, pulled from your own conversations, tagged with whether they support or contradict what your email claimed:

Contradicts"Speed honestly isn't our problem. We can ship in a day. The problem is that four people have to sign off and one of them is always on PTO." — Rollout approval friction

Contradicts"We already tried the fast-deploy pitch internally. It didn't move anyone. What moved people was showing the audit trail." — Prior attempts

Supports"Every release we do by hand costs us most of a Thursday." — Manual release cost

That third quote is the one worth building on. The first two are why Verified specifics failed — your buyers said, in their own words, that speed is not the problem, so "10x faster" has nothing behind it.

The evidence, verbatim
3shown below
Speed honestly isn't our problem. We can ship in a day. The problem is that four people have to sign off and one of them is always on PTO.
Customer · Executive · Negotiation · 12 May 2026Across your customers
We already tried the fast-deploy pitch internally. It didn't move anyone. What moved people was showing the audit trail.
CustomerAcross your customers
Every release we do by hand costs us most of a Thursday.
Across your customers

These are cohort quotes, not Northwind quotes. Every quote carries a tier saying whose voice it is, and here all three are corpus: this run passed no audience, so retrieval searched your whole corpus for the themes your message is about. On a cold account it can do nothing narrower — Northwind has never spoken to you. What comes back is what buyers like Dana say, which is exactly what makes it usable on an account you have no history with.

Pass an audience and the run additionally draws that cohort's own words, tagged segment — a tighter licence than corpus and the right one for a first touch. See Evidence tiers.

That also sets a hard line on how you may write it. "The eng leaders we talk to say…" is supported. "Your engineers said…" is not. Naming the wrong scope turns a well-grounded line into a claim of a conversation that never happened — the most common way a good email becomes a false one, and the eval grades it as a grounding failure rather than a style note.

Pass an account and, when that company IS in your data, their own buyer-side words come back tagged tier: "account" — the one tier that licenses "you told us". See Evidence tiers.

3. The improved prompt — the thing you actually keep

This is the durable output. Not the email. The prompt.

text
You are writing a first-touch email to one specific account. Before you write a word:

1. RESEARCH, in this order. Take everything each tier gives you.
   a. THE COHORT - always available. In Amdahl, find how people in this role
      and this industry talk about the problem area, and note their exact
      language. This is the tier that works on an account nobody has spoken to.
   b. THE TRIGGER - sometimes. Something that just happened to them: funding,
      a hire, a launch, a move. Lead with it when it exists. Never invent one.
   c. THE ACCOUNT - only when they are already in our data. Their own words and
      their own history outrank the cohort. Most first-touch accounts will not
      have this, and that is fine - it is a bonus tier, not a gate.
   If the COHORT comes back empty, say so and stop; do not write from the
   category.

2. ATTRIBUTE at the scope you actually have. "The eng leaders we talk to say"
   is supported by cohort evidence. "Your engineers said" is not, unless tier
   (c) turned up that exact conversation. Never imply a prior interaction that
   did not happen.

3. POSITION. Frame the offer against the problem they named, in their words.
   If the stated problem is approvals and yours is a speed product, lead with
   approvals or do not send. Never open with how fast we are unless a buyer
   said slowness was costing them something.

4. VERIFY. Every number, outcome, or capability claim must be checkable against
   a real call or a real customer. If you cannot point at the evidence, CUT the
   claim. Do not soften it - a hedged unverifiable claim is still unverifiable.

5. DISQUALIFY. If this prospect is not who we sell to - wrong role, wrong
   segment, or a problem we do not solve - stop and say why. If the cohort has
   already tried and rejected this framing, pick a different one.

6. ASK. One next step, sized to the relationship. A 15-minute call is fine for
   a warm account and too much for a cold one - offer to send the thing instead.

Keep it under 120 words. Plain sentences. No category adjectives
("best-in-class", "industry-leading", "revolutionary").

Notice what it is: a template with slots, not a rewrite of this one email. Paste it for the next account tomorrow and you get an equally specific result. That is what it is graded on.

4. The example output — and what it is not

This is not an email to send. It is a specimen produced so the score difference could be measured — evidence that the improved prompt works. On the wire it is typed usage: "illustration_only", and every Amdahl surface labels it "Example output — what the improved prompt produces". Send it verbatim and you are sending a message written by a model that has never met Dana.

text
Subject: The four sign-offs

Hi Dana - the eng leaders we talk to keep saying shipping isn't the
bottleneck, approvals are: four sign-offs, and someone's always out.

We built the approval path for exactly that - the audit trail is the
product, not a side effect. One customer went from four serial
approvals to two parallel ones.

Want me to send the two-page teardown of how they did it? No call needed.

Graded the same way, by the same judge, against the same quotes: 5/5 — all five checks passed. Relevant positioning went fail → pass (it leads on approvals, which is what these buyers said), grounding fail → pass (it uses their exact language, attributed to the cohort it came from), verified specifics fail → pass (the one number is one Amdahl can point at), differentiation fail → pass, and CTA clarity was the single line that already passed and still does (an artifact instead of a meeting is lower friction for a cold account).

Note what it does not say: "one of your engineers told us." Nobody at Northwind has spoken to you. The line is specific and true at the same time because it is attributed to the cohort the evidence actually came from.

Both sides, as you would read them in the console. The reusable prompt is the artifact worth keeping; the message is labelled an illustration everywhere it appears:

Yours and improved, side by side
Improved version4.4of 5example -- what the improved prompt produces, not a message to send

Subject: The four sign-offs Hi Dana - the eng leaders we talk to keep saying shipping isn't the bottleneck, approvals are: four sign-offs, and someone's always out. We built the approval path for exactly that - the audit trail is the product, not a side effect. One customer went from four serial approvals to two parallel ones. Want me to send the two-page teardown of how they did it? No call needed.

An example, written to show the difference. Regenerate from the improved prompt for anything you actually send.

Improved prompt4.6of 5improved -- reusable

You are writing a first-touch email to one specific account. Before you write a word: 1. RESEARCH, in this order. Take everything each tier gives you. a. THE COHORT - always available. In Amdahl, find how people in this role and this industry talk about the problem area, and note their exact language. This is the tier that works on an account nobody has spoken to. b. THE TRIGGER - sometimes. Something that just happened to them: funding, a hire, a launch, a move. Lead with it when it exists. Never invent one. c. THE ACCOUNT - only when they are already in our data. Their own words and their own history outrank the cohort. Most first-touch accounts will not have this, and that is fine - it is a bonus tier, not a gate. If the COHORT comes back empty, say so and stop; do not write from the category. 2. ATTRIBUTE at the scope you actually have. "The eng leaders we talk to say" is supported by cohort evidence. "Your engineers said" is not, unless tier (c) turned up that exact conversation. Never imply a prior interaction that did not happen. 3. POSITION. Frame the offer against the problem they named, in their words. If the stated problem is approvals and yours is a speed product, lead with approvals or do not send. Never open with how fast we are unless a buyer said slowness was costing them something. 4. VERIFY. Every number, outcome, or capability claim must be checkable against a real call or a real customer. If you cannot point at the evidence, CUT the claim. Do not soften it - a hedged unverifiable claim is still unverifiable. 5. DISQUALIFY. If this prospect is not who we sell to - wrong role, wrong segment, or a problem we do not solve - stop and say why. If the cohort has already tried and rejected this framing, pick a different one. 6. ASK. One next step, sized to the relationship. A 15-minute call is fine for a warm account and too much for a cold one - offer to send the thing instead. Keep it under 120 words. Plain sentences. No category adjectives ("best-in-class", "industry-leading", "revolutionary").

A template with slots, not a rewrite of this one email. Run it against the next account and you get an equally specific result.

5. The lift, and the transition

code
Your message      1.8/5   1 of 5 checks    fail
Improved          5/5     5 of 5 checks    pass     fail -> pass

Your message failing is the finding, not an error. It is the thing you ran the eval to learn. The eval reports the two verdicts separately (input_verdict and improved_verdict) precisely so a working run never reports "fail" because the draft it was asked to diagnose was the bad one.

6. What changes when the account is warm

Everything above is the cold case: Dana has never spoken to you, so tier (c) is empty and the evidence is her cohort's. When the account is in your data, exactly one thing changes:

  • Changes — tier (c) fires. Their own words and their own history outrank the cohort, and you can attribute directly to them, because now you actually can.
  • Does not change — verify, disqualify, and attribute. A warm account is not a licence for a claim you cannot point at.

You never tell the eval which case you are in. It grades what you wrote against what your buyers said either way. The tiers exist so the prompt still produces something specific when tier (c) is empty — which, for first-touch outbound, is most of the time.


Advisory: when your prompt is a living document

Plenty of teams do not have a one-line prompt. They have a living rules document — 40 pages of positioning guidance, objection handling, and things legal will not let you say. Nobody is going to replace that because an eval suggested a new one. advisory is the mode for exactly that case.

mode: "advisory" leaves your document alone. Instead of a rewrite you get anchored, surgical suggestions:

json
{
  "kind": "add",
  "facet": "prompt",
  "anchor_quote": "Always mention our SOC 2 certification in the first email.",
  "anchor_offset": 8214,
  "section_id": "s6",
  "title": "Gate the SOC 2 line on the buyer actually asking",
  "detail": "Change this to: mention SOC 2 only when the account has raised security, compliance, or procurement. Otherwise cut it.",
  "why": "Across the retrieved quotes, security comes up from procurement and never from the engineering buyer this sequence targets. Leading with it reads as boilerplate to the person receiving it.",
  "dimension": "Relevant positioning",
  "quotes": [
    {
      "text": "Honestly the SOC 2 stuff is procurement's problem, not mine.",
      "stance": "contradicts"
    }
  ]
}

Five kinds of suggestion come back:

KindWhat it means
keepThis is working. Do not lose it in the next edit. A report that is all criticism does not tell you what to protect.
addA missing instruction, stated so you can paste it in
strengthenA vague instruction that is not doing its job yet
removeAn instruction that is actively hurting the output
reorderThe content is right, the sequence is the problem

anchor_quote is verified server-side. It is only ever a literal substring of your document — if the model paraphrases, the anchor is dropped and you get a document-level suggestion instead. A suggestion can never quote you a line you did not write. anchor_offset is the verified position, which is what makes an anchor findable when the same phrase repeats.

Long documents are sectioned, never quietly truncated

Past 12,000 characters your prompt is split on your own headings and packed head-and-tail (the opening frames the task; the closing usually carries the hard rules). Every run reports coverage:

json
"coverage": {
  "total_chars": 48219,
  "graded_chars": 23904,
  "truncated": true,
  "strategy": "sectioned",
  "sections": [
    { "id": "s1", "heading": "Who we sell to", "chars": 1840, "included": true },
    { "id": "s2", "heading": "Positioning by segment", "chars": 6210, "included": true },
    { "id": "s7", "heading": "Legacy objection scripts", "chars": 9022, "included": false }
  ]
}

You always know what was read. Past 250,000 characters the run refuses rather than returning a confident number computed over a sliver of your document.


Who uses this, and how

Reps: before you send a sequence

Paste step 1 and the prompt you used. You get the specific lines that are not landing and the quotes that prove it. Keep the improved prompt; reuse it per account. Two minutes, and you stop sending the "10x faster" email to a team whose problem is approvals.

Managers: audit the sequence, not the rep

Run every step of a live sequence through it. The pattern across the grades is the coaching: if grounding fails on every step, the problem is the template, not the person sending it. The suggestions tell you exactly which lines to change.

Enablement: keep the rules doc honest

Run your positioning document in advisory mode on a cadence. As what customers say changes, the suggestions change with it - and keep tells you which parts are still earning their place.

Ops: wire it into the review step

It is one API call and a poll. Gate sequence publishing on a passing improved verdict, or just log the transition so you can see whether outbound quality is moving.

A sequence review, end to end

Run every step

One evals.run per sequence step, each with the step's copy as message and the shared sequence prompt as prompt. They are independent - fire them all and poll.

Read the dimensions, not the headline

A headline of 2.6/5 tells you nothing on its own. Grounding failing on all five steps tells you the sequence was written from the category, not from customers.

Take the prompt, not the emails

The improved prompt is the artifact. Put it in the sequence brief so the next person to write a step starts from it.

Fix the anchored lines

Every remove and strengthen suggestion points at a specific line with the quote that justifies it. Work the list.

Re-run and watch the transition

fail -> pass on the improved side means the prompt is doing its job. fail -> fail with an explanation means the evidence does not support a stronger version - which is usually telling you something real about the segment.


How it works under the hood

Five stages. Four of them exist to make the number trustworthy rather than merely produced:

  1. Read the intent. One cheap pass extracts what you are trying to do — the offer, the ask, and each specific claim your draft makes. Retrieval is seeded from that, never from your draft. (Seeding on a weak draft returns generic themes, and those generic themes then become the evidence the draft is held to. That is circular.)
  2. Gather evidence. Several retrieval passes run at once — one per intent seed, one per specific claim — so "cuts onboarding from six weeks to two" is checked against the themes that speak to that, not against the average of your whole message. The set is then frozen for the entire run — every candidate, including a revision round, is judged against the same quotes. Across two separate runs it is not frozen: retrieval is seeded from what you submitted, so editing your draft also changes what gets searched for. That is what evidence_from_run is for.
  3. Write. A model produces the improvement, the suggestions, and the examples. It scores nothing.
  4. Grade, blind. A separate pass scores both candidates in one call, unlabelled, with the presentation order derived from a content hash. The judge does not know which one it wrote, so it cannot be kind to itself.
  5. Revise, if needed. If the improved version misses the bar (4.2/5), it gets exactly one more attempt, guided by the judge's own critique. On a full run that same 4.2 is the bar BOTH sides are read against: one threshold decides the submitted verdict and the improved one. The eval's other bar, pass_threshold 3.5, is what mode: "gate" compares against, and on this rubric the two select the same drafts anyway — five binary lines put score_15 on six values (1, 1.8, 2.6, 3.4, 4.2, 5), so nothing lands between 3.5 and 4.2. Then it stops and tells you honestly why it could not do better.

A quote in a verdict is always a real utterance. The model may only cite quote ids it was handed; the server swaps those ids back to verbatim text and drops any it does not recognise. Inventing a customer saying something is structurally impossible, not merely discouraged.

Honest failure

Nothing to grade against? The run comes back not_applicable — never a false fail. An empty corpus means there is nothing to ground the message in, and saying so is more useful than a fabricated score.

Not the kind of writing this reads? Also not_applicable, under not_applicable_reason: "not_outreach". The rubric scores commercial writing at any stage — prospecting, an open deal, a proposal or signature, a pilot, a renewal, a re-engagement — but not correspondence with no persuasive job (scheduling, receipts, support replies, internal notes), even inside an active deal. Your work was not assessed, so there is no score to read into it. Always branch on not_applicable_reason rather than assuming an empty corpus; the reason table says what each one asks of you.

The improvement did not clear the bar? You get the score, the transition, and an explanation naming what still holds it back. No silent rounding up.


Inputs

FieldRequiredNotes
promptThe instruction, brief, or rules document behind the message. Up to 250,000 characters; sectioned past 12,000. Send this alone and it writes a specimen draft to grade.
messageThe drafted message (email, LinkedIn note, whatever). Send this alone and it also suggests a reusable prompt. The retired key outbound_message is still accepted.
audienceWho it is going to. Name a seniority (executive, manager, individual contributor) or a title (VP of Sales, Head of L&D) and the run scopes to that cohort — see Audience scoping.
accountThe company you are writing to. When they are in your data, their own words are retrieved and tagged account, which is the only evidence that licenses an account-specific claim — see Evidence tiers. Abstains honestly when they are not.
moderewriterewrite (the default), advisory, or gate. rewrite returns a full improved prompt and message. advisory leaves your writing alone and returns anchored suggestions[] against it — the right shape when your prompt is a large document you are not going to replace. gate grades only what you sent and stops.
evidence_from_runA previous run id. Grades this submission against that run's quotes instead of retrieving its own — which is what makes two runs comparable. See Comparing two versions.
scopeTyped filters naming the slice of your corpus to grade against — role_level, objection raised, deal outcome, qualification band. The general form of audience, and it sits beside inputs rather than inside it. See Filter scoping.

At least one of prompt / message must be supplied. Send both when you have both — grading the pair is what produces a useful improved prompt, because the eval can see what the prompt asked for and what it actually got.

Run it

bash
curl -X POST https://app.amdahl.ai/api/platform/v1/evals/run \
  -H "X-API-Key: $AMDAHL_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "eval": "prompt-and-message-eval",
    "inputs": {
      "prompt": "Write a short cold email to a VP of Engineering about our rollout tooling. Keep it under 100 words and friendly.",
      "message": "Hi Dana - saw Northwind is scaling fast. We help engineering teams ship 10x faster with best-in-class rollout tooling. Worth a quick 15 minutes next week?",
      "audience": "VP Engineering at a 500-person SaaS company",
      "mode": "rewrite"
    }
  }'

mode is the one field worth naming deliberately. Omit it and you get rewrite — the replacement prompt and specimen message the walkthrough above shows. Name "mode": "advisory" when you want anchored suggestions against a document you are keeping, or "mode": "gate" for a pass/fail on your own copy with no rewrite at all.

The response is a handle. Poll it until the run settles:

bash
curl https://app.amdahl.ai/api/platform/v1/eval-runs/$RUN_ID \
  -H "X-API-Key: $AMDAHL_KEY"

Comparing two versions

The obvious next move after reading a report is to edit the draft and re-grade it. Do that naively and you have changed two things at once.

Retrieval is seeded from what you submitted — that is what makes the evidence relevant to your claims rather than to the average of your workspace. The consequence is that a different message searches for different quotes. So the second run is scored against a different set, and the score difference mixes your edit with the evidence move.

This is not a rounding effect. On a real pair we measured — the same cold email, edited once — the two runs shared none of their four retrieval queries and only 5 of 12 recorded quotes. The second scored worse. None of that difference was attributable to the edit.

Pin the evidence

Grade v1 normally, then grade v2 against v1's quotes:

bash
curl -X POST https://app.amdahl.ai/api/platform/v1/evals/run \
  -H "X-API-Key: $AMDAHL_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "inputs": { "message": "<your edited draft>" },
    "evidence_from_run": "'"$V1_RUN_ID"'"
  }'

The second run retrieves nothing and reuses the first run's frozen pool, so the only thing that changed is your copy. The response echoes the pin and its vintage:

json
{
  "data": {
    "run_id": "…",
    "reused": false,
    "status": "queued",
    "evidence_pin": { "from_run_id": "…", "frozen_at": "2026-07-31T22:06:13.154Z", "quotes": 16 }
  }
}

The source run must be in your workspace and must have reached retrieval. A run that cannot supply evidence is refused rather than silently graded on fresh quotes — a dropped pin would return a well-formed report that answered a different question than the one you asked.

Check a comparison you already ran

bash
curl https://app.amdahl.ai/api/platform/v1/eval-runs/$V1_RUN_ID/compare/$V2_RUN_ID \
  -H "X-API-Key: $AMDAHL_KEY"
json
{
  "data": {
    "comparison": {
      "a": { "run_id": "…", "transition": "fail_to_pass", "grounded": { "cited": 3, "pool": 16 } },
      "b": { "run_id": "…", "transition": "fail_to_fail", "grounded": { "cited": 1, "pool": 16 } },
      "evidence_overlap": {
        "shared": 5,
        "only_a": 7,
        "only_b": 7,
        "jaccard": 0.263,
        "measured": true
      },
      "delta_attributable": false,
      "caveat": "These two runs were graded against DIFFERENT customer evidence — 5 shared quotes, 7 only in the first and 7 only in the second (26% overlap). A score difference here mixes your edit with a different set of quotes, so it is not a measurement of the edit. No delta is reported.",
      "remedy": "Re-run the second version with evidence_from_run set to … "
    }
  }
}

When delta_attributable is false there is no score_delta in the response at all. That is deliberate rather than cautious: a number next to a warning still gets read, quoted, and pasted into a decision, and the warning does not travel with it.

evidence_overlap.measured is false when at least one run did not record the evidence it was graded against — a run from before this shipped, or one that refused before retrieval. That means the overlap is unknown, not zero, and delta_attributable stays false: a comparison that cannot demonstrate its own controls is not a weak finding, it is not a finding.

Audience scoping

If you pass an audience, the run does two things with it before grading anything.

First it resolves it to a seniority — executive, manager, or individual contributor. Titles resolve directly ("VP of Sales", "Head of L&D", "SDRs", "Sr. Mgr., Talent"), and so do the bare words. Something that names a function, a vertical, or a company size ("marketing", "healthcare", "mid-market logistics") names no seniority, and the run says so rather than guessing one.

Then it checks your data for that cohort. The report is only scoped to an audience once your workspace has enough recorded conversations with them to be worth grading positioning against:

FloorWhy
3+ distinct peopleOne person is an anecdote.
25+ utterancesThree people who each said one line in passing is not a voice.
2+ distinct companiesWithout this, one talkative account clears the other two on its own and its vocabulary gets reported as an audience-wide finding.

All three must pass. If they do not, the run still completes and still grades your prompt and message — it just grades them against your whole customer corpus, and the report says which of five things happened:

What the report saysWhat it means
not_providedYou did not pass an audience. Nothing to do.
unresolvableWe could not match your text to a seniority. Naming a level scopes it.
no_evidenceWe understood the cohort; your workspace has no recorded conversations with them yet.
thin_evidenceSome conversations, below the floors. The counts are in the report so you can see how close.
lookup_failedThe check itself failed on our side. This is not a statement about your data.

A scoped run carries the counts it decided on, so you can see the evidence behind the framing rather than taking it on trust.

And a cleared cohort changes the evidence, not just the framing. The run draws that cohort's own utterances and tags them segment — the tier that licenses "teams like yours". The draw is spread across companies, no more than a couple from any one of them, because the distinct-companies floor above is about the cohort and says nothing about which quotes come back: a cohort can clear it and still have its most recent conversations come from one talkative customer, and a cohort-wide licence over one company's words is the same over-claim one rung down. segment_companies on the report says how many the draw actually spanned.

Until eval version 2.6.0 this did not happen — audience reached the judge as an instruction to assume a cohort while retrieval stayed corpus-wide, so a scoped run and an unscoped one graded the same pool.

Filter scoping

audience cuts on one axis — seniority — and takes free text to do it. scope is the general form: typed predicates over the same field vocabulary Search advertises, so "grade this against what people at enterprise closed-won accounts actually said" is expressible without a bespoke input for every axis someone might want.

It sits beside inputs, not inside it — same level as reuse and evidence_from_run:

bash
curl -X POST https://app.amdahl.ai/api/platform/v1/evals/run \
  -H "X-API-Key: $AMDAHL_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "eval": "prompt-and-message-eval",
    "inputs": { "message": "<your draft>" },
    "scope": {
      "audience": "customer_voice",
      "filters": [
        { "field": "role_level", "op": "eq", "value": "executive" },
        { "field": "pushback_type", "op": "in", "value": ["pricing", "build_vs_buy"] },
        { "surface": "deals", "field": "deal_stage_status", "op": "eq", "value": "won" }
      ]
    }
  }'
FieldTypeDefaultWhat it does
filtersarray[]Typed predicates, ANDed. Each is { surface?, field, op, value? }. At most 25 per run.
audiencestringallcustomer_voice splices the workspace's canonical customer-voice predicate — a disjunction the ANDed filter grammar cannot state.
allow_thin_evidencebooleanfalseGrade a slice that falls below the evidence floors anyway, with its thinness reported rather than hidden.

A filter is the same shape Search's filter lane takes, plus a surface per filter. op is one of eq · neq · in · not_in · contains · gt · gte · lt · lte · between · is_null · not_null — but which operators a given field admits is a property of that field, so read the catalog rather than assuming.

The three surfaces

surface defaults to interactions. Naming it on a filter is how you predicate on a fact that lives at a different grain than the quotes do:

SurfaceGrainFilter on it to say
interactionsone row per utteranceSomething about the speech itself or the person who spoke it — seniority, channel, objection raised, sentiment. This is the surface the quotes come from.
dealsone row per dealSomething about the CRM opportunity the conversation hangs off — stage, outcome, amount, close date.
deal_qualificationone row per companySomething about how well an account is qualified — MEDDPICC/SPICED coverage, close-likelihood scoring, the binding constraint.

Mixing surfaces in one scope is the point, and it is the one deliberate difference from search.query, which takes a single surface per call. "Executives at closed-won accounts" is one utterance-grain cut AND one deal-grain cut; forcing a choice would make the most useful scope inexpressible.

Deal-grain filters cannot slice utterances directly, so the run resolves them to the accounts they match and scopes the quote draw to those accounts. That resolution is bounded: a deal-grain filter matching 500 accounts or more abstains with too_many_accounts rather than carrying a clipped set and calling the result "enterprise closed-won accounts" — the grade would be real and the label would be a lie. The bound is on the read, not on the count, so exactly 500 is on the abstaining side of it.

Discover the field names — never guess

Field names, types, admitted operators and (for coded fields) the exact spellings live in the Search field catalog, derived from the same schema the eval's filter compiler validates against, so the two cannot drift:

bash
curl "https://app.amdahl.ai/api/platform/v1/search/fields" \
  -H "X-API-Key: $AMDAHL_KEY"

Over MCP it is the search tool's fields action ({ "action": "fields" }), or the search_field://list resource. Full walkthrough, including why sample_values matters before you filter a coded field, in Search → Step 0.

department is not filterable, and no spelling of it will work. It is the axis a GTM team names first, and it does not exist as a column on the warehouse view the eval reads — it lives on a contacts table that is not one of the three surfaces above, so search cannot filter on it either. Cut on role_level (ic · manager · executive · unknown) instead. That is not a substitution made for convenience: hand-labelled against production titles, the role_level mapper scores 92.3% and the department mapper 69.5%, and the failure modes are asymmetric — role_level misses abstain to unknown, which nothing grades against, while department misses pooled into a real, high-volume bucket that would contaminate anything cut by it.

The evidence floors

A resolved scope is checked against your corpus before it is used, against the same three floors audience applies — imported from it rather than restated, so the two can never disagree about what "enough evidence" means:

FloorWhy
3+ distinct external peopleOne person is an anecdote.
25+ utterancesThree people who each said one line in passing is not a voice.
2+ distinct companiesWithout this, one talkative account clears the other two by itself and its vocabulary gets reported as a slice-wide finding.

All three must pass. The floors exist because of what an eval returns that a search does not: a search hands you four rows and you judge them; an eval hands you a grade with the slice's name stapled to it. The same four rows become "this is what enterprise champions say", asserted with a score behind it. The grade is real and the framing is a lie, and the framing is the part people quote.

What is refused before the run starts, and what is not

Two SHAPE problems come back as invalid_argument from the run call itself, before a run id is minted. They are refusals rather than in-run abstains because at dispatch you can still be told and can still fix it:

  • A key we do not recognise. scope and each filter are strict: operator for op, column for field, a stray values is refused, not ignored. The usual default is to strip an unknown key, and a stripped filter widens the slice you thought you had asked for while returning a perfectly well-formed report about it.
  • More than 25 filters, for the same reason: narrowing is still available to you at this point.

A scope that narrows nothing is accepted, not refused. {}, { "filters": [] }, { "audience": "all" } and a bare { "allow_thin_evidence": true } — an override with nothing to override — all pass validation, and all mean the same thing: the run is graded unscoped and the scope outcome is not_provided, whose line ("No filters were supplied…") is true of every one of them. Nothing is hidden by accepting them, so refusing would buy no honesty, and it would break a UI that sends an empty array the moment someone clears the filter panel.

Field and operator SEMANTICS are not checked at dispatch either. A misspelled field, an operator the field does not admit, an unknown surface value — the compiler that owns the field catalog sees those inside the run, and they surface as the invalid abstain below rather than as a dispatch-time error. The dispatch gate checks the shape; the run checks the vocabulary.

The one thing that must never happen is a scope that is silently DROPPED or trimmed — that is the single case where an honest-sounding report would be describing a wider slice than you asked for. Being loose about an empty filter array costs nothing; being loose about an unrecognised key would cost exactly that.

When a scope is not used, the run still completes

A scope is an enrichment of the report, never a precondition for it. If it cannot be applied the run still grades your prompt and message, and names which of seven things happened:

ReasonWhat it meansWhat to do
not_providedYou passed no filters. Nothing went wrong.Nothing.
invalidThe filters did not compile — unknown field, operator, or surface.Re-read the field catalog for the surface you filtered.
no_evidenceThe filters are valid; your workspace has no conversations in that cut.Widen the cut, or accept the unscoped grade.
thin_evidenceSome conversations, below one or more floors.Widen the cut, or pass allow_thin_evidence to grade it anyway.
lookup_failedThe evidence check itself failed. Not a statement about your data.Retry. If it persists, tell us.
no_matching_accountsDeal-grain filters matched no accounts, so the slice is empty.Loosen the deal-level filters.
too_many_accountsDeal-grain filters matched 500 accounts or more.Narrow the deal-level filters and re-run.

What it falls back TO is not always your whole corpus. With no other scoping input, it is: the segment tier draws nothing and the run is graded corpus-wide. But inputs.audience is gated separately and independently, so if you sent one and it cleared its own floors, an abstained scope leaves THAT cohort in place — the filters are gone, the run is not unscoped. Read the two together before quoting what a run spoke for.

lookup_failed and no_evidence are kept apart on purpose. Reporting an infrastructure failure as "you have no conversations matching that" is a lie about the customer's data, and it sends someone off to check an integration that is working fine.

Where the outcome shows up

The decision is made once, before any grading, and lands on the run's progress trailprogress.steps[] on eval_run://<id> (GET /eval-runs/<id>), as one scope_resolved step. It is emitted on every run, including one that passed no scope at all, so "this run stood on no filter slice" is a recorded fact rather than an absence you have to infer:

FieldOnValue
statusalwaysresolved or abstained. Branch on this before reading anything else.
labelalwaysThe human line — the slice and its counts on the resolved path, the reason and its consequence sentence on the abstained one.
sliceresolvedThe name of the cut actually drawn from. Composite when filters and an inputs.audience both resolved (see below).
evidenceresolved, and on a thin_evidence / no_evidence abstain{ distinct_speakers, utterances, distinct_companies } — the counts behind the decision.
below_floorsresolvedtrue when allow_thin_evidence carried a thin slice through.
reasonabstainedOne of the seven above.
messageabstainedThe written line for that reason, stating the consequence as well as the cause. This is the one to show a person.
detailabstained, when there is oneThe actionable specifics — the filter compiler's own complaint on invalid, the failure on lookup_failed.
requested_sliceabstained, whenever a scope was suppliedWhat you asked to be graded against, whether or not the gate honoured it — so an abstain still tells you which cut was requested.
provenanceresolvedThe value-free rendering of the cut (surface, field, operator, value counts) that is stamped on every quote and shown to the judge. Your filter values never appear here.

Branch on reason; show a reader message; read detail when you need to fix the filters.

This step is the only place the scope outcome surfaces today. It is not on the verdict or the report card, where audience already is — so a report on a run whose filters were turned down reads the same as a report on a run that never asked for a slice. If you scope runs, read the progress trail; do not infer the slice from a score that came back lower than you expected.

Grading a deliberately narrow slice

allow_thin_evidence: true is the escape hatch for the case where the narrow cut is the question — a new segment with four recorded calls, and you want to know how the message reads against those four.

The run then proceeds scoped, and the scope_resolved step carries below_floors: true beside the real evidence counts. Honest by default, overridable on purpose, never silent either way: what the flag buys you is a thin grade that is labelled as one, not a thin grade that reads like a cohort finding. (Labelled on the progress trail — see the caveat above about the report card.)

A scope changes the evidence, not just the framing

Same rule as a cleared audience: the run draws that slice's own utterances and tags them segment — the tier that licenses "teams like yours". The draw is spread across companies, no more than a couple from any one of them, because the distinct-companies floor is about the slice and says nothing about which quotes come back.

Filters and inputs.audience COMPOSE — they do not race. When both resolve, the slice is their INTERSECTION: the cohort predicate is ANDed into the filters before the floors are measured, so the counts describe the same cut the quotes are drawn from, and the slice name carries both halves. Naming only the filters would describe a wider slice than the one that was graded — the same class of lie as widening, pointed the other way.

The composition is why an inputs.audience that resolved is still standing when a scope abstains: it was gated on its own path, so the run falls back to that cohort rather than to your whole corpus.

Scope is part of the run's identity

scope is a named dimension of the run fingerprint, so a filtered request and an unfiltered one are different runs under the default reuse: "cached". Without that they would be identical in content and the second would be served the first one's verdict — a corpus-wide grade returned under a slice's name, which is the same lie the abstain rules exist to prevent, arriving by a different door.

Pinning evidence also compares the slice — but on a different rendering of it, and the difference is worth knowing before you script an A/B. The fingerprint uses a canonical key: the filters are sorted, so re-ordering them does not fork the address. The pin check compares the DISPLAY LABEL — the line the report shows, rendered in the order you listed the filters in.

The consequence: grade v2 against v1's quotes under a contradicting scope and the run fails rather than grading, which is the point — a pinned pool drawn from one slice cannot honestly answer a question asked about another. But grade against the SAME filters listed in a different order and it fails too, on a pin that was in fact compatible. That is the safe direction to be wrong in — you re-run and pay for a fresh draw, rather than one slice's words being reported under another slice's name — so if you are pinning across runs, build the filter array once and reuse it rather than rebuilding it per call.

Runnable calls in the report

research_steps are real Amdahl calls you can paste to close the evidence gaps the report found. Every one is validated before it ships — an operation that does not exist, one that would mutate your workspace, or SQL our query gate would refuse is dropped rather than shown.

They are also bounded by what your key can actually run: a step you would get a 403 for is not a useful suggestion, so the report only proposes calls covered by your own scopes. The tool_kit field carries the arithmetic — callable is how many operations your key could run, out_of_scope how many exist beyond it. When out_of_scope is 0 there is nothing to caveat; when it is not, the report says how many further calls exist without listing them.

Advisory mode, on a document you are keeping

rewrite is the default, so advisory has to be named explicitly:

bash
curl -X POST https://app.amdahl.ai/api/platform/v1/evals/run \
  -H "X-API-Key: $AMDAHL_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "eval": "prompt-and-message-eval",
    "inputs": {
      "mode": "advisory",
      "prompt": "<your 40-page outbound rules document>",
      "message": "<a message it produced>"
    }
  }'

The response, annotated

The report lives at verdict.cases[].graders[].improvement. Trimmed to the fields worth knowing:

json
{
  "mode": "rewrite",             // the default; "advisory" and "gate" are the alternatives

  "transition": {
    "input_verdict": "fail",        // your draft did not clear the bar - the FINDING
    "improved_verdict": "pass",     // the improvement did
    "transition": "fail_to_pass",
    "threshold": 4.2,               // the bar, config-driven and seeded by Amdahl
    "iterations": 1                 // no revision round was needed
  },

  "facets": [                       // prompt and message NEVER share a score
    {
      "facet": "message",
      "before": {
        "usage": "as_provided",     // exactly what you sent, untouched
        "score_15": 1.8,            // 1 + 4*(1/5) - one rubric line passed
        "checks_passed": 1,
        "checks_total": 5,
        "dimensions": [ { "name": "Relevant positioning", "pass": false, "score": 1, "reasoning": "..." } ],
        //            ^ abridged - 4 of the 5 lines omitted; 1 of the 5 passed
        "quotes":     [ { "text": "Speed honestly isn't our problem...", "stance": "contradicts" } ],
        "good_examples": []
      },
      "after": {
        "usage": "illustration_only",   // <- a specimen, NOT a message to send
        "score_15": 5,                  // all five passed - NOT "perfect"
        "checks_passed": 5,
        "checks_total": 5,
        "dimensions": [ ... ],
        "quotes":     [ ... ],
        "good_examples": [
          { "label": "Opening line", "text": "...", "why": "...", "quotes": [ ... ] }
        ]
      },
      "lift": 0.8                   // the [0,1] gap: 1.0 - 0.2
    },
    {
      "facet": "prompt",
      "before": { "usage": "as_provided",    "score_15": 1.8, "dimensions": [ ... ] },
      "after":  { "usage": "reusable_prompt", "score_15": 4.2, "dimensions": [ ... ] },
      "lift": 0.6
    }
  ],

  "suggestions": [                  // anchored, surgical - see Advisory mode above
    { "kind": "keep", "facet": "message", "title": "Keep the low-friction ask", "...": "..." }
  ],

  "research_steps": [               // validated against the live op registry before you see them
    {
      "op": "data.cluster_search",
      "params": { "query": "release approval friction" },
      "purpose": "Pull the approval theme so the next email can lead with the cohort's own language"
    }
  ],

  "coverage": { "total_chars": 118, "graded_chars": 118, "truncated": false, "strategy": "verbatim" },

  "audience": {                     // discriminated on `status` - narrow before reading `dimensions`
    "status": "resolved",
    "dimensions": { "role_level": "executive", "raw": "VP Engineering at a 500-person SaaS company" },
    "evidence": { "distinct_speakers": 7, "utterances": 210, "distinct_companies": 4 }
  },
  // ...or, when it could not be scoped:
  // "audience": { "status": "abstained", "reason": "thin_evidence", "message": "..." }

  "tool_kit": { "callable": 12, "out_of_scope": 5 },   // research_steps are bounded by YOUR scopes

  "confidence": {                   // how much weight the numbers carry, and why
    "level": "moderate",
    "reasons": ["Only 4 customer quotes backed this grade."]
  },

  "prompt_patch": {                 // the prompt edits above, as one pasteable block
    "heading": "Prompt patch - apply these to your reusable prompt",
    "applies_to_submitted_prompt": true,
    "lines": [ { "kind": "add", "verb": "ADD", "instruction": "...", "why": "..." } ],
    "text": "# Prompt patch - apply these...\n- ADD: ... (...)"
  },

  "grader_meta": {
    "model_calls": 3,
    "blinded": true,                // both candidates scored unlabelled, in one call
    "blinding": {                   // ...and the recorded mechanics behind that claim
      "paired": true,
      "order_shuffled": true,
      "improved_shown_first": false,      // which way it landed on THIS run
      "separate_generate_and_grade": true,
      "generator_model": "claude-sonnet-4-6",
      "judge_model": "claude-sonnet-4-6"
    },
    "evidence_quotes": 14,
    "evidence_frozen": true,        // DEPRECATED - within-run only; read evidence_scope
    "evidence_scope": {             // what the evidence was held fixed ACROSS
      "kind": "run"                 // retrieved for this run; another run has its own
      // when pinned: { "kind": "pinned", "from_run_id": "...", "frozen_at": "...", "as_of": null }
    },
    "evidence_provenance": {        // the three counts, labelled
      "pool": 14,                   // retrieved and frozen for the run
      "cited": 3,                   // distinct ids the judge cited across every graded block
      "resolved": 3,                // of those, how many resolved to a real quote
      "errored_legs": 0
      // when pinned, also: "pinned_from": { "run_id", "frozen_at", "as_of" }
    }
  },

  "before": { "...": "back-compat mirror of the message facet" },
  "after":  { "...": "back-compat mirror of the message facet" },
  "lift": 0.8,
  "what_changed": "Biggest gain: Relevant positioning (1/5 -> 5/5)."
}

What the score counts

Each graded side carries checks_passed and checks_total next to score_15, because they are what the score IS — the fraction, not a rating. The worked example at the top of this page has the arithmetic and the six values the dial takes; the short version is that a score_15 of 5 means "cleared all five checks", not "perfect", which is how a bare 5 beside a threshold reads.

json
"before": { "score_15": 1.8, "checks_passed": 1, "checks_total": 5 },
"after":  { "score_15": 5,   "checks_passed": 5, "checks_total": 5 }

The older spelling passed / total still ships alongside and carries the same integers. Prefer checks_*: passed one level out — on a case, on a grader result — is a boolean verdict, and a shipped grading script that reached for the obvious key on a dimension got undefined and printed every dimension as a failure.

Quote the fraction rather than the rating. Both are the same number; only one says what it measured.

How much the numbers are worth

confidence is a derived read on the run itself — level (high / moderate / low) plus the reasons behind anything below high. It never changes a score; it tells you how firmly to hold one. A run drops below high when the evidence pool was thin, a retrieval leg errored, the run stopped on a budget rather than on the threshold, a long prompt was truncated, or a side was scored on fewer rubric lines than the rubric declares. high carries an empty reasons and is the quiet common case.

Evidence tiers

Every retrieved quote carries a tier, and the tier is a licence for what a claim built on it may say:

tierWhose voiceWhat it licensesEmitted when
accountThe named account you passed in"You told us…" — the only tier that backs an account-specific claim. Carries account.You pass an account and it resolves
segmentThe cohort the run was scoped toA claim at the cohort's own level — "Teams like yours…". A band: never "you" (needs account), and never "one lead told us" (a pattern is not an anecdote).You pass an audience and it clears the evidence floors
corpusA recurring pattern across your conversations"The leaders we talk to…"Always

segment is emitted as of eval version 2.6.0. Before that it was a declared value with no producer: a run came back account or corpus only, so "teams like yours" copy had no tier that could back it. A resolved audience now draws that cohort's own utterances — spread across companies, so one talkative account cannot stand in for the group — and they arrive tagged segment.

segment_status and segment_quotes on the improvement report tell you which happened. A run that scoped to a cohort your workspace has too little of comes back segment_status: "abstained" with the reason, and the absence of segment quotes then IS a statement about your data — the opposite of what it used to mean.

A quote with no tier predates tiering and reads as corpus — the weakest standing, never the strongest. Reaching past a quote's tier (attributing a corpus quote to the account) is graded as a grounding failure, not a style note.

Account-tier evidence only exists when you pass an account AND that account is in your data with buyer-side conversation. Otherwise the run says which of those was missing, on account_abstain_reason, rather than quietly grading you on cohort evidence under an account heading:

account_abstain_reasonWhat it meansWhat to do
not_providedYou sent no account. The run was never asked to scope to one.Nothing — this is the unscoped default.
unresolvableThe name carries characters the lookup will not accept, so it was refused rather than escaped.Re-send the plain company name.
not_in_corpusNo company matched that name at all.Check the spelling — names match as written, so use the one your CRM records use.
no_quotable_utterancesThe company matched, and nothing it said clears the citable band — every utterance is internal, under 40 characters, or over 2,000.Nothing to fix on your side; there is no buyer-side quote to ground on.
lookup_failedOur check broke.Re-run. This is never a statement about your data.

The citable band has three exclusions, not two, and the third is the one people do not expect: a buyer-side utterance longer than 2,000 characters is dropped as well, because an un-segmented transcript block is not a quotable line. So an account whose conversations are held as long, unbroken turns can abstain no_quotable_utterances while holding thousands of utterances — the count in account_abstain_detail is what makes that visible.

Every abstain that has something specific to say carries it on account_abstain_detail, beside the reason: which name found nothing, or the account's real utterance count when nothing it said was citable. The reason is the branch to code against; the detail is the sentence to read. The same text rides the live account_resolved progress step, so a caller watching a run sees it at the moment it is decided.

not_in_corpus and no_quotable_utterances split what shipped through eval version 2.7.0 as one no_conversations label. They mean opposite things and ask opposite things of you — "we hold nothing under that name" versus "we hold this account and it has never said anything we can quote" — and the collapsed label read as the second when it was usually the first. Matching is a case-insensitive substring on the company name as your records spell it, so a brand name that differs from the CRM's (LlamaIndex where the record says Runllama) lands on not_in_corpus with a full conversation history sitting under the canonical name. If you switch on this field, add the two new values; no_conversations is no longer emitted.

"Scored blind", checkable

blinded: true on its own is an assertion. grader_meta.blinding is the set of facts the run recorded while executing, which is what makes it a disclosure: both candidates scored in one call against one rubric and one evidence set (paired), presentation order taken from a content hash of the text you submitted (order_shuffled), which way that landed on this run (improved_shown_first), and whether writing and scoring were separate calls (separate_generate_and_grade) rather than one model marking its own homework. The two model ids are recorded too.

The prompt patch

prompt_patch is the prompt-facet suggestions assembled into ONE pasteable block. It is composed server-side from the same suggestions[] the report itemizes, so it can never claim an edit the itemized list does not; text is the plain-ASCII body to paste, lines[] the same edits structured. Absent when there are no prompt suggestions — you get null, never an empty block. applies_to_submitted_prompt: false means you sent no prompt, so the lines seed one rather than edit one.

The prompt half is the half that compounds: a fixed message helps one send, a fixed prompt helps every send after it.

usage is the field to read. It tells you what each artifact is, so you never present a specimen as a ready-to-send email:

ValueWhat it is
as_providedExactly what you sent. Untouched.
simulated_specimenA draft Amdahl wrote from your prompt so there was something to score. Nobody sent it.
reusable_promptThe takeaway. A template you keep and re-run.
illustration_onlyAn example produced so the score difference could be measured. Evidence, not an email.

The scores are a coach's before/after read, grounded in your cited quotes. They are not a measured reply rate or a conversion lift. If you quote a number from this report to someone, say which it is.

Check lift_reportable before you quote the lift

A judge scoring the same text twice does not return exactly the same number, so a small lift can be the instrument moving rather than the writing improving. Every report carries a lift_reportable boolean answering the only question that matters here: is this difference bigger than the grader's own run-to-run noise?

json
"lift": 0.7,
"lift_reportable": true

When it is false, the lift is still returned — it is useful in aggregate across many runs — but it should not be quoted as a figure for a single message. The per-dimension reasoning and the cited quotes are still fully valid either way; it is only the single headline number that is unresolvable.

The threshold is not one fixed value. A side's score is the mean of its per-dimension scores, so more rubric lines make a steadier mean, and the threshold scales as 1/sqrt(n) — a measured relationship on this grader, not an assumed one. Two things follow:

  • An eval you author with a wider rubric can legitimately resolve a smaller lift than one with a narrow rubric.
  • If the judge returned fewer dimensions than the rubric declares — visible whenever dimensions_scored is below the length of dimensions — the score rests on fewer real judgements than it looks like it does, and the threshold widens to match. The same lift can be reportable on a fully-scored run and not on a short-scored one.

Full request/response shapes, the polling contract, and the grader-kind catalog are in Evals.

Renamed. This eval shipped first as message-grader, then as outreach-eval. The slug is now prompt-and-message-eval — it grades two artifacts, a prompt and a message, and it is not tied to a channel. Both old slugs still resolve and the outbound_message input key is still accepted as an alias for message, so existing integrations keep working — but new code should use prompt-and-message-eval with message.