Docs

Prompt and Message Eval

Paste an email and the prompt behind it. Get back a grade against what your buyers actually said, the specific lines to change, and a reusable prompt your whole team can run - with a worked example of a real cold email graded end to end

Slug: prompt-and-message-eval. The eval Amdahl ships. It answers the question every rep asks and no tool has ever answered honestly: is this email any good, and what specifically would make it better?

Not "is it well written". Not "does it match our tone". Does it say something these buyers actually care about, and can we prove it? The grade is scored against verbatim quotes pulled from your own calls and emails, so the feedback is your customers' words, not a model's taste.

Needs no setup. It is the default, so a run with no eval param uses it.


Start here: a real one, graded

Here is a cold email a rep sent last week. It is not a strawman — it is the shape of most first-touch outbound.

text
Subject: Quick question

Hi Dana - saw Northwind is scaling fast. We help engineering teams
ship 10x faster with best-in-class rollout tooling. Worth a quick
15 minutes next week?

And the prompt behind it:

text
Write a short cold email to a VP of Engineering about our rollout tooling.
Keep it under 100 words and friendly.

Run both through the eval and this is the shape of what comes back.

1. Your message, graded — with the reason for every number

DimensionScoreWhy
Relevant positioning1/5Nothing here is about Northwind. "Scaling fast" applies to every company in the segment; the rest is a product description.
Grounding1/5No claim maps to anything these buyers said. The retrieved quotes talk about release approvals, not speed.
Verified specifics1/5"10x faster" is not supported by any evidence in the corpus, and it is the kind of number a buyer will ask you to defend.
Differentiation2/5"Best-in-class rollout tooling" is a category, not a position. Three of your competitors say it.
CTA clarity3/5The ask is clear and low-friction. It is the strongest part of the email.

Headline: 1.6/5. That number is the mean of the five above — it is never a separate holistic guess, so you can always take it apart.

2. The quotes it graded you against

These are verbatim, pulled from your own conversations, tagged with whether they support or contradict what your email claimed:

Contradicts"Speed honestly isn't our problem. We can ship in a day. The problem is that four people have to sign off and one of them is always on PTO." — Rollout approval friction

Contradicts"We already tried the fast-deploy pitch internally. It didn't move anyone. What moved people was showing the audit trail." — Prior attempts

Supports"Every release we do by hand costs us most of a Thursday." — Manual release cost

That third quote is the one worth building on. The first two are why "10x faster" landed at 1/5 — your buyers said, in their own words, that speed is not the problem.

These are cohort quotes, not Northwind quotes. Retrieval searches your whole corpus for the themes your message is about. It does not filter to the account you are writing to — and on a genuinely cold account it could not, because they have never spoken to you. What comes back is what buyers like Dana say, which is exactly what makes it usable on an account you have no history with.

That also sets a hard line on how you may write it. "The eng leaders we talk to say…" is supported. "Your engineers said…" is not. Naming the wrong scope turns a well-grounded line into a claim of a conversation that never happened — the most common way a good email becomes a false one.

3. The improved prompt — the thing you actually keep

This is the durable output. Not the email. The prompt.

text
You are writing a first-touch email to one specific account. Before you write a word:

1. RESEARCH, in this order. Take everything each tier gives you.
   a. THE COHORT - always available. In Amdahl, find how people in this role
      and this industry talk about the problem area, and note their exact
      language. This is the tier that works on an account nobody has spoken to.
   b. THE TRIGGER - sometimes. Something that just happened to them: funding,
      a hire, a launch, a move. Lead with it when it exists. Never invent one.
   c. THE ACCOUNT - only when they are already in our data. Their own words and
      their own history outrank the cohort. Most first-touch accounts will not
      have this, and that is fine - it is a bonus tier, not a gate.
   If the COHORT comes back empty, say so and stop; do not write from the
   category.

2. ATTRIBUTE at the scope you actually have. "The eng leaders we talk to say"
   is supported by cohort evidence. "Your engineers said" is not, unless tier
   (c) turned up that exact conversation. Never imply a prior interaction that
   did not happen.

3. POSITION. Frame the offer against the problem they named, in their words.
   If the stated problem is approvals and yours is a speed product, lead with
   approvals or do not send. Never open with how fast we are unless a buyer
   said slowness was costing them something.

4. VERIFY. Every number, outcome, or capability claim must be checkable against
   a real call or a real customer. If you cannot point at the evidence, CUT the
   claim. Do not soften it - a hedged unverifiable claim is still unverifiable.

5. DISQUALIFY. If this prospect is not who we sell to - wrong role, wrong
   segment, or a problem we do not solve - stop and say why. If the cohort has
   already tried and rejected this framing, pick a different one.

6. ASK. One next step, sized to the relationship. A 15-minute call is fine for
   a warm account and too much for a cold one - offer to send the thing instead.

Keep it under 120 words. Plain sentences. No category adjectives
("best-in-class", "industry-leading", "revolutionary").

Notice what it is: a template with slots, not a rewrite of this one email. Paste it for the next account tomorrow and you get an equally specific result. That is what it is graded on.

4. The example output — and what it is not

This is not an email to send. It is a specimen produced so the score difference could be measured — evidence that the improved prompt works. On the wire it is typed usage: "illustration_only", and every Amdahl surface labels it "Example output — what the improved prompt produces". Send it verbatim and you are sending a message written by a model that has never met Dana.

text
Subject: The four sign-offs

Hi Dana - the eng leaders we talk to keep saying shipping isn't the
bottleneck, approvals are: four sign-offs, and someone's always out.

We built the approval path for exactly that - the audit trail is the
product, not a side effect. One customer went from four serial
approvals to two parallel ones.

Want me to send the two-page teardown of how they did it? No call needed.

Graded the same way, by the same judge, against the same quotes: 4.4/5. Relevant positioning went 1 → 5 (it leads on approvals, which is what these buyers said), grounding 1 → 5 (it uses their exact language, attributed to the cohort it came from), verified specifics 1 → 4 (the one number is one Amdahl can point at), differentiation 2 → 4, CTA 3 → 4 (an artifact instead of a meeting is lower friction for a cold account).

Note what it does not say: "one of your engineers told us." Nobody at Northwind has spoken to you. The line is specific and true at the same time because it is attributed to the cohort the evidence actually came from.

5. The lift, and the transition

code
Your message      1.6/5   fail
Improved          4.4/5   pass     fail -> pass

Your message failing is the finding, not an error. It is the thing you ran the eval to learn. The eval reports the two verdicts separately (input_verdict and improved_verdict) precisely so a working run never reports "fail" because the draft it was asked to diagnose was the bad one.

6. What changes when the account is warm

Everything above is the cold case: Dana has never spoken to you, so tier (c) is empty and the evidence is her cohort's. When the account is in your data, exactly one thing changes:

  • Changes — tier (c) fires. Their own words and their own history outrank the cohort, and you can attribute directly to them, because now you actually can.
  • Does not change — verify, disqualify, and attribute. A warm account is not a licence for a claim you cannot point at.

You never tell the eval which case you are in. It grades what you wrote against what your buyers said either way. The tiers exist so the prompt still produces something specific when tier (c) is empty — which, for first-touch outbound, is most of the time.


The other mode: you already have a prompt, and it is enormous

Plenty of teams do not have a one-line prompt. They have a living rules document — 40 pages of positioning guidance, objection handling, and things legal will not let you say. Nobody is going to replace that because an eval suggested a new one.

Set mode: "advisory" and the eval leaves your document alone. Instead of a rewrite you get anchored, surgical suggestions:

json
{
  "kind": "add",
  "facet": "prompt",
  "anchor_quote": "Always mention our SOC 2 certification in the first email.",
  "anchor_offset": 8214,
  "section_id": "s6",
  "title": "Gate the SOC 2 line on the buyer actually asking",
  "detail": "Change this to: mention SOC 2 only when the account has raised security, compliance, or procurement. Otherwise cut it.",
  "why": "Across the retrieved quotes, security comes up from procurement and never from the engineering buyer this sequence targets. Leading with it reads as boilerplate to the person receiving it.",
  "dimension": "Relevant positioning",
  "quotes": [
    { "text": "Honestly the SOC 2 stuff is procurement's problem, not mine.", "stance": "contradicts" }
  ]
}

Five kinds of suggestion come back:

KindWhat it means
keepThis is working. Do not lose it in the next edit. A report that is all criticism does not tell you what to protect.
addA missing instruction, stated so you can paste it in
strengthenA vague instruction that is not doing its job yet
removeAn instruction that is actively hurting the output
reorderThe content is right, the sequence is the problem

anchor_quote is verified server-side. It is only ever a literal substring of your document — if the model paraphrases, the anchor is dropped and you get a document-level suggestion instead. A suggestion can never quote you a line you did not write. anchor_offset is the verified position, which is what makes an anchor findable when the same phrase repeats.

Long documents are sectioned, never quietly truncated

Past 12,000 characters your prompt is split on your own headings and packed head-and-tail (the opening frames the task; the closing usually carries the hard rules). Every run reports coverage:

json
"coverage": {
  "total_chars": 48219,
  "graded_chars": 23904,
  "truncated": true,
  "strategy": "sectioned",
  "sections": [
    { "id": "s1", "heading": "Who we sell to", "chars": 1840, "included": true },
    { "id": "s2", "heading": "Positioning by segment", "chars": 6210, "included": true },
    { "id": "s7", "heading": "Legacy objection scripts", "chars": 9022, "included": false }
  ]
}

You always know what was read. Past 250,000 characters the run refuses rather than returning a confident number computed over a sliver of your document.


Who uses this, and how

Reps: before you send a sequence

Paste step 1 and the prompt you used. You get the specific lines that are not landing and the quotes that prove it. Keep the improved prompt; reuse it per account. Two minutes, and you stop sending the "10x faster" email to a team whose problem is approvals.

Managers: audit the sequence, not the rep

Run every step of a live sequence through it. The pattern across the grades is the coaching: if grounding is 1/5 on every step, the problem is the template, not the person sending it. The suggestions tell you exactly which lines to change.

Enablement: keep the rules doc honest

Run your positioning document in advisory mode on a cadence. As what customers say changes, the suggestions change with it - and keep tells you which parts are still earning their place.

Ops: wire it into the review step

It is one API call and a poll. Gate sequence publishing on a passing improved verdict, or just log the transition so you can see whether outbound quality is moving.

A sequence review, end to end

Run every step

One evals.run per sequence step, each with the step's copy as message and the shared sequence prompt as prompt. They are independent - fire them all and poll.

Read the dimensions, not the headline

A 2.4/5 tells you nothing. Grounding 1/5 across all five steps tells you the sequence was written from the category, not from customers.

Take the prompt, not the emails

The improved prompt is the artifact. Put it in the sequence brief so the next person to write a step starts from it.

Fix the anchored lines

Every remove and strengthen suggestion points at a specific line with the quote that justifies it. Work the list.

Re-run and watch the transition

fail -> pass on the improved side means the prompt is doing its job. fail -> fail with an explanation means the evidence does not support a stronger version - which is usually telling you something real about the segment.


How it works under the hood

Five stages. Four of them exist to make the number trustworthy rather than merely produced:

  1. Read the intent. One cheap pass extracts what you are trying to do — the offer, the ask, and each specific claim your draft makes. Retrieval is seeded from that, never from your draft. (Seeding on a weak draft returns generic themes, and those generic themes then become the evidence the draft is held to. That is circular.)
  2. Gather evidence. Several retrieval passes run at once — one per intent seed, one per specific claim — so "cuts onboarding from six weeks to two" is checked against the themes that speak to that, not against the average of your whole message. The set is then frozen for the entire run.
  3. Write. A model produces the improvement, the suggestions, and the examples. It scores nothing.
  4. Grade, blind. A separate pass scores both candidates in one call, unlabelled, with the presentation order derived from a content hash. The judge does not know which one it wrote, so it cannot be kind to itself.
  5. Revise, if needed. If the improved version misses the bar (4.2/5 — deliberately higher than the 3.5 your draft is read against), it gets exactly one more attempt, guided by the judge's own critique. Then it stops and tells you honestly why it could not do better.

A quote in a verdict is always a real utterance. The model may only cite quote ids it was handed; the server swaps those ids back to verbatim text and drops any it does not recognise. Inventing a customer saying something is structurally impossible, not merely discouraged.

Honest failure

No customer data yet? The run comes back not_applicable — never a false fail. An empty corpus means there is nothing to grade against, and saying so is more useful than a fabricated score.

The improvement did not clear the bar? You get the score, the transition, and an explanation naming what still holds it back. No silent rounding up.


Inputs

FieldRequiredNotes
promptThe instruction, brief, or rules document behind the message. Up to 250,000 characters; sectioned past 12,000. Send this alone and it writes a specimen draft to grade.
messageThe drafted message (email, LinkedIn note, whatever). Send this alone and it also suggests a reusable prompt. The retired key outbound_message is still accepted.
audienceWho it is going to. Name a seniority (executive, manager, individual contributor) or a title (VP of Sales, Head of L&D) and the run scopes to that cohort — see Audience scoping.
moderewrite (default) or advisory. Use advisory when your prompt is a large document you are not going to replace.

At least one of prompt / message must be supplied. Send both when you have both — grading the pair is what produces a useful improved prompt, because the eval can see what the prompt asked for and what it actually got.

Run it

bash
curl -X POST https://app.amdahl.co/api/platform/v1/evals/run \
  -H "X-API-Key: $AMDAHL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "eval": "prompt-and-message-eval",
    "inputs": {
      "prompt": "Write a short cold email to a VP of Engineering about our rollout tooling. Keep it under 100 words and friendly.",
      "message": "Hi Dana - saw Northwind is scaling fast. We help engineering teams ship 10x faster with best-in-class rollout tooling. Worth a quick 15 minutes next week?",
      "audience": "VP Engineering at a 500-person SaaS company"
    }
  }'

The response is a handle. Poll it until the run settles:

bash
curl https://app.amdahl.co/api/platform/v1/eval-runs/$RUN_ID \
  -H "X-API-Key: $AMDAHL_API_KEY"

Audience scoping

If you pass an audience, the run does two things with it before grading anything.

First it resolves it to a seniority — executive, manager, or individual contributor. Titles resolve directly ("VP of Sales", "Head of L&D", "SDRs", "Sr. Mgr., Talent"), and so do the bare words. Something that names a function, a vertical, or a company size ("marketing", "healthcare", "mid-market logistics") names no seniority, and the run says so rather than guessing one.

Then it checks your data for that cohort. The report is only scoped to an audience once your workspace has enough recorded conversations with them to be worth grading positioning against:

FloorWhy
3+ distinct peopleOne person is an anecdote.
25+ utterancesThree people who each said one line in passing is not a voice.
2+ distinct companiesWithout this, one talkative account clears the other two on its own and its vocabulary gets reported as an audience-wide finding.

All three must pass. If they do not, the run still completes and still grades your prompt and message — it just grades them against your whole customer corpus, and the report says which of five things happened:

What the report saysWhat it means
not_providedYou did not pass an audience. Nothing to do.
unresolvableWe could not match your text to a seniority. Naming a level scopes it.
no_evidenceWe understood the cohort; your workspace has no recorded conversations with them yet.
thin_evidenceSome conversations, below the floors. The counts are in the report so you can see how close.
lookup_failedThe check itself failed on our side. This is not a statement about your data.

A scoped run carries the counts it decided on, so you can see the evidence behind the framing rather than taking it on trust.

Runnable calls in the report

research_steps are real Amdahl calls you can paste to close the evidence gaps the report found. Every one is validated before it ships — an operation that does not exist, one that would mutate your workspace, or SQL our query gate would refuse is dropped rather than shown.

They are also bounded by what your key can actually run: a step you would get a 403 for is not a useful suggestion, so the report only proposes calls covered by your own scopes. If your key is narrower than the full surface, the report notes how many further calls exist without listing them.

Advisory mode

bash
curl -X POST https://app.amdahl.co/api/platform/v1/evals/run \
  -H "X-API-Key: $AMDAHL_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "eval": "prompt-and-message-eval",
    "inputs": {
      "mode": "advisory",
      "prompt": "<your 40-page outbound rules document>",
      "message": "<a message it produced>"
    }
  }'

The response, annotated

The report lives at verdict.cases[].graders[].improvement. Trimmed to the fields worth knowing:

json
{
  "mode": "rewrite",

  "transition": {
    "input_verdict": "fail",        // your draft did not clear the bar - the FINDING
    "improved_verdict": "pass",     // the improvement did
    "transition": "fail_to_pass",
    "threshold": 4.2,               // the bar, config-driven and seeded by Amdahl
    "iterations": 1                 // no revision round was needed
  },

  "facets": [                       // prompt and message NEVER share a score
    {
      "facet": "message",
      "before": {
        "usage": "as_provided",     // exactly what you sent, untouched
        "score_15": 1.6,
        "dimensions": [ { "name": "Relevant positioning", "score": 1, "reasoning": "..." } ],
        "quotes":     [ { "text": "Speed honestly isn't our problem...", "stance": "contradicts" } ],
        "good_examples": []
      },
      "after": {
        "usage": "illustration_only",   // <- a specimen, NOT a message to send
        "score_15": 4.4,
        "dimensions": [ ... ],
        "quotes":     [ ... ],
        "good_examples": [
          { "label": "Opening line", "text": "...", "why": "...", "quotes": [ ... ] }
        ]
      },
      "lift": 0.7
    },
    {
      "facet": "prompt",
      "before": { "usage": "as_provided",    "score_15": 2.0, "dimensions": [ ... ] },
      "after":  { "usage": "reusable_prompt", "score_15": 4.6, "dimensions": [ ... ] },
      "lift": 0.65
    }
  ],

  "suggestions": [                  // anchored, surgical - see Advisory mode above
    { "kind": "keep", "facet": "message", "title": "Keep the low-friction ask", "...": "..." }
  ],

  "research_steps": [               // validated against the live op registry before you see them
    {
      "op": "data.cluster_search",
      "params": { "query": "release approval friction" },
      "purpose": "Pull the approval theme so the next email can lead with the cohort's own language"
    }
  ],

  "coverage": { "total_chars": 118, "graded_chars": 118, "truncated": false, "strategy": "verbatim" },

  "grader_meta": {
    "model_calls": 3,
    "blinded": true,                // both candidates scored unlabelled, in one call
    "evidence_quotes": 14,
    "evidence_frozen": true         // same evidence for every round - the lift is like-for-like
  },

  "before": { "...": "back-compat mirror of the message facet" },
  "after":  { "...": "back-compat mirror of the message facet" },
  "lift": 0.7,
  "what_changed": "Biggest gain: Relevant positioning (1/5 -> 5/5)."
}

usage is the field to read. It tells you what each artifact is, so you never present a specimen as a ready-to-send email:

ValueWhat it is
as_providedExactly what you sent. Untouched.
simulated_specimenA draft Amdahl wrote from your prompt so there was something to score. Nobody sent it.
reusable_promptThe takeaway. A template you keep and re-run.
illustration_onlyAn example produced so the score difference could be measured. Evidence, not an email.

The scores are a coach's before/after read, grounded in your cited quotes. They are not a measured reply rate or a conversion lift. If you quote a number from this report to someone, say which it is.

Full request/response shapes, the polling contract, and the grader-kind catalog are in Evals.

Renamed. This eval shipped first as message-grader, then as outreach-eval. The slug is now prompt-and-message-eval — it grades two artifacts, a prompt and a message, and it is not tied to a channel. Both old slugs still resolve and the outbound_message input key is still accepted as an alias for message, so existing integrations keep working — but new code should use prompt-and-message-eval with message.