Docs

Grade a cold email

One cold email and the prompt behind it, run through prompt-and-message-eval and read end to end: the per-line verdicts, the quotes, the improved prompt, the illustration, and a sequence review built from the same call.

Level: first run. Instrument: prompt-and-message-eval. You need: a Customer agent key and a workspace with synced conversations.

This walkthrough grades one email and reads every part of the result. The email, the company and the quotes are an illustrative example in the shape of most first-touch outbound, not data from a real workspace; your own run returns your customers' words.

The email and the prompt

Here is the cold email:

text
Subject: Quick question

Hi Dana - saw Northwind is scaling fast. We help engineering teams
ship 10x faster with best-in-class rollout tooling. Worth a quick
15 minutes next week?

And the prompt behind it:

text
Write a short cold email to a VP of Engineering about our rollout tooling.
Keep it under 100 words and friendly.

What comes back

Run both through the eval (the request is under Run it yourself) and this is the shape of what comes back. rewrite is the default mode, and it produces the replacement prompt and the specimen message this walkthrough shows. Send "mode": "advisory" instead when your prompt is a document you are keeping; it grades the same way but returns anchored suggestions against what you already have — see advisory mode.

What the report opens with
failpass

Your email argued speed. These buyers said the problem is approvals.

Your draft scored 1.6 and did not clear the bar. The improved version scored 4.4 against the same quotes, under the same blind rubric.

fail to pass

It leads on the approval bottleneck the cohort actually named, attributes the claim to the cohort it came from, drops the unverifiable 10x, and asks for a teardown instead of a call.

One run. Re-run a few times and read the median.
Scored blind3 real quotesGraded against your whole corpus

1. Your message, graded — with the reason for every number

DimensionVerdictWhy
Relevant positioningfailNothing here is about Northwind. "Scaling fast" applies to every company in the segment; the rest is a product description.
GroundingfailNo claim maps to anything these buyers said. The retrieved quotes talk about release approvals, not speed.
Verified specificsfail"10x faster" is not supported by any evidence in the corpus, and it is the kind of number a buyer will ask you to defend.
Differentiationfail"Best-in-class rollout tooling" is a category, not a position. Three of your competitors say it.
CTA claritypassThe ask is clear and low-friction. It is the strongest part of the email.

Headline: 1.8/5 — one of five checks passed.

Every rubric line is a binary pass/fail, so the headline is not a holistic guess and not an average of five opinions. It is the pass fraction placed on a five-point axis:

code
score_15 = 1 + 4 × (passed / total)

Five lines means the dial takes exactly six values — 1, 1.8, 2.6, 3.4, 4.2, 5 — and nothing in between. 1.8 is exactly one line passed. The response ships checks_passed and checks_total next to score_15 for that reason: "1 of 5 checks" says what the number measured; "1.8" on its own reads like a rating on a scale that does not exist. Quote the fraction.

Here are the same five as the console renders them, each row judged on both sides and expandable to the reasoning behind its verdict:

The dimensions, expandable
Relevant positioningYours1/5Improved5/5 pass
Your draftNothing here is about Northwind. "Scaling fast" applies to every company in the segment; the rest is a product description.
ImprovedIt leads on approvals, which is what these buyers said the problem was.
GroundingYours1/5Improved5/5 pass
Your draftNo claim maps to anything these buyers said. The retrieved quotes talk about release approvals, not speed.
ImprovedIt uses their exact language, attributed to the cohort the evidence actually came from.
Verified specificsYours1/5Improved4/5 pass
Your draft"10x faster" is not supported by any evidence in the corpus, and it is the kind of number a buyer will ask you to defend.
ImprovedThe one number in it is a number Amdahl can point at.
DifferentiationYours2/5Improved4/5 pass
Your draft"Best-in-class rollout tooling" is a category, not a position. Three of your competitors say it.
ImprovedThe audit trail as the product is a position, not a category claim.
CTA clarityYours3/5Improved4/5 pass
Your draftThe ask is clear and low-friction. It is the strongest part of the email.
ImprovedAn artifact instead of a meeting is lower friction for a cold account.

2. The quotes it graded you against

These are verbatim, pulled from your own conversations, tagged with whether they support or contradict what your email claimed:

Contradicts — "Speed honestly isn't our problem. We can ship in a day. The problem is that four people have to sign off and one of them is always on PTO." — Rollout approval friction

Contradicts — "We already tried the fast-deploy pitch internally. It didn't move anyone. What moved people was showing the audit trail." — Prior attempts

Supports — "Every release we do by hand costs us most of a Thursday." — Manual release cost

That third quote is the one worth building on. The first two are why Verified specifics failed — your buyers said, in their own words, that speed is not the problem, so "10x faster" has nothing behind it.

The evidence, verbatim
3shown below
Speed honestly isn't our problem. We can ship in a day. The problem is that four people have to sign off and one of them is always on PTO.
Customer · Executive · Negotiation · 12 May 2026Across your customers
We already tried the fast-deploy pitch internally. It didn't move anyone. What moved people was showing the audit trail.
CustomerAcross your customers
Every release we do by hand costs us most of a Thursday.
Across your customers

These are cohort quotes, not Northwind quotes. Every quote carries a tier saying whose voice it is, and here all three are corpus: this run passed no audience, so retrieval searched your whole corpus for the themes your message is about. On a cold account it can do nothing narrower — Northwind has never spoken to you. What comes back is what buyers like Dana say, which is exactly what makes it usable on an account you have no history with.

Pass an audience and the run additionally draws that cohort's own words, tagged segment — a tighter licence than corpus and the right one for a first touch. See Evidence tiers.

That also sets a hard line on how you may write it. "The eng leaders we talk to say…" is supported. "Your engineers said…" is not. Naming the wrong scope turns a well-grounded line into a claim of a conversation that never happened — the most common way a good email becomes a false one, and the eval grades it as a grounding failure rather than a style note.

Pass an account and, when that company IS in your data, their own buyer-side words come back tagged tier: "account" — the one tier that licenses "you told us". See Evidence tiers.

3. The improved prompt — the thing you actually keep

This is the durable output. Not the email. The prompt.

text
You are writing a first-touch email to one specific account. Before you write a word:

1. RESEARCH, in this order. Take everything each tier gives you.
   a. THE COHORT - always available. In Amdahl, find how people in this role
      and this industry talk about the problem area, and note their exact
      language. This is the tier that works on an account nobody has spoken to.
   b. THE TRIGGER - sometimes. Something that just happened to them: funding,
      a hire, a launch, a move. Lead with it when it exists. Never invent one.
   c. THE ACCOUNT - only when they are already in our data. Their own words and
      their own history outrank the cohort. Most first-touch accounts will not
      have this, and that is fine - it is a bonus tier, not a gate.
   If the COHORT comes back empty, say so and stop; do not write from the
   category.

2. ATTRIBUTE at the scope you actually have. "The eng leaders we talk to say"
   is supported by cohort evidence. "Your engineers said" is not, unless tier
   (c) turned up that exact conversation. Never imply a prior interaction that
   did not happen.

3. POSITION. Frame the offer against the problem they named, in their words.
   If the stated problem is approvals and yours is a speed product, lead with
   approvals or do not send. Never open with how fast we are unless a buyer
   said slowness was costing them something.

4. VERIFY. Every number, outcome, or capability claim must be checkable against
   a real call or a real customer. If you cannot point at the evidence, CUT the
   claim. Do not soften it - a hedged unverifiable claim is still unverifiable.

5. DISQUALIFY. If this prospect is not who we sell to - wrong role, wrong
   segment, or a problem we do not solve - stop and say why. If the cohort has
   already tried and rejected this framing, pick a different one.

6. ASK. One next step, sized to the relationship. A 15-minute call is fine for
   a warm account and too much for a cold one - offer to send the thing instead.

Keep it under 120 words. Plain sentences. No category adjectives
("best-in-class", "industry-leading", "revolutionary").

Notice what it is: a template with slots, not a rewrite of this one email. Paste it for the next account tomorrow and you get an equally specific result. That is what it is graded on.

4. The example output — and what it is not

This is not an email to send. It is a specimen produced so the score difference could be measured — evidence that the improved prompt works. On the wire it is typed usage: "illustration_only", and every Amdahl surface labels it "Example output — what the improved prompt produces". Send it verbatim and you are sending a message written by a model that has never met Dana.

text
Subject: The four sign-offs

Hi Dana - the eng leaders we talk to keep saying shipping isn't the
bottleneck, approvals are: four sign-offs, and someone's always out.

We built the approval path for exactly that - the audit trail is the
product, not a side effect. One customer went from four serial
approvals to two parallel ones.

Want me to send the two-page teardown of how they did it? No call needed.

Graded the same way, by the same judge, against the same quotes: 5/5 — all five checks passed. Relevant positioning went fail → pass (it leads on approvals, which is what these buyers said), grounding fail → pass (it uses their exact language, attributed to the cohort it came from), verified specifics fail → pass (the one number is one Amdahl can point at), differentiation fail → pass, and CTA clarity was the single line that already passed and still does (an artifact instead of a meeting is lower friction for a cold account).

Note what it does not say: "one of your engineers told us." Nobody at Northwind has spoken to you. The line is specific and true at the same time because it is attributed to the cohort the evidence actually came from.

Both sides, as you would read them in the console. The reusable prompt is the artifact worth keeping; the message is labelled an illustration everywhere it appears:

Yours and improved, side by side
Improved version4.4of 5example -- what the improved prompt produces, not a message to send

Subject: The four sign-offs Hi Dana - the eng leaders we talk to keep saying shipping isn't the bottleneck, approvals are: four sign-offs, and someone's always out. We built the approval path for exactly that - the audit trail is the product, not a side effect. One customer went from four serial approvals to two parallel ones. Want me to send the two-page teardown of how they did it? No call needed.

An example, written to show the difference. Regenerate from the improved prompt for anything you actually send.

Improved prompt4.6of 5improved -- reusable

You are writing a first-touch email to one specific account. Before you write a word: 1. RESEARCH, in this order. Take everything each tier gives you. a. THE COHORT - always available. In Amdahl, find how people in this role and this industry talk about the problem area, and note their exact language. This is the tier that works on an account nobody has spoken to. b. THE TRIGGER - sometimes. Something that just happened to them: funding, a hire, a launch, a move. Lead with it when it exists. Never invent one. c. THE ACCOUNT - only when they are already in our data. Their own words and their own history outrank the cohort. Most first-touch accounts will not have this, and that is fine - it is a bonus tier, not a gate. If the COHORT comes back empty, say so and stop; do not write from the category. 2. ATTRIBUTE at the scope you actually have. "The eng leaders we talk to say" is supported by cohort evidence. "Your engineers said" is not, unless tier (c) turned up that exact conversation. Never imply a prior interaction that did not happen. 3. POSITION. Frame the offer against the problem they named, in their words. If the stated problem is approvals and yours is a speed product, lead with approvals or do not send. Never open with how fast we are unless a buyer said slowness was costing them something. 4. VERIFY. Every number, outcome, or capability claim must be checkable against a real call or a real customer. If you cannot point at the evidence, CUT the claim. Do not soften it - a hedged unverifiable claim is still unverifiable. 5. DISQUALIFY. If this prospect is not who we sell to - wrong role, wrong segment, or a problem we do not solve - stop and say why. If the cohort has already tried and rejected this framing, pick a different one. 6. ASK. One next step, sized to the relationship. A 15-minute call is fine for a warm account and too much for a cold one - offer to send the thing instead. Keep it under 120 words. Plain sentences. No category adjectives ("best-in-class", "industry-leading", "revolutionary").

A template with slots, not a rewrite of this one email. Run it against the next account and you get an equally specific result.

5. The lift, and the transition

code
Your message      1.8/5   1 of 5 checks    fail
Improved          5/5     5 of 5 checks    pass     fail -> pass

Your message failing is the finding, not an error. It is the thing you ran the eval to learn. The eval reports the two verdicts separately (input_verdict and improved_verdict) precisely so a working run never reports "fail" because the draft it was asked to diagnose was the bad one.

6. What changes when the account is warm

Everything above is the cold case: Dana has never spoken to you, so tier (c) is empty and the evidence is her cohort's. When the account is in your data, exactly one thing changes:

  • Changes — tier (c) fires. Their own words and their own history outrank the cohort, and you can attribute directly to them, because now you actually can.
  • Does not change — verify, disqualify, and attribute. A warm account is not a licence for a claim you cannot point at.

You never tell the eval which case you are in. It grades what you wrote against what your buyers said either way. The tiers exist so the prompt still produces something specific when tier (c) is empty — which, for first-touch outbound, is most of the time.

Run it yourself

bash
curl -X POST https://app.amdahl.ai/api/platform/v1/evals/run \
  -H "X-API-Key: $AMDAHL_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "eval": "prompt-and-message-eval",
    "inputs": {
      "prompt": "Write a short cold email to a VP of Engineering about our rollout tooling. Keep it under 100 words and friendly.",
      "message": "Hi Dana - saw Northwind is scaling fast. We help engineering teams ship 10x faster with best-in-class rollout tooling. Worth a quick 15 minutes next week?"
    }
  }'

Then wait for it and read the card:

bash
curl "https://app.amdahl.ai/api/platform/v1/eval-runs/$RUN_ID?wait_ms=30000" -H "X-API-Key: $AMDAHL_KEY"
curl "https://app.amdahl.ai/api/platform/v1/eval-runs/$RUN_ID/report" -H "X-API-Key: $AMDAHL_KEY"

Repeat the first read until data.status is complete, failed or canceled. This request passes no audience, like the walkthrough above; add "audience": "VP of Engineering" inside inputs to draw that cohort's own quotes as well.

Who uses this, and how

Reps: before you send a sequence

Paste step 1 and the prompt you used. You get the specific lines that are not landing and the quotes that prove it. Keep the improved prompt; reuse it per account. Two minutes, and you stop sending the "10x faster" email to a team whose problem is approvals.

Managers: audit the sequence, not the rep

Run every step of a live sequence through it. The pattern across the grades is the coaching: if grounding fails on every step, the problem is the template, not the person sending it. The suggestions tell you exactly which lines to change.

Enablement: keep the rules doc honest

Run your positioning document in advisory mode on a cadence. As what customers say changes, the suggestions change with it - and keep tells you which parts are still earning their place.

Ops: wire it into the review step

It is one API call and a poll. Gate sequence publishing on gate.passed (whether the submitted copy cleared the bar), never on the improved side, or log the submitted checks fraction so you can see whether outbound quality is moving.

A sequence review, end to end

Run every step

One evals.run per sequence step, each with the step's copy as message and the shared sequence prompt as prompt. They are independent - fire them all and poll.

Read the dimensions, not the headline

A headline of 2.6/5 tells you nothing on its own. Grounding failing on all five steps tells you the sequence was written from the category, not from customers.

Take the prompt, not the emails

The improved prompt is the artifact. Put it in the sequence brief so the next person to write a step starts from it.

Fix the anchored lines

Every remove and strengthen suggestion points at a specific line with the quote that justifies it. Work the list.

Re-run and watch the transition

Grade the edited step with evidence_from_run set to the first run, so it is scored against the same quotes, and read the submitted checks fraction. fail -> fail with an explanation means the evidence does not support a stronger version, which is usually telling you something real about the segment.

Next