Patterns
Call sequences that work: optimize before send, check before you query, cohort then voice, gate before send, pinned-evidence A/B, eval as a trend, and keeping two verdicts apart.
Each pattern below is a short sequence of Amdahl's calls, with the field that makes it safe. They assume you have read Evidence and trust.
| Pattern | Calls | Buys you |
|---|---|---|
| Optimize before send | The Optimizer, between your writer and your send queue | Every draft rewritten or confirmed, with no data connected and nothing waiting on Amdahl |
| Check before you query | GET /search/overview, then Search | An empty result you can explain |
| Cohort, then voice | Search (filter), then Search (semantic) | A population with its own words attached, no agent run |
| Gate before send | Eval (mode: "gate"), then GET /eval-runs/{id}/gate | A send/hold bit on copy an agent wrote |
| Pinned-evidence A/B | Eval, then Eval with evidence_from_run, then compare | A score difference you can attribute to the edit |
| Eval as a trend | Many Evals, then the KPI read | A week-over-week measure of writing quality |
| Keep two verdicts apart | Any two graders | A well-written claim nobody made stays visible |
Optimize before send
Code that writes outbound runs the Optimizer on each draft before it is queued.
- Generate the final draft, merge fields included.
POST /messages/optimizewith the draft, in a background job with a 180 second timeout. It usually answers in under 30.- On
ok: true, queuemessageexactly as returned: a rewrite whenunchangedisfalse, your own draft when it istrue. - On
ok: falseor a timeout, queue the draft as it was (or hold it) and logreason.
Never post-process message with another model: the checks that kept names, numbers and merge fields exact ran on that text. With no data connected this is a style pass, so it does not check claims; add evidence: "workspace" once your conversations sync, or gate on an eval (below) where claims matter. The full recipe is Optimize before you send.
Check before you query
A workspace that has not synced yet answers every search with an empty, well-formed result. GET /search/overview returns the row count and time span for each warehouse surface, with null meaning the read failed rather than zero rows.
GET /search/overview. If everyrow_countis0, stop and say the workspace has no data yet.- Run the search. An empty result is now a real "nothing matched".
The onboarding skill runs this check first for the same reason.
Cohort, then voice
Find the population with typed filters, then read what it said, without starting an agent.
POST /search/querywithsurface: "deals"and filters such asdeal_stage_status = "won". Collect thecompany_idvalues.POST /search/querywithmode: "semantic",filters: [{ "field": "company_id", "op": "in", "value": [...] }]andhydrate: true. You get the utterances from those accounts, ranked by meaning, with quotable text.
Both calls are synchronous, so the whole sequence finishes in two round trips. The semantic lane holds customer-side speech only (corpus.speaker_scope: "external_only"), which is what you want here.
Gate before send
An agent that writes outbound copy grades it before it acts.
- Generate the message.
POST /evals/runwith{ "eval": "prompt-and-message-eval", "inputs": { "message": "…", "mode": "gate" } }. Gate mode grades only your copy and makes one judge call.GET /eval-runs/{id}/gate?wait_ms=30000untilstatusis terminal.- Branch on
gate.passed:truesends,falseholds,nullabstains and passes through.
Do not gate on the improved side, on lift, or on verdict. The gate rules say why each one misleads. On a hold, run a full eval once and regenerate from its improved prompt; never send the illustration_only message it produced.
Pinned-evidence A/B
Two eval runs each retrieve their own quotes, seeded from what you sent. Edit a draft and re-grade it and you have changed the copy and the evidence at once.
- Grade version A normally.
- Grade version B with
evidence_from_runset to A's run id. B reuses A's frozen quotes, so only the copy changed. GET /eval-runs/{a}/compare/{b}. It returns a delta only when the two were graded on the same basis and the gap is larger than the judge's own spread; otherwise it names which control failed indelta_withheld_reason.
For a small edit, grade each version several times and compare medians. One run is one draw.
Eval as a trend
One run is a report card. The trend over many runs is the metric.
- Grade every draft that matters, or a fixed sample each week.
- Read
GET /evals/prompt-and-message-eval/kpi?window_days=90&granularity=week. It tracks the submitted side only, abstains on windows with fewer than three scored runs, and counts refusals separately rather than as zeros.
Read eval_version before you compare windows: two versions are two instruments. A Routine can post the movement to your team each week.
Keep two verdicts apart
A message can be well written and assert something no customer ever said. A single blended number hides that.
When you combine an eval with any other grader, a house style check, or a person's review, report each verdict on its own. Within one eval, read dimensions[] line by line: Grounding and Verified specifics answer "can we back this up", and CTA clarity answers something else.
See also
- Use cases: GTM jobs mapped to these patterns.
- Optimize before you send: the Optimizer in a send pipeline, with code.
- Grade a cold email: one run, read end to end.
- The grading loop: the full write, ground, grade, fix loop.