Docs

Patterns

Call sequences that work: optimize before send, check before you query, cohort then voice, gate before send, pinned-evidence A/B, eval as a trend, and keeping two verdicts apart.

Each pattern below is a short sequence of Amdahl's calls, with the field that makes it safe. They assume you have read Evidence and trust.

PatternCallsBuys you
Optimize before sendThe Optimizer, between your writer and your send queueEvery draft rewritten or confirmed, with no data connected and nothing waiting on Amdahl
Check before you queryGET /search/overview, then SearchAn empty result you can explain
Cohort, then voiceSearch (filter), then Search (semantic)A population with its own words attached, no agent run
Gate before sendEval (mode: "gate"), then GET /eval-runs/{id}/gateA send/hold bit on copy an agent wrote
Pinned-evidence A/BEval, then Eval with evidence_from_run, then compareA score difference you can attribute to the edit
Eval as a trendMany Evals, then the KPI readA week-over-week measure of writing quality
Keep two verdicts apartAny two gradersA well-written claim nobody made stays visible

Optimize before send

Code that writes outbound runs the Optimizer on each draft before it is queued.

  1. Generate the final draft, merge fields included.
  2. POST /messages/optimize with the draft, in a background job with a 180 second timeout. It usually answers in under 30.
  3. On ok: true, queue message exactly as returned: a rewrite when unchanged is false, your own draft when it is true.
  4. On ok: false or a timeout, queue the draft as it was (or hold it) and log reason.

Never post-process message with another model: the checks that kept names, numbers and merge fields exact ran on that text. With no data connected this is a style pass, so it does not check claims; add evidence: "workspace" once your conversations sync, or gate on an eval (below) where claims matter. The full recipe is Optimize before you send.

Check before you query

A workspace that has not synced yet answers every search with an empty, well-formed result. GET /search/overview returns the row count and time span for each warehouse surface, with null meaning the read failed rather than zero rows.

  1. GET /search/overview. If every row_count is 0, stop and say the workspace has no data yet.
  2. Run the search. An empty result is now a real "nothing matched".

The onboarding skill runs this check first for the same reason.

Cohort, then voice

Find the population with typed filters, then read what it said, without starting an agent.

  1. POST /search/query with surface: "deals" and filters such as deal_stage_status = "won". Collect the company_id values.
  2. POST /search/query with mode: "semantic", filters: [{ "field": "company_id", "op": "in", "value": [...] }] and hydrate: true. You get the utterances from those accounts, ranked by meaning, with quotable text.

Both calls are synchronous, so the whole sequence finishes in two round trips. The semantic lane holds customer-side speech only (corpus.speaker_scope: "external_only"), which is what you want here.

Gate before send

An agent that writes outbound copy grades it before it acts.

  1. Generate the message.
  2. POST /evals/run with { "eval": "prompt-and-message-eval", "inputs": { "message": "…", "mode": "gate" } }. Gate mode grades only your copy and makes one judge call.
  3. GET /eval-runs/{id}/gate?wait_ms=30000 until status is terminal.
  4. Branch on gate.passed: true sends, false holds, null abstains and passes through.

Do not gate on the improved side, on lift, or on verdict. The gate rules say why each one misleads. On a hold, run a full eval once and regenerate from its improved prompt; never send the illustration_only message it produced.

Pinned-evidence A/B

Two eval runs each retrieve their own quotes, seeded from what you sent. Edit a draft and re-grade it and you have changed the copy and the evidence at once.

  1. Grade version A normally.
  2. Grade version B with evidence_from_run set to A's run id. B reuses A's frozen quotes, so only the copy changed.
  3. GET /eval-runs/{a}/compare/{b}. It returns a delta only when the two were graded on the same basis and the gap is larger than the judge's own spread; otherwise it names which control failed in delta_withheld_reason.

For a small edit, grade each version several times and compare medians. One run is one draw.

Eval as a trend

One run is a report card. The trend over many runs is the metric.

  1. Grade every draft that matters, or a fixed sample each week.
  2. Read GET /evals/prompt-and-message-eval/kpi?window_days=90&granularity=week. It tracks the submitted side only, abstains on windows with fewer than three scored runs, and counts refusals separately rather than as zeros.

Read eval_version before you compare windows: two versions are two instruments. A Routine can post the movement to your team each week.

Keep two verdicts apart

A message can be well written and assert something no customer ever said. A single blended number hides that.

When you combine an eval with any other grader, a house style check, or a person's review, report each verdict on its own. Within one eval, read dimensions[] line by line: Grounding and Verified specifics answer "can we back this up", and CTA clarity answers something else.

See also