Docs

Models and limits

Which models grade, draft and rewrite, how eval instruments are versioned and what you can pin, latency budgets, the two caches, data handling, and a dated table of known limits with what to do instead.

This page collects the facts about how the engine behaves that are otherwise spread across the reference pages. Last reviewed 2026-09-29.

Instrument versions

Every eval run records the version of the instrument that graded it, as eval_version on the run handle, the verdict and every export row. The current prompt-and-message-eval version is 2.35.0. The version moves whenever the grading work changes (the judge, the prompts, the evidence draw) or a stored field comes to mean something else.

  • Two versions are two instruments. Do not average scores or count verdicts across a version change. Comparing two runs refuses a pair that spans one; a trend or an export does not, so check eval_version yourself.
  • You cannot request an older version. Every new run grades under the current one. What you can hold fixed across runs is the evidence (evidence_from_run) and the graded rewrite (candidate_from_run). See Pinned-evidence A/B.
  • Training exports pin one version per file. A filter that spans two is refused with mixed_eval_versions.

The REST API itself is versioned separately, in the URL. See Versioning.

Models

You never choose a model on the Optimizer, Search or Eval. The engine uses Anthropic Claude models, and every eval run records which ones it used in grader_meta.blinding.

WhereModelNotes
Eval judge (scores both sides)claude-sonnet-5Scores your draft and the improved version in one blinded call.
Eval generator (writes the improved prompt and message)claude-sonnet-5A separate call from the judge. It scores nothing.
Eval preparation steps (reading your intent, resolving the audience, planning an outside search) and the judge grader kindclaude-haiku-4-5-20251001None of these produce the score on prompt-and-message-eval.
Chat, depth: "quick"claude-sonnet-5
Chat, depth: "standard" (the default)The agent profile's own model
Chat, depth: "deep"claude-opus-5
Optimizer judge (scores your draft and every version it tries)claude-haiku-4-5Picks the version to return.
Optimizer re-check (of the version it returns)claude-sonnet-5-5A different model from the judge. On a style pass it answers yes/no questions about the draft and the rewrite; when the run read your conversations it produces lift, which names it in lift.grader.
Search, filter and lexical lanesNoneCompiled predicates and full-text matching.
Search, fuzzy laneA model writes the SQLThe SQL it wrote comes back in compiled.sql.
Search, semantic laneAn embedding of your queryRanked against customer-side utterances.

Sampling. The eval judge and generator run at the model's default temperature, because the Claude 5-series models reject the parameter. grader_meta.blinding.sampling records what each stage actually sampled at. The practical consequence is run-to-run variance, covered under Known limits.

Customizing. There is no fine-tuning. You shape a grade with what you send (audience, account, reference_accounts, scope, artifact_type) and, when the built-in rubric is the wrong question, by authoring your own eval.

Latency and budgets

CallBudget
POST /messages/optimizeAnswers on the same call. The median in production was about 25 seconds in October 2026. A hard deadline of 170 seconds ends a run with ok: false and optimizer_timeout, so give your client 180. evidence: "workspace" is slower.
POST /search/query, synchronousSized to finish under a 15 second ceiling. A partial answer with typed fields beats a timeout.
POST /search/query, async: true (fuzzy lane)Up to about 180 seconds. You get a job_id to poll.
POST /evals/runReturns a handle at once. A full run takes minutes: the improvement loop alone is budgeted at 150 seconds. mode: "gate" makes one judge call and finishes sooner.
Eval and Chat reads with wait_msBlock for up to 30 seconds, then return the run as it stands.
Chat turns10 turns at quick, 75 at deep, the profile's own budget at standard (50 for the Master agent). A run that runs out pauses and asks whether to continue.

Rate limits are on Rate limits: per credential (per IP for a request with no credential) and per minute, 600 MCP protocol calls, 300 reads, 30 expensive calls (optimize, eval run, research start, Chat start, connect) and 60 of everything else, inside a ceiling of 300 per minute per client IP (1,200 for protocol-only requests), and 10 warehouse reads per minute per user underneath it.

Two caches

Both can hand you back an earlier answer. That is usually what you want, and it is the wrong thing when you are measuring a change.

CacheWhat it servesHow to bypass
Search result cacheAn identical warehouse query within the last five minutes, at no cost to your read budget. An implementation detail, not a contract.Change the query.
Eval run reuse (reuse: "cached", the default)A run for the same inputs that is still in flight, or completed within the last 15 minutes. reused: true on the handle says so.reuse: "force", then check reused came back false.

Data handling

What the code guarantees today:

  • Search and Eval never change your data. An eval run dispatches read operations only.
  • The Optimizer hands back text and sends nothing. It keeps a record of each run with your workspace: the draft you sent, the versions it tried and how each scored, and the per-sender record of how each sender writes. Facts you send in context are the exception: they are used for that request only, never stored as evidence and never written to logs as text.
  • Reads are scoped server-side. Every warehouse read carries your workspace filter and any data-scope rule your admin has set; your credential's scopes cap what it may call.
  • Quotes come from your corpus. A quote in a verdict is a stored utterance, looked up by id. Speaker names are not included.
  • Model calls go to Anthropic. The drafts you submit and the quotes retrieved for them are sent to Claude models to grade and rewrite.
  • Outside search is opt-in. An eval with include_external: true, or a Chat at depth: "deep" or with external_search: true, sends a search query derived from your request to an outside web search. Nothing else leaves for the web.
  • Runs are stored. Eval runs, with their verdicts and cited quotes, are kept in your workspace and are what the export reads.

For contractual terms (retention periods, subprocessors, regions), ask your Amdahl contact; this page does not state them.

Known limits

Last reviewed 2026-09-29.

LimitWhat happensDo this instead
No shared demo workspaceA workspace with nothing synced returns empty searches and not_applicable grades (empty_corpus).Start with the Optimizer, which needs no data. Run GET /search/overview before you judge Search or Eval results, and connect a CRM or call recorder.
The Optimizer's default pass does not check claimsWith no evidence, a draft is judged on how it reads and how clearly it asks. A wrong claim in the draft is still wrong in the rewrite.Send evidence: "workspace" once your conversations sync, or grade the draft with the eval.
The Optimizer's default pass prefers a question to a meeting askA rewrite usually closes on a one-line question about whether the problem is real for the reader, not a request for a call.Put the ask you need in rules, inside context. A rewrite that breaks a rule is never returned.
One draft per Optimizer call, 1,000 a monthThere is no batch call. API and connector calls are capped at 1,000 a month per workspace by default; console calls are not counted.Loop one draft at a time (see Optimize a campaign), and ask Amdahl to raise the cap before a launch that needs more.
One eval run is one drawGrading identical text twice can differ. The last measured 95th-percentile spread was 0.40 on the 0 to 1 axis (about two rubric lines), taken on the claude-sonnet-5 judge on 2026-08-05. The instrument has changed since, and the spread has not been re-measured on the current version.Grade each version several times and compare medians. The compare endpoint withholds deltas under 0.45.
lift is not a quality measureThe improved side is written to the rubric it is graded on and clears the bar on most runs, so lift mostly reflects how low your draft scored.Report headline.submitted.checks. Never gate or chart on lift.
No grading as of a past dateEvals grade against current data.Pin a run's evidence with evidence_from_run to hold its quotes fixed while you iterate.
The rubric is written for persuasive outreachScheduling notes, receipts and support replies return not_outreach. Landing copy fails "relevant positioning" unless you say what it is.Send artifact_type: "landing_copy" for web copy, or author your own eval.
Account names match as your CRM spells themA brand name that differs from the CRM record comes back not_in_corpus.Send the name exactly as the CRM records it. requires_input suggests one when it can verify it.
department is not filterableNo surface has the column.Filter on role_level.
The semantic mirror ignores excluded interactionsA semantic match can come from an interaction your workspace excluded. The filter and fuzzy lanes respect exclusions.Confirm a semantic match on the filter lane before acting on it when exclusions matter.
The lexical lane reads a 400-character previewA term past the first 400 characters of a long utterance is not findable by exact match.Use the fuzzy or filter lane with contains for long text.
A quoted phrase changes the search laneUnder mode: "auto", any query with a quoted phrase runs on the lexical lane. REST cannot force mode: "lexical"; MCP can.Read mode_ran. Drop the quotes, or set mode, to reach another lane.
The fuzzy lane reports throttling as a failed answerRunning out of warehouse reads returns 200 with detail.internal.status: "failed", not 429.Treat failed during a burst as a throttle and back off.
Unknown top-level Search keys are droppedA typo such as "filter" for "filters" is ignored, and the call runs unfiltered.Check compiled and corpus.filters_applied.
No server-side idempotencyRetrying a timed-out POST /chat or POST /routines can create a duplicate. There is no Idempotency-Key header.Deduplicate on the client. See Idempotency.
A default Chat does not reach the marketdepth: "standard" answers from your corpus only.Pass depth: "deep", or external_search: true.
No scenario projection endpointThere is no simulate or what-if operation.Measure the baseline with Search's group_by and metrics, and reason over it in a Chat.
The calendar Subscription source is pausedNew Google Calendar connections are unavailable, so calendar-triggered Chats do not fire.Use a Routine on a cron.