- model adapters for Anthropic, OpenAI, and OpenAI-compatible services;
- conversation sessions and context policies, which decide what the model sees next: scripted user messages, tool results, or custom logic;
- retries for rate limits and server errors;
- conversation records that are fxtr entities, so they’re stored, linked, and shown in the viewer.
Set up
Add behaviors, with its provider SDKs, to your project. Like fxtr, it comes from your fxtr checkout by path. Inpyproject.toml:
pyproject.toml
uv sync and commit the updated uv.lock.
Set your provider’s API key in the environment: OPENAI_API_KEY or ANTHROPIC_API_KEY. Never
store credentials in experiment inputs or entities.
Construct an adapter with the provider’s model ID, such as
OpenAIResponsesAPI("gpt-6-luna").
The OpenAI adapters also accept base_url for compatible services.
A conversation as a step
This step holds a two-turn conversation: it asks a question, then asks the model to explain its answer. The model and sampling settings arrive as a configuration entity, so every conversation points at the exact settings that produced it.chat_session_model_from_apiwraps the adapter as a conversation model. Passretry=RetryConfig()to retry rate limits, overloads, and server errors; without it, calls aren’t retried.ScriptedUserSimulatorsends one scripted message per turn and then ends the conversation.context_policy_from_user_simulatorturns it into a context policy, here starting with a system message.run_chat_rolloutalternates between the policy and the model, and returns aChatRollout: the whole conversation, with its status. The step’s output is an entity, so fxtr stores it.- Building the adapter in one helper,
make_api, lets you substitute a fake model for credential-free checks. - The adapter is closed in
finally, becauserun_chat_rolloutcloses the conversation sessions but not the provider client.
replicas={"sample": range(n)} to the
run_step call (see Replicas).
A judge step
Judging needs no conversation state, so a judge calls the model once withgenerate_with_retries.
This step grades a conversation’s final answer and returns a Verdict entity, which the viewer
can link to. Put it in the same module as the conversation step, whose make_api it uses:
judge_config: JudgeConfig and rollout: ChatRollout annotations load
the entities before the body runs:
Verdict for outcomes every attempt would reproduce, such as a conversation
that never completed or a reply without a score. It raises when the model call still fails after
its retries, so resuming the job calls the judge again.
Summarize the verdicts
A reduction step takes one model’s verdict references and loads them:Rollouts, settings, and testing
Check a rollout's status before scoring it
Check a rollout's status before scoring it
A rollout ends as
RolloutCompleted, RolloutMaxTurns, or RolloutFailed. max_turns
counts model turns: a script with N prompts needs max_turns=N + 1 to end as completed. A
model stop reason such as StopMaxTokens is separate from the rollout’s status, and may mean
the content was truncated even in a completed rollout.Recorded failures are cached as results
Recorded failures are cached as results
By default, a rollout that exceeds the model’s context length is returned as
RolloutFailed,
keeping the conversation so far; other failures raise. A returned failed rollout is a
successful step result, so it’s cached. Decide deliberately how scoring treats such outcomes.Sampling settings vary by provider
Sampling settings vary by provider
SamplingParams covers max_tokens, temperature, top_p, stop_sequences, reasoning,
and provider-specific extra. Which settings take effect depends on the adapter and model:
for example, OpenAIResponsesAPI doesn’t apply stop_sequences. Leave a setting unset
unless your adapter and model support it.Test without credentials
Test without credentials
Any object with an async
generate(request) method returning a ModelResponse, and an
aclose() method, can stand in for a provider adapter. Return it from make_api in tests, and
run against a scratch schema.Tools and custom conversation flows
Tools and custom conversation flows
context_policy_from_user_and_tools combines a user simulator with a tool policy you
implement, which defines the tools and executes their calls. For fully custom flows,
context_policy_from_generator turns an async generator into a context policy: it yields each
turn’s messages and receives the model’s reply.