Skip to main content
behaviors is a Python library for calling language models from fxtr experiments. It provides:
  • model adapters for Anthropic, OpenAI, and OpenAI-compatible services;
  • conversation sessions and context policies, which decide what the model sees next: scripted user messages, tool results, or custom logic;
  • retries for rate limits and server errors;
  • conversation records that are fxtr entities, so they’re stored, linked, and shown in the viewer.

Set up

Add behaviors, with its provider SDKs, to your project. Like fxtr, it comes from your fxtr checkout by path. In pyproject.toml:
pyproject.toml
Then run uv sync and commit the updated uv.lock. Set your provider’s API key in the environment: OPENAI_API_KEY or ANTHROPIC_API_KEY. Never store credentials in experiment inputs or entities. Construct an adapter with the provider’s model ID, such as OpenAIResponsesAPI("gpt-6-luna"). The OpenAI adapters also accept base_url for compatible services.

A conversation as a step

This step holds a two-turn conversation: it asks a question, then asks the model to explain its answer. The model and sampling settings arrive as a configuration entity, so every conversation points at the exact settings that produced it.
The pieces:
  • chat_session_model_from_api wraps the adapter as a conversation model. Pass retry=RetryConfig() to retry rate limits, overloads, and server errors; without it, calls aren’t retried.
  • ScriptedUserSimulator sends one scripted message per turn and then ends the conversation. context_policy_from_user_simulator turns it into a context policy, here starting with a system message.
  • run_chat_rollout alternates between the policy and the model, and returns a ChatRollout: the whole conversation, with its status. The step’s output is an entity, so fxtr stores it.
  • Building the adapter in one helper, make_api, lets you substitute a fake model for credential-free checks.
  • The adapter is closed in finally, because run_chat_rollout closes the conversation sessions but not the provider client.
Launch it with a stored configuration per model:
For repeated samples of each conversation, add replicas={"sample": range(n)} to the run_step call (see Replicas).

A judge step

Judging needs no conversation state, so a judge calls the model once with generate_with_retries. This step grades a conversation’s final answer and returns a Verdict entity, which the viewer can link to. Put it in the same module as the conversation step, whose make_api it uses:
Schedule it in the workflow, mapped over the conversations, with the judge configuration as a single stored input. The judge_config: JudgeConfig and rollout: ChatRollout annotations load the entities before the body runs:
Notice how the judge follows the return or raise rule. It returns a Verdict for outcomes every attempt would reproduce, such as a conversation that never completed or a reply without a score. It raises when the model call still fails after its retries, so resuming the job calls the judge again.

Summarize the verdicts

A reduction step takes one model’s verdict references and loads them:

Things to know

A rollout ends as RolloutCompleted, RolloutMaxTurns, or RolloutFailed. max_turns counts model turns: a script with N prompts needs max_turns=N + 1 to end as completed. A model stop reason such as StopMaxTokens is separate from the rollout’s status, and may mean the content was truncated even in a completed rollout.
By default, a rollout that exceeds the model’s context length is returned as RolloutFailed, keeping the conversation so far; other failures raise. A returned failed rollout is a successful step result, so it’s cached. Decide deliberately how scoring treats such outcomes.
SamplingParams covers max_tokens, temperature, top_p, stop_sequences, reasoning, and provider-specific extra. Which settings take effect depends on the adapter and model: for example, OpenAIResponsesAPI doesn’t apply stop_sequences. Leave a setting unset unless your adapter and model support it.
Any object with an async generate(request) method returning a ModelResponse, and an aclose() method, can stand in for a provider adapter. Return it from make_api in tests, and run against a scratch schema.
context_policy_from_user_and_tools combines a user simulator with a tool policy you implement, which defines the tools and executes their calls. For fully custom flows, context_policy_from_generator turns an async generator into a context policy: it yields each turn’s messages and receives the model’s reply.