> ## Documentation Index
> Fetch the complete documentation index at: https://docs.transluce.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Calling language models

> Use the behaviors library to hold conversations and judge outputs in fxtr steps.

**behaviors** is a Python library for calling language models from fxtr experiments. It provides:

* model adapters for Anthropic, OpenAI, and OpenAI-compatible services;
* conversation sessions and **context policies**, which decide what the model sees next:
  scripted user messages, tool results, or custom logic;
* retries for rate limits and server errors;
* conversation records that are fxtr entities, so they're stored, linked, and shown in the
  viewer.

## Set up

Add behaviors, with its provider SDKs, to your project. Like fxtr, it comes from your fxtr
checkout by path. In `pyproject.toml`:

```toml pyproject.toml theme={null}
[project]
dependencies = ["fxtr[runner,debug-web]", "behaviors[models]"]

[tool.uv.sources]
fxtr = { path = "/path/to/fxtr/fxtr", editable = true }
behaviors = { path = "/path/to/fxtr/behaviors", editable = true }
```

Then run `uv sync` and commit the updated `uv.lock`.

Set your provider's API key in the environment: `OPENAI_API_KEY` or `ANTHROPIC_API_KEY`. Never
store credentials in experiment inputs or entities.

| Provider | Adapter |
| - | - |
| OpenAI Responses | `from behaviors.implementations.models.openai_responses import OpenAIResponsesAPI` |
| Anthropic Messages | `from behaviors.implementations.models.anthropic import AnthropicMessagesAPI` |
| OpenAI Chat Completions-compatible | `from behaviors.implementations.models.openai_compatible import OpenAICompatibleAPI` |

Construct an adapter with the provider's model ID, such as `OpenAIResponsesAPI("gpt-6-luna")`.
The OpenAI adapters also accept `base_url` for compatible services.

## A conversation as a step

This step holds a two-turn conversation: it asks a question, then asks the model to explain its
answer. The model and sampling settings arrive as a configuration entity, so every conversation
points at the exact settings that produced it.

```python theme={null}
from dataclasses import dataclass
from typing import Annotated, cast

import anyio

from behaviors.implementations.context_policies import (
    ScriptedUserSimulator,
    context_policy_from_user_simulator,
)
from behaviors.implementations.models import chat_session_model_from_api
from behaviors.implementations.models.openai_responses import OpenAIResponsesAPI
from behaviors.rollouts.chat import run_chat_rollout
from behaviors.types import (
    ChatRollout, ContentText, RetryConfig, SamplingParams, SetSystemMessage,
    SystemMessage, UserMessage,
)
from fxtr.core.entities import BoundID
from fxtr.entity_defns.dataclass_entity import DataclassEntity
from fxtr.experiment.handles import ArrayHandle
from fxtr.experiment.steps import StepContext, step
from fxtr.experiment.workflows import WorkflowContext, workflow


@dataclass(frozen=True)
class ConversationConfig(DataclassEntity, fxtr_type="org.example.my_experiment.ConversationConfig.v1"):
    model_name: str
    sampling: SamplingParams
    system_prompt: str


def make_api(model_name: str) -> OpenAIResponsesAPI:
    return OpenAIResponsesAPI(model_name)


@step(name="my_experiment.conversation")
async def converse(context: StepContext, config: ConversationConfig, prompt: str) -> ChatRollout:
    """Answer a prompt, then explain the answer in a follow-up turn."""
    api = make_api(config.model_name)
    try:
        model = chat_session_model_from_api(api, config.sampling, retry=RetryConfig())
        user = ScriptedUserSimulator(messages=(
            UserMessage(content=(ContentText(text=prompt),)),
            UserMessage(content=(ContentText(text="Explain your answer."),)),
        ))
        policy = context_policy_from_user_simulator(
            user,
            initialization=SetSystemMessage(
                message=SystemMessage(content=(ContentText(text=config.system_prompt),)),
            ),
        )
        # One more turn than there are scripted prompts, so the rollout can end as completed.
        return await run_chat_rollout(model, policy, max_turns=3)
    finally:
        with anyio.CancelScope(shield=True):
            await api.aclose()


@workflow(name="my_experiment.run")
async def experiment(
    context: WorkflowContext,
    configs: Annotated[ArrayHandle[BoundID[ConversationConfig]], "[model: str]"],
    prompts: Annotated[ArrayHandle[str], "[prompt: str]"],
) -> Annotated[ArrayHandle[BoundID[ChatRollout]], "[model: str, prompt: str]"]:
    """Hold one conversation per model and prompt."""
    rollouts = context.run_step(
        "conversations",
        converse,
        {"config": configs, "prompt": prompts},
        map_over=["model", "prompt"],
    )
    return cast(ArrayHandle[BoundID[ChatRollout]], rollouts)
```

The pieces:

* `chat_session_model_from_api` wraps the adapter as a conversation model. Pass
  `retry=RetryConfig()` to retry rate limits, overloads, and server errors; without it, calls
  aren't retried.
* `ScriptedUserSimulator` sends one scripted message per turn and then ends the conversation.
  `context_policy_from_user_simulator` turns it into a context policy, here starting with a system
  message.
* `run_chat_rollout` alternates between the policy and the model, and returns a `ChatRollout`:
  the whole conversation, with its status. The step's output is an entity, so fxtr stores it.
* Building the adapter in one helper, `make_api`, lets you substitute a fake model for
  credential-free checks.
* The adapter is closed in `finally`, because `run_chat_rollout` closes the conversation
  sessions but not the provider client.

Launch it with a stored configuration per model:

```python theme={null}
import anyio

from fxtr.entity_defns.array import Array
from fxtr.project.running import open_local_client, run_job
from my_project.conversations import ConversationConfig, experiment


async def main() -> None:
    async with open_local_client(__file__) as client:
        luna = await client.store(ConversationConfig(
            model_name="gpt-6-luna",
            sampling=SamplingParams(max_tokens=2000),
            system_prompt="Answer concisely.",
        ))
        configs = Array.from_items([(("luna",), luna)], ("[model: str]", BoundID[ConversationConfig]))
        prompts = Array.from_items([(("greeting",), "Hello!")], ("[prompt: str]", str))
        await run_job(client, experiment, {"configs": configs, "prompts": prompts}, root="conv-v1")


if __name__ == "__main__":
    anyio.run(main)
```

For repeated samples of each conversation, add `replicas={"sample": range(n)}` to the
`run_step` call (see [Replicas](/fxtr/concepts/mapping#replicas)).

## A judge step

Judging needs no conversation state, so a judge calls the model once with `generate_with_retries`.
This step grades a conversation's final answer and returns a `Verdict` entity, which the viewer
can link to. Put it in the same module as the conversation step, whose `make_api` it uses:

```python theme={null}
import re
from dataclasses import dataclass
from typing import Annotated

import anyio

from behaviors.implementations.models import (
    DEFAULT_FAILURE_CATEGORIES_TO_RETRY,
    generate_with_retries,
)
from behaviors.types import (
    AssistantMessage, ChatMessage, ChatRollout, ContentText, ModelCallFailure,
    ModelRequest, RetryConfig, RolloutCompleted, SamplingParams, UserMessage,
)
from fxtr.core.entities import BoundID
from fxtr.entity_defns.array import Array
from fxtr.entity_defns.dataclass_entity import DataclassEntity
from fxtr.experiment.steps import StepContext, step


@dataclass(frozen=True)
class JudgeConfig(DataclassEntity, fxtr_type="org.example.my_experiment.JudgeConfig.v1"):
    model_name: str
    sampling: SamplingParams
    rubric: str


@dataclass(frozen=True)
class Verdict(DataclassEntity, fxtr_type="org.example.my_experiment.Verdict.v1"):
    score: int | None  # None when there was nothing to grade or the reply had no score
    judge_reply: str


def answer_text(messages: tuple[ChatMessage, ...]) -> str:
    """The assistant's text, without reasoning or tool-use blocks."""
    return "".join(
        block.text
        for message in messages
        if isinstance(message, AssistantMessage)
        for block in message.content
        if isinstance(block, ContentText)
    )


@step(name="my_experiment.judge")
async def judge(context: StepContext, judge_config: JudgeConfig, rollout: ChatRollout) -> Verdict:
    """Grade the conversation's final answer against the rubric from 0 to 10."""
    if not isinstance(rollout.status, RolloutCompleted) or not rollout.model_turns:
        return Verdict(score=None, judge_reply=f"not graded: {rollout.status.kind}")
    prompt = (
        f"{judge_config.rubric}\n\n"
        f"Answer to grade:\n{answer_text(rollout.model_turns[-1].messages)}\n\n"
        "End your reply with a line of the form 'Score: N', where N is from 0 to 10."
    )
    request = ModelRequest(
        context=(UserMessage(content=(ContentText(text=prompt),)),),
        tools=(),
        tool_choice=None,
        parallel_tool_calls=None,
        sampling=judge_config.sampling,
    )
    api = make_api(judge_config.model_name)
    try:
        response = await generate_with_retries(
            api,
            request,
            retry=RetryConfig(),
            failure_categories_to_retry=DEFAULT_FAILURE_CATEGORIES_TO_RETRY,
        )
    finally:
        with anyio.CancelScope(shield=True):
            await api.aclose()
    if isinstance(response, ModelCallFailure):
        # Raise rather than return: a returned failure would be cached as the result.
        raise RuntimeError(f"judge call failed ({response.category}): {response.description}")
    reply = answer_text(response.output)
    match = re.search(r"Score:\s*(\d+)\s*$", reply)
    score = int(match.group(1)) if match else None
    return Verdict(score=score if score is not None and 0 <= score <= 10 else None, judge_reply=reply)
```

Schedule it in the workflow, mapped over the conversations, with the judge configuration as a
single stored input. The `judge_config: JudgeConfig` and `rollout: ChatRollout` annotations load
the entities before the body runs:

```python theme={null}
verdicts = context.run_step(
    "judge_final_answers",
    judge,
    {"judge_config": judge_config, "rollout": rollouts},
    map_over=["model", "prompt"],
)
```

Notice how the judge follows the [return or raise](/fxtr/concepts/steps-and-workflows#return-or-raise)
rule. It returns a `Verdict` for outcomes every attempt would reproduce, such as a conversation
that never completed or a reply without a score. It raises when the model call still fails after
its retries, so resuming the job calls the judge again.

## Summarize the verdicts

A reduction step takes one model's verdict references and loads them:

```python theme={null}
@step(name="my_experiment.mean_score")
async def mean_score(
    context: StepContext,
    verdicts: Annotated[Array[BoundID[Verdict]], "[prompt: str]"],
) -> float | None:
    """Average the judge's scores across prompts."""
    scores: list[int] = []
    for ref in verdicts.values():
        verdict = await context.load(ref)
        if verdict.score is not None:
            scores.append(verdict.score)
    return sum(scores) / len(scores) if scores else None


# In the workflow:
mean_scores = context.run_step("mean_scores", mean_score, {"verdicts": verdicts}, map_over=["model"])
```

## Things to know

<AccordionGroup>
  <Accordion title="Check a rollout's status before scoring it">
    A rollout ends as `RolloutCompleted`, `RolloutMaxTurns`, or `RolloutFailed`. `max_turns`
    counts model turns: a script with N prompts needs `max_turns=N + 1` to end as completed. A
    model stop reason such as `StopMaxTokens` is separate from the rollout's status, and may mean
    the content was truncated even in a completed rollout.
  </Accordion>

  <Accordion title="Recorded failures are cached as results">
    By default, a rollout that exceeds the model's context length is returned as `RolloutFailed`,
    keeping the conversation so far; other failures raise. A returned failed rollout is a
    successful step result, so it's cached. Decide deliberately how scoring treats such outcomes.
  </Accordion>

  <Accordion title="Sampling settings vary by provider">
    `SamplingParams` covers `max_tokens`, `temperature`, `top_p`, `stop_sequences`, `reasoning`,
    and provider-specific `extra`. Which settings take effect depends on the adapter and model:
    for example, `OpenAIResponsesAPI` doesn't apply `stop_sequences`. Leave a setting unset
    unless your adapter and model support it.
  </Accordion>

  <Accordion title="Test without credentials">
    Any object with an async `generate(request)` method returning a `ModelResponse`, and an
    `aclose()` method, can stand in for a provider adapter. Return it from `make_api` in tests, and
    run against a [scratch schema](/fxtr/guides/running-jobs#before-you-launch).
  </Accordion>

  <Accordion title="Tools and custom conversation flows">
    `context_policy_from_user_and_tools` combines a user simulator with a tool policy you
    implement, which defines the tools and executes their calls. For fully custom flows,
    `context_policy_from_generator` turns an async generator into a context policy: it yields each
    turn's messages and receives the model's reply.
  </Accordion>
</AccordionGroup>
