> ## Documentation Index
> Fetch the complete documentation index at: https://docs.transluce.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Steps and workflows

> Write the functions that compute results and the workflows that connect them.

## Steps

A **step** is an async function decorated with `@step`. It receives concrete data, does one piece
of work, and returns a result, which fxtr stores and caches.

```python theme={null}
from fxtr.experiment.steps import StepContext, step


@step(name="word_count.count")
async def count_words(context: StepContext, prompt: str) -> int:
    """Count the whitespace-separated words in a prompt."""
    return len(prompt.split())
```

Use a step for anything expensive or nondeterministic: model calls, judging, external I/O,
sampling. Use one too for a deterministic transformation or reduction whose dependencies you
want to see in the viewer. Steps can store and load entities and write log messages with
`context.log(...)`, but they can't schedule further work.

## Workflows

A **workflow** is an async function decorated with `@workflow`. It schedules steps and child
workflows, and so builds the experiment's graph:

```python theme={null}
from typing import Annotated

from fxtr.experiment.handles import ArrayHandle
from fxtr.experiment.workflows import WorkflowContext, workflow


@workflow(name="word_count.experiment")
async def experiment(
    context: WorkflowContext,
    prompts: Annotated[ArrayHandle[str], "[prompt: str]"],
) -> Annotated[ArrayHandle[int], "[prompt: str]"]:
    """Measure the length of each prompt."""
    return context.run_step("count_words", count_words, {"prompt": prompts}, map_over=["prompt"])
```

`context.run_step(step_name, step, inputs, ...)` and
`context.run_child_workflow(workflow_name, workflow, inputs, ...)` return a **handle** to the
result immediately, without waiting for the work. Any workflow can be the root of a job, and a
workflow can call others as children, each building its own graph.

### Rules for both

* The first parameter must be named `context`, typed `StepContext` or `WorkflowContext`.
* Give each step and workflow a stable registered name with `name=`, such as
  `"word_count.count"`. It defaults to the function's module and qualified name, so setting it
  keeps it stable if you move the function.
* Put the module that defines them in `[tool.fxtr] modules`, so that resuming a job, `fxtr run`,
  and the viewer can find them.
* The docstring's first sentence appears on the function's card in the viewer. Make it a short
  imperative sentence, such as "Compute mean scores across judges."
* Don't call a decorated function directly. `await count_words(context, "hi")` runs only its
  Python body, skipping scheduling and caching. Always schedule it through a workflow context.

## Signatures

A signature describes the arrays one call accepts and returns. Annotations are the usual way to
write one:

| Annotation | Means |
| - | - |
| `prompt: str`, `-> int` | A single value. The step body receives and returns the bare value. |
| `Annotated[Array[float], "[sample: int]"]` | A concrete array with a `sample` dimension, in a step. |
| `Annotated[ArrayHandle[str], "[prompt: str]"]` | A handle with a `prompt` dimension, in a workflow. |
| `ScalarHandle[str]` | A handle to a single value; short for `Annotated[ArrayHandle[str], "[]"]`. |

`Array[T]` or `ArrayHandle[T]` alone doesn't say which dimensions the array has, so always wrap
them in `Annotated[..., "[...]"]`.

For an entity input, the annotation decides what the step receives:

* `document: Document` loads the entity before calling the body.
* `document: BoundID[Document]` passes a reference without loading it.
* `Annotated[Array[BoundID[Document]], "[doc: str]"]` passes references for a whole slice.

A step declared `-> Document` can simply return a `Document(...)`: fxtr stores it and records its
reference. A step with a dimensioned output returns an `Array`.

## Handles

A workflow works with **handles**: references to arrays whose type is known before their data
exists. Handles let a workflow describe the whole graph up front, and they're how the viewer knows
what depends on what.

| Operation | Purpose |
| - | - |
| `context.add_source_array(name, array)` | Add fixed data, such as configuration, to the workflow's graph. |
| `handle.field("score")` | Select one field of a struct-valued array, keeping its dimensions. |
| `handle.sel({"model": "baseline"})` | Select keys along some dimensions, removing those dimensions. |
| `handle.rename({"sample": "judge_sample"})` | Rename dimensions. |
| `context.combine({"answer": answers, "score": scores})` | Join arrays into one struct-valued array. |
| `context.stack({"baseline": a, "variant": b}, dim="condition")` | Stack same-typed arrays along a new dimension. |
| `handle.expect_type(("[task: str]", float))` | Check the handle's type while building the graph. |
| `await handle.observe()` | Wait for the concrete array. |
| `await handle.item()` | Wait for a single value and return it. |

### Data-dependent control flow

A workflow can wait for a result and decide what to schedule next. For example, an iterative
refinement loop keeps the current result's handle, observes a scalar stopping condition, and
passes the handle on to the next round:

```python theme={null}
current = draft
for i in range(max_rounds):
    score = context.run_step(f"score_{i}", score_draft, {"draft": current})
    if await score.item() >= target:
        break
    current = context.run_step(f"refine_{i}", refine_draft, {"draft": current})
return current
```

Give each round a stable name, such as `refine_0` and `refine_1`, so that replaying the workflow
rebuilds the same sequence.

<Tip>
  Observe data to make decisions, not to transform it. If you observe an array, change it in
  Python, and add the result back as a source array, the viewer loses the link between the two.
  Put the transformation in a step that takes the original handle instead.
</Tip>

## Keep workflows deterministic

fxtr replays workflow code against recorded results, both when resuming a job and to check the
workflow: by default it runs each workflow body a second time and verifies that it makes the same
decisions. So a workflow must be cheap and deterministic:

* Do expensive or random work in steps.
* Fix datasets and settings before the job starts, and pass them in as inputs.
* Never generate random data, UUIDs, or timestamps inside a workflow.
* Don't read files or call external services from a workflow.

## Return or raise?

A step's return value is its result, cached like any other. A step that raises caches nothing:
work that doesn't depend on it still runs, but the job ends without succeeding, and resuming the
job runs the step again.

So **return** data for an outcome every attempt would reproduce, such as an input with nothing to
grade or a reply without the expected answer. **Raise** for a failure a later attempt might not
hit, such as a network error or an exhausted retry budget.

<Note>
  Caching doesn't make side effects happen exactly once. An attempt that's interrupted before its
  result is saved can run again.
</Note>
