Skip to main content

Steps

A step is an async function decorated with @step. It receives concrete data, does one piece of work, and returns a result, which fxtr stores and caches.
Use a step for anything expensive or nondeterministic: model calls, judging, external I/O, sampling. Use one too for a deterministic transformation or reduction whose dependencies you want to see in the viewer. Steps can store and load entities and write log messages with context.log(...), but they can’t schedule further work.

Workflows

A workflow is an async function decorated with @workflow. It schedules steps and child workflows, and so builds the experiment’s graph:
context.run_step(step_name, step, inputs, ...) and context.run_child_workflow(workflow_name, workflow, inputs, ...) return a handle to the result immediately, without waiting for the work. Any workflow can be the root of a job, and a workflow can call others as children, each building its own graph.

Rules for both

  • The first parameter must be named context, typed StepContext or WorkflowContext.
  • Give each step and workflow a stable registered name with name=, such as "word_count.count". It defaults to the function’s module and qualified name, so setting it keeps it stable if you move the function.
  • Put the module that defines them in [tool.fxtr] modules, so that resuming a job, fxtr run, and the viewer can find them.
  • The docstring’s first sentence appears on the function’s card in the viewer. Make it a short imperative sentence, such as “Compute mean scores across judges.”
  • Don’t call a decorated function directly. await count_words(context, "hi") runs only its Python body, skipping scheduling and caching. Always schedule it through a workflow context.

Signatures

A signature describes the arrays one call accepts and returns. Annotations are the usual way to write one: Array[T] or ArrayHandle[T] alone doesn’t say which dimensions the array has, so always wrap them in Annotated[..., "[...]"]. For an entity input, the annotation decides what the step receives:
  • document: Document loads the entity before calling the body.
  • document: BoundID[Document] passes a reference without loading it.
  • Annotated[Array[BoundID[Document]], "[doc: str]"] passes references for a whole slice.
A step declared -> Document can simply return a Document(...): fxtr stores it and records its reference. A step with a dimensioned output returns an Array.

Handles

A workflow works with handles: references to arrays whose type is known before their data exists. Handles let a workflow describe the whole graph up front, and they’re how the viewer knows what depends on what.

Data-dependent control flow

A workflow can wait for a result and decide what to schedule next. For example, an iterative refinement loop keeps the current result’s handle, observes a scalar stopping condition, and passes the handle on to the next round:
Give each round a stable name, such as refine_0 and refine_1, so that replaying the workflow rebuilds the same sequence.
Observe data to make decisions, not to transform it. If you observe an array, change it in Python, and add the result back as a source array, the viewer loses the link between the two. Put the transformation in a step that takes the original handle instead.

Keep workflows deterministic

fxtr replays workflow code against recorded results, both when resuming a job and to check the workflow: by default it runs each workflow body a second time and verifies that it makes the same decisions. So a workflow must be cheap and deterministic:
  • Do expensive or random work in steps.
  • Fix datasets and settings before the job starts, and pass them in as inputs.
  • Never generate random data, UUIDs, or timestamps inside a workflow.
  • Don’t read files or call external services from a workflow.

Return or raise?

A step’s return value is its result, cached like any other. A step that raises caches nothing: work that doesn’t depend on it still runs, but the job ends without succeeding, and resuming the job runs the step again. So return data for an outcome every attempt would reproduce, such as an input with nothing to grade or a reply without the expected answer. Raise for a failure a later attempt might not hit, such as a network error or an exhausted retry budget.
Caching doesn’t make side effects happen exactly once. An attempt that’s interrupted before its result is saved can run again.