Skip to main content
This tutorial builds a small experiment in the project from the Quickstart. It measures the length of each prompt in a dataset, then averages the lengths. There are no model calls, so it runs instantly, but it has the same shape as most real experiments: a step mapped over a dataset, followed by a step that reduces the results.

Write the experiment

Create src/my_project/word_count.py:
src/my_project/word_count.py
A few things to notice:
  • Steps do the work. count_words takes one prompt and returns one number. mean_length takes a whole array of counts, indexed by the prompt dimension, and returns one number.
  • The workflow connects them. run_step schedules a step and immediately returns a handle to its future result, without waiting for it. Passing counts to the second step is what makes it depend on the first.
  • map_over=["prompt"] makes a sweep. The runner calls count_words once per prompt and stacks the results into an array with the same prompt dimension. The second run_step has no map_over, so mean_length runs once and receives the whole array.
  • Annotations declare array shapes. A plain str or int means a single value. Annotated[..., "[prompt: str]"] names the dimensions, which a Python type alone can’t say.
  • Docstrings show up in the viewer, on each step’s and workflow’s card.
Then register the module in pyproject.toml, so that fxtr can find the experiment by name when it resumes a job or shows it in the viewer:
pyproject.toml

Launch it

Create a launcher at the project root. It builds the input dataset and runs the workflow as a job:
launch_word_count.py
Array.from_items takes (key, value) pairs and a type: here, one prompt dimension with string keys, holding string values. root names the job’s root workflow, and, as you’ll see below, decides which earlier results the job can reuse. Commit, then launch. A launch records the commit it runs, so it refuses uncommitted changes:
With uv run fxtr view running, open the printed link to see the job: the prompts input feeding three count_words calls, one per prompt, feeding mean_length.

Rerun it

Launch it again, unchanged. The job succeeds straight away, because every step finds its result in the cache. You can see the entries with fxtr cache list:
Each entry is at a cache address built from the job’s root, the step’s name in the workflow, and, for a mapped step, the key it was called with. A step reuses a cached result when its address holds a result computed from the same step and the same inputs.

Change an input

Now edit the question prompt, commit, and launch again under the same root. The job stops:
The address word-count-v1/count_words[prompt:"question"] already holds a result computed from the old prompt. Rather than silently reuse a result that no longer matches, or overwrite a result an earlier job depends on, fxtr holds the step and stops the job once nothing else can run. The unchanged prompts still reuse their results. run_job raises JobNotFinishedError for any job that doesn’t succeed, so your launcher can tell. You have two ways forward:
Change root="word-count-v1" to root="word-count-v2" and launch again. Every address in the new job is fresh, so every step runs. This is the usual choice when you change the experiment’s data or settings, and it keeps the old results intact.
Changing a step’s code does not invalidate its cached results: the cache matches a step by its name and inputs, not its implementation. After changing what a step computes, launch under a new root or clear its entries. See Caching and reruns.

Working with a coding agent

A new project includes fxtr’s guides for coding agents in .agents/skills/, where Codex finds them, linked from .claude/skills/, where Claude Code finds them. They cover writing steps and workflows, launching jobs, caching, calling models with behaviors, and building viewer renderers, so you can ask an agent for an experiment in plain language and review the result in the viewer. Commit them with the project. After updating your fxtr checkout, refresh them from the project:
This replaces fxtr’s own guides and leaves any other skills in the project alone.

Next steps

Mapping and sweeps

Sweep over models, prompts, and samples, and reduce over any dimension.

Calling language models

Use behaviors to hold conversations and judge outputs in your steps.