Write the experiment
Createsrc/my_project/word_count.py:
src/my_project/word_count.py
- Steps do the work.
count_wordstakes one prompt and returns one number.mean_lengthtakes a whole array of counts, indexed by thepromptdimension, and returns one number. - The workflow connects them.
run_stepschedules a step and immediately returns a handle to its future result, without waiting for it. Passingcountsto the second step is what makes it depend on the first. map_over=["prompt"]makes a sweep. The runner callscount_wordsonce per prompt and stacks the results into an array with the samepromptdimension. The secondrun_stephas nomap_over, somean_lengthruns once and receives the whole array.- Annotations declare array shapes. A plain
strorintmeans a single value.Annotated[..., "[prompt: str]"]names the dimensions, which a Python type alone can’t say. - Docstrings show up in the viewer, on each step’s and workflow’s card.
pyproject.toml, so that fxtr can find the experiment by name when
it resumes a job or shows it in the viewer:
pyproject.toml
Launch it
Create a launcher at the project root. It builds the input dataset and runs the workflow as a job:launch_word_count.py
Array.from_items takes (key, value) pairs and a type: here, one prompt dimension with string
keys, holding string values. root names the job’s root workflow, and, as you’ll see below,
decides which earlier results the job can reuse.
Commit, then launch. A launch records the commit it runs, so it refuses uncommitted changes:
uv run fxtr view running, open the printed link to see the job: the prompts input feeding
three count_words calls, one per prompt, feeding mean_length.
Rerun it
Launch it again, unchanged. The job succeeds straight away, because every step finds its result in the cache. You can see the entries withfxtr cache list:
root, the step’s name in the
workflow, and, for a mapped step, the key it was called with. A step reuses a cached result when
its address holds a result computed from the same step and the same inputs.
Change an input
Now edit thequestion prompt, commit, and launch again under the same root. The job stops:
word-count-v1/count_words[prompt:"question"] already holds a result computed from
the old prompt. Rather than silently reuse a result that no longer matches, or overwrite a result
an earlier job depends on, fxtr holds the step and stops the job once nothing else can run.
The unchanged prompts still reuse their results. run_job raises JobNotFinishedError for any
job that doesn’t succeed, so your launcher can tell.
You have two ways forward:
- Use a new root
- Clear and resume
Change
root="word-count-v1" to root="word-count-v2" and launch again. Every address in
the new job is fresh, so every step runs. This is the usual choice when you change the
experiment’s data or settings, and it keeps the old results intact.Working with a coding agent
A new project includes fxtr’s guides for coding agents in.agents/skills/, where Codex finds
them, linked from .claude/skills/, where Claude Code finds them. They cover writing steps and
workflows, launching jobs, caching, calling models with behaviors, and building viewer renderers,
so you can ask an agent for an experiment in plain language and review the result in the viewer.
Commit them with the project.
After updating your fxtr checkout, refresh them from the project:
Next steps
Mapping and sweeps
Sweep over models, prompts, and samples, and reduce over any dimension.
Calling language models
Use behaviors to hold conversations and judge outputs in your steps.