Skip to main content

The mental model

An fxtr experiment has two halves:
  • The experiment code is a program describing the computation you want to run: which steps to call, on which data, in what order.
  • The experiment record is what actually happened when it ran: which data each step received, what it returned, and which workflow asked for it.
When you launch an experiment, fxtr runs the code and builds the record as it goes, as a graph connecting every input, operation, and result. The record lives in the project’s database, so you can inspect it in the viewer, resume it after an interruption, and reuse its expensive results in later runs.

Glossary

A uv project whose Python package defines steps and workflows. It’s configured by [tool.fxtr] in pyproject.toml, and by fxtr.local.toml, which says where this machine reaches the project’s Postgres database.
One unit of computation (@step), such as a model call, a judgment, or a reduction. A step receives concrete data, returns a result, and can’t schedule further work. Its result is cached. See Steps and workflows.
Deterministic code (@workflow) that schedules steps and child workflows, building the experiment’s graph. A workflow works with handles to data rather than the data itself, and can wait for a result when it needs one to decide what to do next.
A collection of values of one schema, indexed by named dimensions such as model or prompt. Every input and output of a step or workflow is an array; a single value is an array with no dimensions. See Arrays and entities.
A workflow’s reference to an array that may not be computed yet, such as a step’s future result. Its dimensions and schema are known immediately, so the workflow can pass it on to further steps before the data exists.
A stored record, such as a conversation or a judge’s verdict, with a type and an ID derived from its content. Arrays hold references to entities, and the viewer links to them.
One durable execution of a root workflow on a set of input arrays. A job records its progress as it runs, so it can be inspected while running and resumed if it stops. See Running and managing jobs.
A human-readable address for each workflow, step, and source array within a job, such as root/evaluate/sample[task:"a"]. It is built from the names you give each operation.
Where a step’s result is stored: its pathname, unless you override it. A step reuses a cached result when its address holds one computed from the same step and inputs. See Caching and reruns.
A web app, started with fxtr view, that shows a project’s jobs: each workflow’s graph, the arrays flowing through it, and the entities they reference, through renderers you can customize. See Viewing results.

Why steps and workflows are separate

Splitting an experiment into workflows and steps is what makes it both durable and inspectable:
  • Steps are the expensive, nondeterministic part: sampling a model, calling an API, drawing a random number. fxtr stores every step’s result, so none of this work is repeated when a job resumes or a later job asks for the same thing.
  • Workflows are cheap, deterministic glue. Because a workflow makes the same decisions given the same inputs and results, fxtr can replay it against the recorded results to rebuild the graph exactly, and check that it does.
As a result, fxtr can always tell which step produced a value, what that step received, and which workflow asked for it.