Why fxtr
Experiments on AI systems tend to sprawl. A few scripts become dozens, JSON files pile up on disk, and it gets hard to say which version of the code produced which results, or to check that a judge did something sensible on the fiftieth sample. Coding agents make this worse: they can write an experiment in minutes, but reviewing what they wrote, and what it produced, takes much longer. fxtr keeps an experiment organized by default:Results you can trace
Every result is stored in a database alongside the step that computed it, the inputs it
received, and the code commit the job ran at. The viewer lets you move from a chart to the
records behind it to the code that produced them.
Less code to review
Execution, storage, and visualization are handled for you, so the code left to read is the
experiment itself: what each step computes and how the steps connect.
Caching you can trust
Model calls are expensive. fxtr reuses a cached result only when the step and its inputs
match, and stops rather than silently reusing a result computed from different inputs.
Durable jobs
A job records its progress as it runs. If it stops, whether from a crash, a failed API call,
or a cancellation, resuming it picks up where it left off.
How it fits together
An fxtr project is a Python package of steps and workflows:- A step does one piece of computation, such as a model call, a judgment, or a reduction. Its result is cached.
- A workflow schedules steps and child workflows over arrays: collections of values
indexed by named dimensions such as
model,prompt, orsample. Mapping a step over those dimensions is how you express a sweep.
Built for working with coding agents
fxtr ships guides for coding agents alongside the library. A new project includes them, so Claude Code and Codex know how to write steps and workflows, launch jobs, and build viewer renderers in your project. See Your first experiment.Quickstart
Create a project and launch its first job.
Core concepts
Projects, steps, workflows, jobs, arrays, and entities.