Why fxtr
Experiments on AI systems tend to sprawl. A few scripts become dozens, JSON files pile up on disk, and it gets hard to say which version of the code produced which results, or to check that a judge did something sensible on the fiftieth sample. Coding agents make this worse: they can write an experiment in minutes, but reviewing what they wrote, and what it produced, takes much longer. fxtr keeps an experiment organized as it grows:- Every result is stored in a database with the step that computed it, the inputs it received, and the commit the job ran at. In the viewer you can go from a chart to the records behind it, and from a record to the code that produced it.
- fxtr runs the experiment, stores its data, and displays it, so the code you review is the experiment itself: what each step computes and how the steps connect.
- A step reuses a cached result only when the step and its inputs match. When they don’t, the job stops instead of using a result computed from different inputs.
- A job records its progress as it runs. If it stops, from a crash, a failed API call, or a cancellation, resuming it continues from where it stopped.
How it fits together
An fxtr project is a Python package of steps and workflows:- A step does one piece of computation, such as a model call, a judgment, or a reduction. Its result is cached.
- A workflow schedules steps and child workflows over arrays: collections of values
indexed by named dimensions such as
model,prompt, orsample. Mapping a step over those dimensions is how you express a sweep.
Guides for coding agents
fxtr ships guides for coding agents alongside the library. A new project includes them, so Claude Code and Codex know how to write steps and workflows, launch jobs, and build viewer renderers in your project. See Your first experiment.Quickstart
Create a project and launch its first job.
Core concepts
Projects, steps, workflows, jobs, arrays, and entities.