Skip to main content
fxtr is a framework for running experiments on AI systems: sampling models, holding conversations, judging outputs, and aggregating the results. You write the experiment’s logic in Python. fxtr takes care of running it at scale, caching expensive results, recovering from interruptions, and recording how every result was produced.

Why fxtr

Experiments on AI systems tend to sprawl. A few scripts become dozens, JSON files pile up on disk, and it gets hard to say which version of the code produced which results, or to check that a judge did something sensible on the fiftieth sample. Coding agents make this worse: they can write an experiment in minutes, but reviewing what they wrote, and what it produced, takes much longer. fxtr keeps an experiment organized as it grows:
  • Every result is stored in a database with the step that computed it, the inputs it received, and the commit the job ran at. In the viewer you can go from a chart to the records behind it, and from a record to the code that produced it.
  • fxtr runs the experiment, stores its data, and displays it, so the code you review is the experiment itself: what each step computes and how the steps connect.
  • A step reuses a cached result only when the step and its inputs match. When they don’t, the job stops instead of using a result computed from different inputs.
  • A job records its progress as it runs. If it stops, from a crash, a failed API call, or a cancellation, resuming it continues from where it stopped.
Steps are ordinary async Python functions, and workflows can make data-dependent decisions, such as running another refinement round only when a score is too low.

How it fits together

An fxtr project is a Python package of steps and workflows:
  • A step does one piece of computation, such as a model call, a judgment, or a reduction. Its result is cached.
  • A workflow schedules steps and child workflows over arrays: collections of values indexed by named dimensions such as model, prompt, or sample. Mapping a step over those dimensions is how you express a sweep.
Launching a workflow starts a job on the project’s Postgres database. The job builds a graph of everything the workflow scheduled, runs the steps, and stores their results. The experiment viewer shows the graph, the arrays flowing through it, and the records they reference, through renderers you can customize. fxtr pairs with behaviors, a library for calling language models from fxtr steps. It provides model adapters for OpenAI and Anthropic, multi-turn conversations, retries, and storable conversation records.

Guides for coding agents

fxtr ships guides for coding agents alongside the library. A new project includes them, so Claude Code and Codex know how to write steps and workflows, launch jobs, and build viewer renderers in your project. See Your first experiment.

Quickstart

Create a project and launch its first job.

Core concepts

Projects, steps, workflows, jobs, arrays, and entities.