Skip to main content
fxtr is a framework for running experiments on AI systems: sampling models, holding conversations, judging outputs, and aggregating the results. You write the experiment’s logic in Python. fxtr takes care of running it at scale, caching expensive results, recovering from interruptions, and showing you exactly how every result was produced.

Why fxtr

Experiments on AI systems tend to sprawl. A few scripts become dozens, JSON files pile up on disk, and it gets hard to say which version of the code produced which results, or to check that a judge did something sensible on the fiftieth sample. Coding agents make this worse: they can write an experiment in minutes, but reviewing what they wrote, and what it produced, takes much longer. fxtr keeps an experiment organized by default:

Results you can trace

Every result is stored in a database alongside the step that computed it, the inputs it received, and the code commit the job ran at. The viewer lets you move from a chart to the records behind it to the code that produced them.

Less code to review

Execution, storage, and visualization are handled for you, so the code left to read is the experiment itself: what each step computes and how the steps connect.

Caching you can trust

Model calls are expensive. fxtr reuses a cached result only when the step and its inputs match, and stops rather than silently reusing a result computed from different inputs.

Durable jobs

A job records its progress as it runs. If it stops, whether from a crash, a failed API call, or a cancellation, resuming it picks up where it left off.
None of this limits what an experiment can express. Steps are ordinary async Python functions, and workflows can make data-dependent decisions, such as running another refinement round only when a score is too low.

How it fits together

An fxtr project is a Python package of steps and workflows:
  • A step does one piece of computation, such as a model call, a judgment, or a reduction. Its result is cached.
  • A workflow schedules steps and child workflows over arrays: collections of values indexed by named dimensions such as model, prompt, or sample. Mapping a step over those dimensions is how you express a sweep.
Launching a workflow starts a job on the project’s Postgres database. The job builds a graph of everything the workflow scheduled, runs the steps, and stores their results. The experiment viewer shows the graph, the arrays flowing through it, and the records they reference, through renderers you can customize. fxtr pairs with behaviors, a library for calling language models from fxtr steps. It provides model adapters for OpenAI and Anthropic, multi-turn conversations, retries, and storable conversation records.

Built for working with coding agents

fxtr ships guides for coding agents alongside the library. A new project includes them, so Claude Code and Codex know how to write steps and workflows, launch jobs, and build viewer renderers in your project. See Your first experiment.

Quickstart

Create a project and launch its first job.

Core concepts

Projects, steps, workflows, jobs, arrays, and entities.