Skip to main content
Most experiments are sweeps: run every task with every model, draw a few samples of each, and summarize. In fxtr you express a sweep by mapping a step or workflow over named dimensions.

The mapping rule

run_step and run_child_workflow share one rule. Given the dimensions listed in map_over, the runner:
  1. joins the inputs’ keys along those dimensions;
  2. calls the function once per joint key, passing each input’s slice with those dimensions removed;
  3. stacks the results into one array.
Dimensions not listed in map_over are core dimensions: they’re passed whole to each call. An input that lacks a mapped dimension is broadcast, so every call along that dimension receives the same value. The result’s dimensions are the mapped dimensions, then any replica dimensions, then the function’s own output dimensions. With map_over left empty, the default, the function is called once with the whole input arrays. Mapping is never implicit.

Start from one call

Write the function’s signature for one call, then choose which input dimensions should create multiple calls. Suppose each call of run_task takes one model, one harness, one task, and all the few-shot examples for its model: Map over the model, harness, and task dimensions:
model, harness, and task have no dimensions in common, so they cross: two models, two harnesses, and three tasks make 12 calls. examples shares the model dimension, so each call receives the slice for its own model, an [example] array, and the same slice is broadcast across harnesses and tasks. Different models can have different numbers of examples; nothing needs padding. For example, if examples holds these rows: then each call receives: If run_task returns a single value, results has dimensions [model, harness, task]. The example dimension stays inside each call and doesn’t appear in the result.

Alignment

Inputs line up by dimension name and key:
  • Inputs on different dimensions cross: [model] and [task] produce every model × task.
  • Inputs sharing a dimension align by key; they don’t cross. If two independent axes happen to share a name, rename one with handle.rename(...).
  • A single value broadcasts to every call.
  • join="exact", the default, fails if the join would drop any input’s key, so nothing is silently skipped. Sparse arrays are fine; they don’t need to be full grids.
  • join="inner" allows dropped keys and runs only the keys every input has.

Restrict which combinations run

To run only some combinations, say, each model with the other provider’s harness, put the allowed pairs in one array with joint keys, and select its fields:
Both fields keep the joint (model, harness) keys, so the map runs only the two allowed pairs, crossed with the tasks: six calls instead of 12. join="inner" lets examples contain models that aren’t selected. If the allowed pairs depend on earlier results, compute them in a step that takes those results, rather than observing them and adding them back as a source array.

Reductions

To summarize over some dimensions, leave them inside the reduction step’s signature and map over the rest. For scores shaped [model, task, sample]: float, a step that takes all of one model’s scores:
Each call receives one model’s task × sample scores, and per_model has dimensions [model]: float. With no map_over, a step that accepts the full three-dimensional array can return one overall number.

Replicas

replicas repeats each call with the same inputs, to draw independent samples:
samples has dimensions [model, prompt, sample]. Each replica has its own cache address, which is what makes the samples independent. The step body doesn’t receive the replica number, and nothing seeds the model. If the body needs a sample number or a seed, pass it as an ordinary input array and map over its dimension instead.

Putting it together

A typical experiment chains these operations, each consuming the previous handle so the whole dependency chain shows up in the viewer:
  1. Map a sampling step over the model and prompt inputs, with replicas for independent draws.
  2. Map a scoring step over the resulting cells.
  3. Reduce with a step that keeps the prompt and sample dimensions in its signature, mapped over model.
For a worked example with model calls and an LLM judge, see Calling language models.