The mapping rule
run_step and run_child_workflow share one rule. Given the dimensions listed in map_over,
the runner:
- joins the inputs’ keys along those dimensions;
- calls the function once per joint key, passing each input’s slice with those dimensions removed;
- stacks the results into one array.
map_over are core dimensions: they’re passed whole to each call.
An input that lacks a mapped dimension is broadcast, so every call along that dimension
receives the same value. The result’s dimensions are the mapped dimensions, then any replica
dimensions, then the function’s own output dimensions.
With map_over left empty, the default, the function is called once with the whole input arrays.
Mapping is never implicit.
Start from one call
Write the function’s signature for one call, then choose which input dimensions should create multiple calls. Suppose each call ofrun_task takes one model, one harness, one task, and all the
few-shot examples for its model:
Map over the model, harness, and task dimensions:
model, harness, and task have no dimensions in common, so they cross: two models, two
harnesses, and three tasks make 12 calls. examples shares the model dimension, so each call
receives the slice for its own model, an [example] array, and the same slice is broadcast across
harnesses and tasks. Different models can have different numbers of examples; nothing needs
padding.
For example, if examples holds these rows:
then each call receives:
If
run_task returns a single value, results has dimensions [model, harness, task]. The
example dimension stays inside each call and doesn’t appear in the result.
Alignment
Inputs line up by dimension name and key:- Inputs on different dimensions cross:
[model]and[task]produce every model × task. - Inputs sharing a dimension align by key; they don’t cross. If two independent axes happen to
share a name, rename one with
handle.rename(...). - A single value broadcasts to every call.
join="exact", the default, fails if the join would drop any input’s key, so nothing is silently skipped. Sparse arrays are fine; they don’t need to be full grids.join="inner"allows dropped keys and runs only the keys every input has.
Restrict which combinations run
To run only some combinations, say, each model with the other provider’s harness, put the allowed pairs in one array with joint keys, and select its fields:(model, harness) keys, so the map runs only the two allowed pairs,
crossed with the tasks: six calls instead of 12. join="inner" lets examples contain models
that aren’t selected. If the allowed pairs depend on earlier results, compute them in a step that
takes those results, rather than observing them and adding them back as a source array.
Reductions
To summarize over some dimensions, leave them inside the reduction step’s signature and map over the rest. For scores shaped[model, task, sample]: float, a step that takes all of one
model’s scores:
per_model has dimensions
[model]: float. With no map_over, a step that accepts the full three-dimensional array can
return one overall number.
Replicas
replicas repeats each call with the same inputs, to draw independent samples:
samples has dimensions [model, prompt, sample]. Each replica has its own cache address, which
is what makes the samples independent. The step body doesn’t receive the replica number, and
nothing seeds the model. If the body needs a sample number or a seed, pass it as an ordinary
input array and map over its dimension instead.
Putting it together
A typical experiment chains these operations, each consuming the previous handle so the whole dependency chain shows up in the viewer:- Map a sampling step over the model and prompt inputs, with replicas for independent draws.
- Map a scoring step over the resulting cells.
- Reduce with a step that keeps the prompt and sample dimensions in its signature, mapped over
model.