> ## Documentation Index
> Fetch the complete documentation index at: https://docs.transluce.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Mapping and sweeps

> Run a step once per model, prompt, or sample, and reduce over any dimension.

Most experiments are sweeps: run every task with every model, draw a few samples of each, and
summarize. In fxtr you express a sweep by **mapping** a step or workflow over named dimensions.

## The mapping rule

`run_step` and `run_child_workflow` share one rule. Given the dimensions listed in `map_over`,
the runner:

1. joins the inputs' keys along those dimensions;
2. calls the function once per joint key, passing each input's slice with those dimensions
   removed;
3. stacks the results into one array.

Dimensions not listed in `map_over` are **core dimensions**: they're passed whole to each call.
An input that lacks a mapped dimension is **broadcast**, so every call along that dimension
receives the same value. The result's dimensions are the mapped dimensions, then any replica
dimensions, then the function's own output dimensions.

With `map_over` left empty, the default, the function is called once with the whole input arrays.
Mapping is never implicit.

## Start from one call

Write the function's signature for **one call**, then choose which input dimensions should create
multiple calls. Suppose each call of `run_task` takes one model, one harness, one task, and all the
few-shot examples for its model:

| Parameter | One call expects | Supplied array |
| - | - | - |
| `model` | a single `str` | `[model]: str` |
| `harness` | a single `str` | `[harness]: str` |
| `task` | a single `str` | `[task]: str` |
| `examples` | `[example]: str` | `[model, example]: str` |

Map over the model, harness, and task dimensions:

```python theme={null}
results = context.run_step(
    "run_tasks",
    run_task,
    {"model": models, "harness": harnesses, "task": tasks, "examples": examples},
    map_over=["model", "harness", "task"],
)
```

`model`, `harness`, and `task` have no dimensions in common, so they cross: two models, two
harnesses, and three tasks make 12 calls. `examples` shares the `model` dimension, so each call
receives the slice for its own model, an `[example]` array, and the same slice is broadcast across
harnesses and tasks. Different models can have different numbers of examples; nothing needs
padding.

For example, if `examples` holds these rows:

| model | example | value |
| - | - | - |
| `model-a` | `0` | `A_tools` |
| `model-b` | `1` | `B_tools` |
| `model-a` | `2` | `A_examples` |

then each call receives:

| model | harness | task | `examples` passed to the call |
| - | - | - | - |
| `model-a` | any | any | `{0: A_tools, 2: A_examples}` |
| `model-b` | any | any | `{1: B_tools}` |

If `run_task` returns a single value, `results` has dimensions `[model, harness, task]`. The
`example` dimension stays inside each call and doesn't appear in the result.

## Alignment

Inputs line up by dimension **name** and key:

* Inputs on different dimensions cross: `[model]` and `[task]` produce every model × task.
* Inputs sharing a dimension align by key; they don't cross. If two independent axes happen to
  share a name, rename one with `handle.rename(...)`.
* A single value broadcasts to every call.
* `join="exact"`, the default, fails if the join would drop any input's key, so nothing is
  silently skipped. Sparse arrays are fine; they don't need to be full grids.
* `join="inner"` allows dropped keys and runs only the keys every input has.

## Restrict which combinations run

To run only some combinations, say, each model with the other provider's harness, put the allowed
pairs in one array with **joint** keys, and select its fields:

```python theme={null}
from typing import TypedDict

from fxtr.entity_defns.array import Array


class ModelHarness(TypedDict):
    model_name: str
    harness_name: str


allowed = context.add_source_array(
    "allowed_model_harness_pairs",
    Array.from_records(
        [
            {"model": "m1", "harness": "h2", "model_name": "model-1", "harness_name": "harness-2"},
            {"model": "m2", "harness": "h1", "model_name": "model-2", "harness_name": "harness-1"},
        ],
        ("[model: str, harness: str]", ModelHarness),
    ),
)
results = context.run_step(
    "run_allowed_tasks",
    run_task,
    {
        "model": allowed.field("model_name"),
        "harness": allowed.field("harness_name"),
        "task": tasks,
        "examples": examples,
    },
    map_over=["model", "harness", "task"],
    join="inner",
)
```

Both fields keep the joint `(model, harness)` keys, so the map runs only the two allowed pairs,
crossed with the tasks: six calls instead of 12. `join="inner"` lets `examples` contain models
that aren't selected. If the allowed pairs depend on earlier results, compute them in a step that
takes those results, rather than observing them and adding them back as a source array.

## Reductions

To summarize over some dimensions, leave them **inside** the reduction step's signature and map
over the rest. For scores shaped `[model, task, sample]: float`, a step that takes all of one
model's scores:

```python theme={null}
@step(name="my_experiment.mean_score")
async def mean_score(
    context: StepContext,
    scores: Annotated[Array[float], "[task: str, sample: int]"],
) -> float:
    """Average a model's scores across tasks and samples."""
    values = list(scores.values())
    return sum(values) / len(values)


# In the workflow:
per_model = context.run_step("mean_scores", mean_score, {"scores": scores}, map_over=["model"])
```

Each call receives one model's task × sample scores, and `per_model` has dimensions
`[model]: float`. With no `map_over`, a step that accepts the full three-dimensional array can
return one overall number.

## Replicas

`replicas` repeats each call with the same inputs, to draw independent samples:

```python theme={null}
samples = context.run_step(
    "sample_answers",
    sample_answer,
    {"model": models, "prompt": prompts},
    map_over=["model", "prompt"],
    replicas={"sample": range(3)},
)
```

`samples` has dimensions `[model, prompt, sample]`. Each replica has its own cache address, which
is what makes the samples independent. The step body doesn't receive the replica number, and
nothing seeds the model. If the body needs a sample number or a seed, pass it as an ordinary
input array and map over its dimension instead.

## Putting it together

A typical experiment chains these operations, each consuming the previous handle so the whole
dependency chain shows up in the viewer:

1. Map a sampling step over the model and prompt inputs, with replicas for independent draws.
2. Map a scoring step over the resulting cells.
3. Reduce with a step that keeps the prompt and sample dimensions in its signature, mapped over
   `model`.

For a worked example with model calls and an LLM judge, see
[Calling language models](/fxtr/guides/language-models).
