> ## Documentation Index
> Fetch the complete documentation index at: https://docs.transluce.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Your first experiment

> Write a small experiment with a mapped step and a reduction, launch it, and see how caching works.

This tutorial builds a small experiment in the project from the [Quickstart](/fxtr/quickstart).
It measures the length of each prompt in a dataset, then averages the lengths. There are no model
calls, so it runs instantly, but it has the same shape as most real experiments: a step mapped
over a dataset, followed by a step that reduces the results.

## Write the experiment

Create `src/my_project/word_count.py`:

```python src/my_project/word_count.py theme={null}
from typing import Annotated

from fxtr.entity_defns.array import Array
from fxtr.experiment.handles import ArrayHandle
from fxtr.experiment.steps import StepContext, step
from fxtr.experiment.workflows import WorkflowContext, workflow


@step(name="word_count.count")
async def count_words(context: StepContext, prompt: str) -> int:
    """Count the whitespace-separated words in a prompt."""
    return len(prompt.split())


@step(name="word_count.mean")
async def mean_length(
    context: StepContext,
    counts: Annotated[Array[int], "[prompt: str]"],
) -> float:
    """Average the word counts across prompts."""
    values = list(counts.values())
    return sum(values) / len(values)


@workflow(name="word_count.experiment")
async def experiment(
    context: WorkflowContext,
    prompts: Annotated[ArrayHandle[str], "[prompt: str]"],
) -> Annotated[ArrayHandle[float], "[]"]:
    """Measure the length of each prompt, then average the lengths."""
    counts = context.run_step("count_words", count_words, {"prompt": prompts}, map_over=["prompt"])
    return context.run_step("mean_length", mean_length, {"counts": counts})
```

A few things to notice:

* **Steps do the work.** `count_words` takes one prompt and returns one number. `mean_length`
  takes a whole array of counts, indexed by the `prompt` dimension, and returns one number.
* **The workflow connects them.** `run_step` schedules a step and immediately returns a
  *handle* to its future result, without waiting for it. Passing `counts` to the second step is
  what makes it depend on the first.
* **`map_over=["prompt"]` makes a sweep.** The runner calls `count_words` once per prompt and
  stacks the results into an array with the same `prompt` dimension. The second `run_step` has
  no `map_over`, so `mean_length` runs once and receives the whole array.
* **Annotations declare array shapes.** A plain `str` or `int` means a single value.
  `Annotated[..., "[prompt: str]"]` names the dimensions, which a Python type alone can't say.
* **Docstrings show up in the viewer**, on each step's and workflow's card.

Then register the module in `pyproject.toml`, so that fxtr can find the experiment by name when
it resumes a job or shows it in the viewer:

```toml pyproject.toml theme={null}
[tool.fxtr]
modules = ["my_project.experiment", "my_project.word_count"]
```

## Launch it

Create a launcher at the project root. It builds the input dataset and runs the workflow as a
job:

```python launch_word_count.py theme={null}
import anyio

from fxtr.entity_defns.array import Array
from fxtr.project.running import open_local_client, run_job
from my_project.word_count import experiment


async def main() -> None:
    prompts = Array.from_items(
        [
            (("greeting",), "Hello there"),
            (("question",), "Why is the sky blue?"),
            (("request",), "Please summarize this article in three sentences."),
        ],
        ("[prompt: str]", str),
    )
    async with open_local_client(__file__) as client:
        result = await run_job(client, experiment, {"prompts": prompts}, root="word-count-v1")
        print(f"mean length: {result.item():.2f} words")


if __name__ == "__main__":
    anyio.run(main)
```

`Array.from_items` takes `(key, value)` pairs and a type: here, one `prompt` dimension with string
keys, holding string values. `root` names the job's root workflow, and, as you'll see below,
decides which earlier results the job can reuse.

Commit, then launch. A launch records the commit it runs, so it refuses uncommitted changes:

```bash theme={null}
git add . && git commit -m "Add the word count experiment"
uv run python launch_word_count.py
```

```text theme={null}
job 3eb75a59-911e-44af-a148-411ac5dcd8d1: running
watch it at http://127.0.0.1:8000/#job=3eb75a59-911e-44af-a148-411ac5dcd8d1
job 3eb75a59-911e-44af-a148-411ac5dcd8d1: succeeded
...
mean length: 4.67 words
```

With `uv run fxtr view` running, open the printed link to see the job: the `prompts` input feeding
three `count_words` calls, one per prompt, feeding `mean_length`.

## Rerun it

Launch it again, unchanged. The job succeeds straight away, because every step finds its result
in the cache. You can see the entries with `fxtr cache list`:

```text theme={null}
completed  word-count-v1/count_words[prompt:"greeting"]  (word_count.count)
completed  word-count-v1/count_words[prompt:"question"]  (word_count.count)
completed  word-count-v1/count_words[prompt:"request"]  (word_count.count)
completed  word-count-v1/mean_length  (word_count.mean)
```

Each entry is at a **cache address** built from the job's `root`, the step's name in the
workflow, and, for a mapped step, the key it was called with. A step reuses a cached result when
its address holds a result computed from the same step and the same inputs.

## Change an input

Now edit the `question` prompt, commit, and launch again under the same `root`. The job stops:

```text theme={null}
job 3fa3df92-2edb-4088-9e1f-a99d28b7ec77: stopped
  reason: cache key 'word-count-v1/count_words[prompt:"question"]' is held by a different request; clear it to run this request there
Traceback (most recent call last):
  ...
fxtr.project.client.JobNotFinishedError: Job 3fa3df92-2edb-4088-9e1f-a99d28b7ec77 is stopped
```

The address `word-count-v1/count_words[prompt:"question"]` already holds a result computed from
the old prompt. Rather than silently reuse a result that no longer matches, or overwrite a result
an earlier job depends on, fxtr **holds** the step and stops the job once nothing else can run.
The unchanged prompts still reuse their results. `run_job` raises `JobNotFinishedError` for any
job that doesn't succeed, so your launcher can tell.

You have two ways forward:

<Tabs>
  <Tab title="Use a new root">
    Change `root="word-count-v1"` to `root="word-count-v2"` and launch again. Every address in
    the new job is fresh, so every step runs. This is the usual choice when you change the
    experiment's data or settings, and it keeps the old results intact.
  </Tab>

  <Tab title="Clear and resume">
    Clear the held addresses and resume the stopped job. Steps downstream of a changed input are
    held too (here, `mean_length`), so clear everything under the root:

    ```bash theme={null}
    uv run fxtr cache clear word-count-v1
    uv run fxtr resume 3fa3df92-2edb-4088-9e1f-a99d28b7ec77
    ```

    Clearing affects later requests only. Jobs that already finished keep the results they used.
  </Tab>
</Tabs>

<Warning>
  Changing a step's **code** does not invalidate its cached results: the cache matches a step by
  its name and inputs, not its implementation. After changing what a step computes, launch under
  a new `root` or clear its entries. See [Caching and reruns](/fxtr/concepts/caching).
</Warning>

## Working with a coding agent

A new project includes fxtr's guides for coding agents in `.agents/skills/`, where Codex finds
them, linked from `.claude/skills/`, where Claude Code finds them. They cover writing steps and
workflows, launching jobs, caching, calling models with behaviors, and building viewer renderers,
so you can ask an agent for an experiment in plain language and review the result in the viewer.
Commit them with the project.

After updating your fxtr checkout, refresh them from the project:

```bash theme={null}
uv run fxtr skills install
```

This replaces fxtr's own guides and leaves any other skills in the project alone.

## Next steps

<CardGroup cols={2}>
  <Card title="Mapping and sweeps" icon="table-cells" href="/fxtr/concepts/mapping">
    Sweep over models, prompts, and samples, and reduce over any dimension.
  </Card>

  <Card title="Calling language models" icon="comments" href="/fxtr/guides/language-models">
    Use behaviors to hold conversations and judge outputs in your steps.
  </Card>
</CardGroup>
