> ## Documentation Index
> Fetch the complete documentation index at: https://docs.transluce.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Input data

> Get datasets, configuration, and stored entities into an experiment.

Everything that affects a result should reach its step as an input: data, model IDs, prompts, and
sampling settings. Then changing any of them changes what the step receives, so fxtr never reuses
a result computed from old settings (see [Caching and reruns](/fxtr/concepts/caching)). There are
three ways in:

* **Workflow inputs**, passed by the launcher: the usual choice for datasets.
* **Source arrays**, added inside a workflow with `context.add_source_array(...)`: for
  configuration fixed in the experiment's code.
* **Project arrays**: named datasets kept in the database, which jobs run against as snapshots.

Credentials belong in the environment, never in inputs.

## External datasets

Read and validate external data in the launcher, convert it to a typed array, and pass it to the
workflow. For a CSV with columns `case_id,prompt,expected`:

```python theme={null}
import csv
from pathlib import Path
from typing import TypedDict

from fxtr.entity_defns.array import Array


class Case(TypedDict):
    prompt: str
    expected: str


def load_cases(path: Path):
    with path.open(newline="", encoding="utf-8") as file:
        return Array.from_records(csv.DictReader(file), type=("[case_id: str]", Case))
```

Then, inside an [async launcher](/fxtr/guides/running-jobs#launch-from-python):

```python theme={null}
cases = load_cases(Path("cases.csv"))
result = await run_job(client, experiment, {"cases": cases}, root="cases-v1")
```

Submitting stores the input array with the job, so the job keeps the exact data it ran on even if
the CSV changes later. The workflow receives a `[case_id: str]` handle to map steps over, and can
pass individual fields to steps with `cases.field("prompt")` and `cases.field("expected")`.

A few tips:

* **Key rows by stable IDs** from the source, not by row number, so that reordering or adding
  rows doesn't change which cached result belongs to which row.
* **Convert types yourself.** `Array.from_records` validates the schema and rejects duplicate
  keys, but it doesn't convert values: turn numbers, booleans, and missing CSV values into the
  declared types before building the array.
* **Validate early.** Check dataset-specific requirements, such as nonempty prompts, in the
  loader.
* **Don't read files in workflows.** Workflows must replay deterministically, so external reads
  belong in the launcher or in a step.

## Configuration

For settings fixed in the experiment's code, add a source array inside the workflow:

```python theme={null}
judge_rubric = context.add_source_array("judge_rubric", Array.scalar(RUBRIC, str))
```

Settings used together often belong in one entity, such as a judge configuration holding the
model ID, rubric, and sampling parameters. Pass a reference to it as an input, so every result
points at the exact settings that produced it. See
[Calling language models](/fxtr/guides/language-models).

## Entity-valued inputs

To pass entities, store each one with the client first, then build an array from the returned
references:

```python theme={null}
async with open_local_client(__file__) as client:
    intro = await client.store(Document(title="Intro", text="..."))
    documents = Array.from_items([(("intro",), intro)], ("[doc: str]", BoundID[Document]))
    await run_job(client, experiment, {"documents": documents}, root="docs-v1")
```

Submitting the job, or writing the array into a project array, saves the entities it references.

## Project arrays

A **project array** is a named, editable, versioned array stored in the project's database. It
suits a dataset you curate over time and launch many jobs against. A job always runs against an
immutable **snapshot**, so editing the array later doesn't change what earlier jobs used:

```python theme={null}
await client.create_array("word_count_prompts", ("[prompt: str]", str))
await client.replace_array("word_count_prompts", prompts)
snapshot = await client.snapshot_array("word_count_prompts")
result = await run_job(client, experiment, {"prompts": snapshot.array}, root="word-count-v1")
```

`replace_array` sets the complete contents; `upsert_rows` adds or updates the given rows and keeps
the rest. From the command line, pass a snapshot as an input with `@`:

```bash theme={null}
uv run fxtr run word_count.experiment --input prompts=@word_count_prompts
```

Project arrays are optional. When you don't need a named, editable dataset, pass an ordinary
`Array` to `run_job` directly.
