- Workflow inputs, passed by the launcher: the usual choice for datasets.
- Source arrays, added inside a workflow with
context.add_source_array(...): for configuration fixed in the experiment’s code. - Project arrays: named datasets kept in the database, which jobs run against as snapshots.
External datasets
Read and validate external data in the launcher, convert it to a typed array, and pass it to the workflow. For a CSV with columnscase_id,prompt,expected:
[case_id: str] handle to map steps over, and can
pass individual fields to steps with cases.field("prompt") and cases.field("expected").
A few tips:
- Key rows by stable IDs from the source, not by row number, so that reordering or adding rows doesn’t change which cached result belongs to which row.
- Convert types yourself.
Array.from_recordsvalidates the schema and rejects duplicate keys, but it doesn’t convert values: turn numbers, booleans, and missing CSV values into the declared types before building the array. - Validate early. Check dataset-specific requirements, such as nonempty prompts, in the loader.
- Don’t read files in workflows. Workflows must replay deterministically, so external reads belong in the launcher or in a step.
Configuration
For settings fixed in the experiment’s code, add a source array inside the workflow:Entity-valued inputs
To pass entities, store each one with the client first, then build an array from the returned references:Project arrays
A project array is a named, editable, versioned array stored in the project’s database. It suits a dataset you curate over time and launch many jobs against. A job always runs against an immutable snapshot, so editing the array later doesn’t change what earlier jobs used:replace_array sets the complete contents; upsert_rows adds or updates the given rows and keeps
the rest. From the command line, pass a snapshot as an input with @:
Array to run_job directly.