> ## Documentation Index
> Fetch the complete documentation index at: https://docs.transluce.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Arrays and entities

> The data model: arrays indexed by named dimensions, and stored entities they reference.

## Arrays

An **array** is a collection of values indexed by named **dimensions**. Each dimension has a name
and a key type (`int` or `str`), and each combination of keys appears at most once. Every value
in an array follows one **value schema**.

For example, scores indexed by model and task:

| model (key) | task (key) | value |
| - | - | - |
| `"baseline"` | `"math"` | `0.72` |
| `"baseline"` | `"code"` | `0.64` |
| `"variant"` | `"math"` | `0.81` |

This array's type is written `[model: str, task: str]: float`. Arrays can be sparse: the keys
don't need to form a full grid, so `variant` can lack a `code` score. A **scalar** is an array
with no dimensions and exactly one value.

Named dimensions are what make sweeps natural. Inputs line up by dimension name, so a table of
per-model settings pairs with the matching rows of a per-model, per-task table without any
bookkeeping. See [Mapping and sweeps](/fxtr/concepts/mapping).

### Building arrays

Import `Array` from `fxtr.entity_defns.array`. A type is a pair of the dimensions, written as a
string, and the value schema, written as a Python type:

```python theme={null}
from typing import TypedDict

from fxtr.entity_defns.array import Array

prompts = Array.from_items(
    [(("greeting",), "Hello"), (("question",), "Why?")],
    ("[prompt: str]", str),
)
budget = Array.scalar(200, int)


class Pair(TypedDict):
    left: int
    right: int


pairs = Array.from_records(
    [{"task": "a", "left": 1, "right": 2}],
    ("[task: str]", Pair),
)

assert pairs.field("left")[("a",)] == 1
assert prompts.sel({"prompt": "greeting"}).item() == "Hello"
```

* `Array.from_items` takes `(key, value)` pairs. A key is a tuple, or a dictionary from dimension
  name to key.
* `Array.from_records` builds an array of structs: each record holds the dimension keys and the
  struct's fields.
* `Array.scalar` builds a single-value array.
* `infer_from_items`, `infer_from_records`, and `infer_scalar` work out the type from the data.
  An empty array always needs an explicit type.

<Warning>
  Write array types as a pair, like `("[task: str]", int)`. The single-string form
  `"[task: str]: int"` is only how types are displayed, and the constructors reject it.
</Warning>

### Value schemas

A value can be:

* a primitive: `int`, `float`, `bool`, `str`, or `bytes`;
* a reference to an entity;
* a homogeneous list;
* a struct with fixed fields, written as a `TypedDict`;
* a nullable version of any of these, written `X | None`.

Dataclasses and other unions aren't supported as values; use an entity for records that need
them. Values are validated strictly and stored immutably. Schema constructors such as `id_of`,
`list_of`, `struct_of`, and `optional` are in `fxtr.entity_defns.schema` for building schemas
explicitly.

A list is a single value, not a dimension: you can't map over its elements. When you want one call
per element, make the elements rows of an array with their own dimension.

## Entities

An **entity** is a stored record with a type and an ID derived from its content: identical
content always has the same ID and is stored once. Define entity types as frozen dataclasses:

```python theme={null}
from dataclasses import dataclass

from fxtr.entity_defns.dataclass_entity import DataclassEntity


@dataclass(frozen=True)
class Document(DataclassEntity, fxtr_type="org.example.my_experiment.Document.v1"):
    title: str
    text: str
```

The `fxtr_type` string identifies the type in storage and in the viewer. Include a version, so
that you can introduce a changed type under a new name.

Storing an entity returns a `BoundID`: its ID together with the type needed to load it back.
Both step and workflow contexts, and the project client, provide `store` and `load`:

```python theme={null}
ref = await context.store(Document(title="Intro", text="..."))
document = await context.load(ref)
```

To load a bare ID, name its type: `await context.load(entity_id, as_type=Document)`.

Arrays hold **references** to entities, never the entities themselves. Store each entity first,
then build the array from the references, using `BoundID[Document]` as the value type.

### Structs or entities?

**Default to struct values, and use an entity for a record you want to refer to as a whole.** A
record is worth referring to as a whole when something points at that specific value:

* a person opens it to read it, such as a judge's reasoning or a conversation;
* a view links to it, such as a chart cell linking to the verdict behind it;
* another record cites it, such as a report naming the judge configuration it used;
* one value serves many places, such as a single judge configuration for every judge call.

So conversations, judge verdicts, model and judge configurations, and reports are usually
entities. Structs suit data read column by column: dataset rows, sweep coordinates, extracted
numbers, and summary statistics.

The practical difference:

| | Struct value | Entity |
| - | - | - |
| Identity | Lives inside its array | Has its own ID; identical content stored once |
| In workflows | Fields visible: select one with `handle.field()` | Opaque: a step that takes it depends on all of it |
| In the viewer | Shown as columns | Linked, and shown by a renderer chosen by its type |
| Types | Primitives, lists, `TypedDict`s | Any dataclass fields, including unions |
