Skip to main content

Arrays

An array is a collection of values indexed by named dimensions. Each dimension has a name and a key type (int or str), and each combination of keys appears at most once. Every value in an array follows one value schema. For example, scores indexed by model and task: This array’s type is written [model: str, task: str]: float. Arrays can be sparse: the keys don’t need to form a full grid, so variant can lack a code score. A scalar is an array with no dimensions and exactly one value. Named dimensions are what make sweeps natural. Inputs line up by dimension name, so a table of per-model settings pairs with the matching rows of a per-model, per-task table without any bookkeeping. See Mapping and sweeps.

Building arrays

Import Array from fxtr.entity_defns.array. A type is a pair of the dimensions, written as a string, and the value schema, written as a Python type:
  • Array.from_items takes (key, value) pairs. A key is a tuple, or a dictionary from dimension name to key.
  • Array.from_records builds an array of structs: each record holds the dimension keys and the struct’s fields.
  • Array.scalar builds a single-value array.
  • infer_from_items, infer_from_records, and infer_scalar work out the type from the data. An empty array always needs an explicit type.
Write array types as a pair, like ("[task: str]", int). The single-string form "[task: str]: int" is only how types are displayed, and the constructors reject it.

Value schemas

A value can be:
  • a primitive: int, float, bool, str, or bytes;
  • a reference to an entity;
  • a homogeneous list;
  • a struct with fixed fields, written as a TypedDict;
  • a nullable version of any of these, written X | None.
Dataclasses and other unions aren’t supported as values; use an entity for records that need them. Values are validated strictly and stored immutably. Schema constructors such as id_of, list_of, struct_of, and optional are in fxtr.entity_defns.schema for building schemas explicitly. A list is a single value, not a dimension: you can’t map over its elements. When you want one call per element, make the elements rows of an array with their own dimension.

Entities

An entity is a stored record with a type and an ID derived from its content: identical content always has the same ID and is stored once. Define entity types as frozen dataclasses:
The fxtr_type string identifies the type in storage and in the viewer. Include a version, so that you can introduce a changed type under a new name. Storing an entity returns a BoundID: its ID together with the type needed to load it back. Both step and workflow contexts, and the project client, provide store and load:
To load a bare ID, name its type: await context.load(entity_id, as_type=Document). Arrays hold references to entities, never the entities themselves. Store each entity first, then build the array from the references, using BoundID[Document] as the value type.

Structs or entities?

Default to struct values, and use an entity for a record you want to refer to as a whole. A record is worth referring to as a whole when something points at that specific value:
  • a person opens it to read it, such as a judge’s reasoning or a conversation;
  • a view links to it, such as a chart cell linking to the verdict behind it;
  • another record cites it, such as a report naming the judge configuration it used;
  • one value serves many places, such as a single judge configuration for every judge call.
So conversations, judge verdicts, model and judge configurations, and reports are usually entities. Structs suit data read column by column: dataset rows, sweep coordinates, extracted numbers, and summary statistics. The practical difference: