Skip to main content
An entity is a stored record with a type and an ID derived from its content. Where an array holds many regular values addressed by key, an entity is a single thing you can point at. Arrays hold references to entities, entities can hold references to other entities, and the viewer knows how to follow every reference and show what it finds. Entities are for three kinds of data:
  • Large objects you want to refer to later as a whole. A conversation transcript, a document, a generated report. You pass references around and load the content only where it’s needed.
  • Objects with irregular structure. Array values must follow one fixed schema, but an entity can hold a tagged union, nested optional parts, or a blob of JSON, because it is serialized as a unit rather than laid out in columns.
  • Things other objects should point to. A judge’s verdict about a transcript naturally references the transcript. A result should reference the configuration that produced it. Entity references make these links explicit, and the viewer renders them as links you can click.
Because an entity’s ID is a hash of its content, identical content always has the same ID and is stored once. Storing the same model configuration from a hundred steps costs one record, and two jobs that reference it reference the very same thing.

Defining entities

Most entity types are frozen dataclasses that inherit from DataclassEntity:
The fxtr_type keyword is the entity’s type tag. Every entity type must have one, it must be unique, and it is how the system refers to the type everywhere: in storage, in array schemas, and in the viewer, where renderers are keyed by it. Use a reverse-domain name, and include a version so that a changed type can be introduced under a new name while old records keep theirs. Fields may be primitives, tuples and frozen sets, nested dataclasses, BoundID[T] references to other entities, nullable and union types, and recursive JSON values. The Answer above points at its ModelConfig by reference, so every answer records exactly which configuration produced it without copying the configuration into every record.

Inline tagged unions

A field whose value can be one of several shapes, say a step outcome that is either a score or a refusal, is written as a union of tagged dataclasses. A TaggedDataclass carries a type tag like an entity does, but it is not stored on its own: it is serialized inline, inside whatever entity contains it, and the tag tells the loader which variant it is reading.
This is the kind of structure an array value cannot hold, and one of the main reasons to reach for an entity.
fxtr currently converts dataclasses to and from their stored form with the cattrs library, with strict settings: unknown fields, missing fields, and values of the wrong type are rejected rather than coerced. This implementation detail may change in a future version; the dataclass surface described here is what to rely on.

Advanced use: low-level entity classes

Underneath the dataclass layer, an entity is defined by one ability: it can serialize itself to a canonical form, and something can deserialize that form back into an object. The canonical form is a tree of plain values (None, numbers, booleans, strings, bytes, lists, string-keyed mappings, and entity IDs) with a $type key holding the type tag, encoded as DAG-CBOR and hashed to produce the ID. Any class can be an entity by subclassing Entity and implementing its abstract methods:
_fxtr_serialize produces the canonical form and _fxtr_loader names the loader that turns it back into an object. A loader’s convert receives the stored data and returns the value; StaticTypeLoader is the base for a loader that handles exactly one type tag. Note that the low-level protocol is usually unnecessary; prefer to use DataclassEntity if possible.

Storing and loading entities

Storing an entity returns a BoundID: the entity’s ID together with the loader that knows how to read it back. You can store from a step context, a workflow context, or the project client in a launcher, and all three expose the same two methods:
A BoundID is cheap to pass around and to put in other entities and arrays. A bare EntityID, such as one read out of an array of untyped IDs, needs its type to be loaded: await context.load(entity_id, as_type=ModelConfig). The only way to get an entity back is to have its ID, and storing an entity gives it no name and puts it in no list you can browse. So for an entity to be retrievable later, something that is itself persisted in the database must reference it: a job’s inputs, which the job records; a job’s results, which hold the references of the entities steps return; or a project array. An entity that nothing references can’t be found again, so when you store entities in a launcher, pass the references on as a job input or write them into a project array.

Entities in arrays

Arrays hold references, never entities themselves. The common pattern is to store each entity and then build the array from the references, with BoundID[T] as the value type:
Two shorthands cover the step side. A step annotated to receive an entity, as in config: ModelConfig, gets the loaded object rather than the reference. A step annotated to return an entity type, as in -> Answer, may return the entity itself: fxtr stores it and records the reference as the step’s result, so the mapped results of such a step form an array of BoundID[Answer] without any explicit store calls.

Entities versus array structs

There are two ways to put structured data in an array: a struct value, written as a TypedDict, whose fields live inside the array as columns, or an entity, stored separately and referenced by ID. The rule of thumb:
  • Use a struct when the data is regular and exists to organize values for this step or workflow, without needing an identity of its own. “A collection of an X and a Y”, dataset rows, sweep coordinates, numbers extracted for aggregation, summary statistics.
  • Use an entity when you want to refer to the thing by ID and point at it from elsewhere, when it is naturally a single thing rather than a collection of values (a transcript, a verdict, a configuration, a report), or when its shape is irregular.
The practical differences follow from that: A step can combine the two: store an entity for the record that deserves one, and return a struct row holding the fields you’ll aggregate next to a BoundID pointing at the entity. The aggregation then reads columns, and anyone who wants the detail follows the reference.
Coming soon: We plan to make it possible to efficiently extract parts of entities into an array struct, to make it possible to efficiently process and summarize individual fields of entity data.