Pathnames and cache addresses
Every operation in a job has a pathname: a readable address built from the names you give things. The root workflow’s pathname is theroot you pass when launching ("root" by default).
Below it, each step, child workflow, and source array adds its local name, and a mapped or
replicated call adds its keys:
A step’s pathname is also its cache address, where its result is stored. Addresses are
shared by every job in the project’s database schema, which is what lets a later job reuse an
earlier job’s results.
Because names form addresses, choose specific local names (
comparison_judge_config rather than
config) and never put /, [, or ] in one.
When a step reuses a result
When a job reaches a step, it makes a request at the step’s address. The request’s fingerprint is the step’s registered name plus the exact input data the call receives.
A held step means the address already holds a result computed from different inputs. fxtr won’t
silently reuse a result that doesn’t match, and won’t overwrite one that another job depends on,
so it stops and tells you:
What this means in practice
Changing a step's code does not invalidate its cache
Changing a step's code does not invalidate its cache
The fingerprint includes the step’s name and inputs, not its implementation. After changing
what a step computes, launch under a new
root, or clear the step’s entries. fxtr does
record the commit each job ran at, so you can always see which code produced a result.Changed inputs are held, not rerun
Changed inputs are held, not rerun
Rerunning under the same
root with edited data reuses the unchanged slices and holds the
changed ones, along with any step downstream of them. Use a new root, or clear those
entries.Fresh samples need a new address
Fresh samples need a new address
A new job with the same
root and inputs reuses earlier model samples: the job’s ID isn’t
part of the address. For fresh samples, generate a new root in the launcher, such as
root=f"word-count-{uuid.uuid4()}". Never generate it inside a workflow.Addresses aren't seeds
Addresses aren't seeds
Two different addresses run separately even with identical inputs. That’s how replicas get
independent samples.
A job keeps the results it used
A job keeps the results it used
Clearing a cache entry affects later requests only. A finished job still has the results it
was bound to, even if a later job computes a different result at the same address.
Put everything that matters in the inputs
Everything that affects a result should reach its step as an input: data, model IDs, prompts, and sampling settings. Then changing any of them changes the step’s fingerprint, and a result computed from the old settings is never reused for the new ones. For the same reason, pass file contents or a typed dataset rather than a file name: a step given only a file name has the same fingerprint however the file’s contents change. Keep credentials in the environment, never in inputs.Clearing entries
List entries withfxtr cache list, and clear an address or everything under it with
fxtr cache clear:
uv run fxtr resume JOB. From Python,
use await client.clear_cache_entries(prefix="word-count-v1/count_words"), or
keys=[...] for exact addresses.
Cache namespaces
A job’s cache namespace, the prefix of every address in it, is itsroot. To change the
namespace of part of a job while keeping its pathnames in the viewer unchanged, pass
override_cache_key_prefix to run_step or run_child_workflow:
StepConflictError), which usually means two runs were given the
same prefix.