Mental model
Four facts describe the cache:- Every step invocation has a cache address. Within one job, a cache address always resolves to exactly one value. By default each invocation a workflow schedules gets a unique address built from its position in the job, so you rarely need to think about addresses unless you want more control over how results are shared between jobs.
- Every invocation also has a fingerprint: the step’s registered name together with the exact input data it received. Two invocations with the same fingerprint would compute the same thing (or, for a nondeterministic step, a sample of the same thing).
- There is one shared, mutable cache for the whole project, keyed by address. Every job in the project’s database schema reads and writes it. Separately, each job keeps its own immutable record of which result it took at each address. Clearing or overwriting a cache entry later never changes what a finished job used.
- A lookup compares fingerprints. When a job reaches a step, it looks at the entry at the step’s address:
Cache addresses
Every operation in a job has a pathname, a readable address built from the names you give things. The root workflow’s pathname is theroot you pass when launching the job, "root" by
default. Each step, child workflow, and source array adds its local name under its parent, and a
mapped or replicated invocation adds its keys:
A step’s pathname is its default cache address. That has two consequences worth internalizing:
- Two jobs share results when their pathnames match. The same
root, the same local names along the way, and the same mapped keys give the same address. Rename a step in the workflow, move it into a child workflow, or change a dimension’s name, and its results get new addresses, so the next job computes them afresh. - The root name is a namespace. Because every address starts with
root, keeping the root the same across launches is how you reuse results, and changing it is how you start from a clean slate without touching anything. A new root is the simplest way to draw fresh samples from a nondeterministic step, since the job ID is not part of the address and a new job under the old root would reuse the old samples.
/, [, or ] in a local name, and make names specific:
comparison_judge_config reads better in an address than config.
Customizing an address
Occasionally you want a step’s results to be shared under an address that does not follow from its position, for example a slow preprocessing step whose results several experiments should share whatever their roots are.run_step and run_child_workflow take an
override_cache_address_prefix argument that replaces the pathname-derived prefix. Mapped and replica
keys are still appended to it, and a child workflow passes its prefix down to everything beneath
it.
CacheAddressReuseError), which almost
always means two scheduled runs were given the same override prefix.
Clearing the cache
fxtr cache list shows the cache, and fxtr cache clear forgets entries so that the next job to
ask for them runs the steps afresh. Jobs that already used a cleared result keep it.
*. Give every key of the group, in the
address’s order, and quote the pattern in the shell. --step NAME narrows any selection to the
entries one step computed, and --state to entries in one state. From Python, the project client
has list_cache_entries and clear_cache_entries, taking prefix=, addresses=, or conflicts_in=.
Resolving a job’s conflicts
When a job stops on cache conflicts,fxtr jobs status JOB groups them into patterns and prints
the commands that fix them:
- Clear and resume
- Overwrite
- New root
Clear exactly the job’s conflicts, then resume. You see each conflict before it goes.This is the safest approach and is recommended when you want to avoid clearing expensive work.
A downstream step’s inputs only change once the steps before it have rerun, so it may conflict
on the resume, and a pipeline of several stages can take a few rounds.
Reusing another job’s results
Clearing the cache does not mean that the data is lost! Every job maintains an immutable record of the exact step results that were used when it ran. If you want to restore the cache to a previous state, you can adopt the cache entries that were recorded in any of the jobs in the database using the commandfxtr cache adopt JOB. This
puts the results a job took back into the cache at their addresses, so later jobs reuse them.
You can also use this to import parts of the cache from a separate copy of the experiment database
(e.g. from a collaborator’s copy). The command fxtr jobs import FILE --adopt-cache both makes the
imported job available and adopts all of its cache entries so that they are used by local jobs.
Handling changes to step logic
fxtr tracks the exact version of code used to run any given step, and associates it with the cache entry. However, the step fingerprint only covers a step’s name and inputs, not its code. Editing a step’s body and launching under the same root reuses the cached results even if they are tagged as generated by an older code version. This is a deliberate choice, since it is difficult to predict what code changes can affect behavior, and it is often useful to significantly refactor code without invalidating previous results. We choose to simply track the differing code versions, and allow users to directly control whether or not their changes should invalidate the cache. To make sure that code changes do not lead to invalid results in the cache, we recommend following one of the following two best practices.Iterating on step logic during development
While you are first writing a step, you will often make several breaking changes in a row. After each breaking change, you can clear the cache for that step so the next job does not pick up a stale result:Adding configuration to stable steps
Once a step’s logic has stabilized and real results depend on it, or once you have shared your code with collaborators, you should generally avoid changing its behavior in place. A change in place forces you to remember to clear the cache and to track which version of the code produced each result, and the system cannot protect you if you forget. Instead:- To change or extend what a step does, add a configuration argument rather than editing the logic. Give the new argument a default that preserves the old behavior. Existing results keep matching, because their inputs are unchanged, and new results computed with the new option have different fingerprints and so can never be confused with old ones. (You can also bundle configuration arguments into a config entity, and handle configuration changes by adding a field to that entity.)
- If the change is too large for a flag, give the step a new name, for example by adding a
version suffix:
qa.answerbecomesqa.answer_v2. The new name gets fresh cache entries, old results stay intact and attributable, and jobs that still reference the old step continue to work.