Blog
A task-board export needs an experiment index
A task board can show that work moved while losing the experiment behind it. Add an index that links each test to its question, evidence and result so later work can tell what is safe to reuse.

A task board can show that work moved while losing the experiment behind it. Add an index that links each test to its question, evidence and result so later work can tell what is safe to reuse. This becomes obvious when a project is exported or archived. The cards still have titles, assignees, dates and status changes. Attachments may survive. Yet the reason for a test, the configuration used and the meaning of "done" are often scattered through comments or left in someone's working notes. Six months later, a completed card can look like proof for a claim it never tested. An experiment needs an identifier that remains usable outside the task-board product. Board card numbers are convenient while the board exists, but they can collide across projects or disappear during migration. A short project prefix and sequence is enough if the team applies it consistently. The index entry should start with the question being tested. "Try new retrieval settings" describes an activity. "Does adding document date to retrieval reduce answers based on superseded guidance?" states what the team expects to learn. I would keep the initial record to the fields that support later interpretation:
- experiment identifier and date;
- question or hypothesis;
- owner and review status;
- input set or evidence source;
- configuration and relevant versions;
- result, limits and linked artefacts.
The task card can still carry delivery detail, discussion and subtasks. The index is the compact record that connects the question to the evidence. Each system links to the other by identifier. A result is easier to overstate when the original plan was never written down. Before running the test, record which cases will be used, what will change, what will stay fixed and how the result will be judged. This does not require laboratory ceremony for every product decision. It does require enough detail to separate a planned comparison from an interesting output found afterwards. If a team tries several prompts and keeps the best-looking example, the record should say that. It should not be described later as a controlled evaluation. Configuration includes more than the prompt. Model version, retrieval index, tool permissions, rule set, source snapshot and application release may affect the outcome. Save stable references where possible. If a live dependency cannot be fixed, record the date and how it may have changed. The input set also needs a name or version. "Tested on customer questions" is unsafe and too vague. Use an approved fixture set or a restricted reference that shows exactly which cases ran without copying private content into the index.
The index should point to outputs, logs, notebooks, screenshots, evaluation reports or pull requests. It should not duplicate every artefact. Duplication creates stale copies and makes access control harder. Each link needs enough description to survive outside its original interface. A bare attachment named results-final.csv will be difficult to interpret later. The index can state what the file contains, which run produced it and where the authoritative copy lives. Access matters. A board export may be shared more widely than the project workspace. Keep private datasets, credentials, personal information and restricted logs in their approved systems. The index can use a secure record identifier and note the required access. If the evidence cannot be retained, record the retention limit and avoid presenting the experiment as indefinitely reproducible. For generated outputs, retain the raw result where policy allows. A summary written after review can hide failures or omit awkward cases. The raw output, scorer result and reviewer note let a later reader understand how the conclusion was reached. "Successful" is rarely enough. The result should answer the original question and include the boundary of the test. If five defined cases improved and two regressed, say that. If the result was a qualitative preference from one reviewer, identify it as such.
A useful result note covers the observed comparison, the acceptance rule and anything that made interpretation uncertain. It can also state the immediate decision: adopt the change, reject it, run another test or keep it behind a controlled release. Negative and inconclusive results belong in the index. Otherwise later teams repeat them because the board only preserves the work that reached production. An inconclusive test may still reveal that the fixtures were too easy, a dependency was unstable or the proposed measurement did not reflect the user task. Avoid converting a local result into a broad claim. A prompt that worked on one approved assessment set has evidence for that set and configuration. Reuse in another workflow should begin by checking whether its inputs, consequences and review rules are comparable. Task status describes workflow. Experiment outcome describes evidence. A card can be complete because the test ran, even when the change failed. It can be cancelled because priorities moved, despite useful partial evidence. Combining these meanings makes both records less reliable. I use a small set of outcome labels that fit the work, such as supported, not supported, mixed or inconclusive. The note carries the detail. Review status is separate again because a result may be recorded before someone independent has checked the artefacts.
Corrections should append to the record rather than quietly replacing the old conclusion. If a scoring bug invalidates a result, mark it and link the corrected run. This retains the path that explains why an earlier decision changed. An archive process should export the experiment index in a plain, searchable format with stable identifiers and readable dates. It should also retain task-card references and resolvable links to artefacts that remain available. Test the export before the board is retired. Choose a few experiments and ask someone outside the immediate work to find the question, setup, evidence, outcome and resulting decision. Check one failed test as well as one adopted change. If they have to search chat history or ask the original owner what "successful" meant, the index is missing information. Fix those records while the work is still familiar, then make creation of the experiment entry part of starting the next test.
Discussion
Continue the thinking.
Comments are public and hosted in an open-source GitHub Discussions repository.
Loading comments connects your browser to GitHub. A GitHub account is required to post.