The research problem
Research legitimately creates many datasets. Confusion begins when a derived file cannot be connected to the exact source state, transformation script or decision that produced it.
How the workflow works
- Choose the exact source file or historical version.
- Create a named derived dataset and record its research purpose.
- Review the Dataset family tree and, when code is linked, the reproducibility map.
What SepiaLog provides
- Original data remains unchanged
- Derived files stored separately
- Source file and exact version recorded
- Family tree and script-to-data connections
Common questions
Is branching only for code-based projects?
No. Branches are useful whenever a new dataset is derived, including spreadsheet cleaning and anonymization.
What should a branch reason contain?
State the purpose, important transformation and intended analysis, collaborator or access level.
Can one source have several valid branches?
Yes. Public, restricted, teaching and analysis-specific branches can all have distinct legitimate purposes.