Understand · SepiaLog feature

Dataset branching and lineage from source to every derivative

Create cleaned, anonymized and analysis-specific datasets without overwriting the original, while preserving each branch’s parent and purpose.

What is dataset lineage in research?

Dataset lineage is the documented relationship between a source dataset and every cleaned, anonymized, harmonized or analysis-specific dataset derived from it. It preserves the parent file, exact source version, transformation and research purpose behind each branch.

The research problem

Research legitimately creates many datasets. Confusion begins when a derived file cannot be connected to the exact source state, transformation script or decision that produced it.

Research actionConnected documentationRecoverable output

How the workflow works

  1. Choose the exact source file or historical version.
  2. Create a named derived dataset and record its research purpose.
  3. Review the Dataset family tree and, when code is linked, the reproducibility map.

What SepiaLog provides

  • Original data remains unchanged
  • Derived files stored separately
  • Source file and exact version recorded
  • Family tree and script-to-data connections

Common questions

Is branching only for code-based projects?

No. Branches are useful whenever a new dataset is derived, including spreadsheet cleaning and anonymization.

What should a branch reason contain?

State the purpose, important transformation and intended analysis, collaborator or access level.

Can one source have several valid branches?

Yes. Public, restricted, teaching and analysis-specific branches can all have distinct legitimate purposes.