What should researchers ask before using AI with sensitive data?

Ask where the data is processed, what happens to it after processing, and who controls access, storage and deletion. If those answers are unclear—or conflict with consent, ethics approval or institutional policy—do not upload the original research material.

You finish transcribing a set of sensitive interviews. The material is rich, complex, and potentially useful for analysis with AI.

Then one question stops you:

Where is my data actually going?

Is it leaving your computer? Is it being retained? Who can access it? Could it be used for another purpose? Are you still able to honour the conditions under which participants trusted you with their stories?

That hesitation is real, and it is reasonable.

Many researchers are interested in what AI can do but are unwilling to trade away control of sensitive material. This is not a lack of technical confidence. It is an ethical response to the responsibility researchers carry.

The upload is a research decision

Uploading a transcript, field note, image, or dataset to an online service is not merely a software action. It can change where the data is processed, which organisation handles it, how long it may be retained, and which policies apply.

Before using any AI service with research data, three questions deserve clear answers:

  1. Where is the data processed? On your device, on an institutional server, or in a third-party cloud environment?
  2. What happens after processing? Is the input retained, logged, reviewed, or used to improve another system?
  3. Who remains in control? Can you determine access, storage location, deletion, and the model being used?

If the answers are difficult to find, the uncertainty itself matters. Researchers should not have to infer privacy behaviour from a marketing promise or a vague settings screen.

Trust is part of the data

Participants do not share only words or numbers. They share experiences, identities, locations, health information, political views, workplace details, and relationships.

Even when direct identifiers are removed, combinations of details can remain revealing. A transcript may be technically pseudonymized and still contain a distinctive life history. A small geography, rare occupation, or unusual event can make a person recognizable.

That is why good research data practice goes beyond checking whether a name has been deleted. It asks whether the complete workflow respects the trust on which the research depends.

The right question is not simply, “Can this AI summarize the transcript?” It is also, “Should this transcript be sent to this system under these conditions?”

Local-first changes the default

A local-first tool begins from a different assumption: the data should remain on the researcher's computer unless the researcher explicitly decides otherwise.

When AI runs locally, the model and the data are processed on the same machine. There is no automatic need to send the transcript to an external service. The researcher can work with the model while keeping the original material inside the existing storage and access environment.

Local processing is not a complete ethics policy. Researchers still need appropriate consent, lawful processing, access controls, secure storage, and institutional guidance. But local-first design removes one important source of uncertainty: the silent movement of data into infrastructure the researcher does not control.

It also makes the software's behaviour easier to explain. “The material remains on this computer” is a clearer starting point than a chain of service providers, retention exceptions, and account settings.

A researcher using SepiaLog Local AI on her laptop while sensitive interview materials remain beside her
The application view is genuine SepiaLog footage composited into the editorial photograph. Local AI can run with a model on the researcher's own machine.

A practical pause before using AI

Before placing research material into any AI workflow, pause and check:

  • Does the material contain direct or indirect identifiers?
  • Does the consent process cover this kind of processing?
  • Where will the input be processed and stored?
  • Is the input retained or used beyond the immediate task?
  • Can a local model provide the assistance you need?
  • Would you be comfortable explaining the workflow to a participant or ethics reviewer?

If any answer is unclear, use a lower-risk version of the material, remove unnecessary detail, choose a local workflow, or seek guidance before proceeding.

The pause is not an obstacle to innovation. It is how responsible innovation happens.

Control should not be a premium feeling

At SepiaLog, the philosophy is simple: you stay in control, and your data stays with you.

The app is designed to run on your computer. There is no forced account just to begin. SepiaLog does not silently upload your research files. Its Local AI workflow can use a model running on your own machine, keeping project history and generated summaries inside the local environment you control.

This approach is not about adding a privacy label after the product is built. It is about treating participant trust as a design requirement from the beginning.

AI should help researchers understand their work. It should not make them uncertain about who else may now hold it.

Good research starts with trust. Good research software should protect it.

Sources and further reading

About the visuals: editorial photographs in this article are AI-generated illustrations. The SepiaLog interface described in the caption is a genuine product screen composited into the scene; the person is illustrative and is not an endorsement or research participant. Read our editorial and image policy.