Data handling
Do our documents get uploaded to your servers?#
No. Files are opened, unpacked and read in your browser, on your own machine. PST and zip unpacking, text extraction, OCR and the redaction burn all run locally. The original documents never leave your computer.
Then what does leave the machine?#
Extracted text, sent to a commercial AI API for entity detection. Not the files, not images, not attachments. Under the commercial terms that apply, that text is never used to train models and is deleted from the provider's systems within 30 days. The sub-processor is named in our disclosure and in the DPA.
What do you retain after the case closes?#
Nothing of your documents, because we never hold a copy at any point. What persists is the case record in your own browser storage and whatever you downloaded. Clear the browser data and the case is gone.
Is there a data processing agreement?#
Yes, published for UK GDPR, EU GDPR and US CCPA/CPRA. It is executed with you during onboarding and in place before your first upload. Where your own paper is required instead, that is handled as a negotiated project DPA.
Where is the service hosted?#
The application is served from a US-hosted platform and the detection API is US-operated. The distinction that usually matters for a data residency review is that your documents stay on your own machine, so the personal data in the files is not transferred to either. Extracted text is, under the terms described above. There is no UK-hosted or EU-hosted deployment at present, so if residency is a hard requirement rather than a preference, raise it before the pilot.
What authentication do you support?#
Email and password today. Access is governed at the organisation level: your named users are provisioned by us and inherit your organisation's executed agreements, so joiners and leavers are controlled through us rather than through a shared login. Single sign-on, MFA and role-based access control are not in the product yet. If your security review treats any of those as a requirement rather than a preference, raise it before the pilot and we will tell you plainly where it sits.
Can anyone at SafeRedact see our case?#
No. You run the case yourself, and there is no operator view into its contents. If you want help with a specific document, you decide what to send us.
What it does
What problem does this actually solve?#
The bottleneck in a subject access request is not finding the documents. It is removing third-party personal data from them before disclosure, at a volume where reading every page is not realistic. SafeRedact finds that data across the whole collection, puts it in front of you for a decision, and burns the approved redactions into the output.
How does detection work?#
Two layers. Pattern matching catches structured identifiers. An AI layer reads context and catches what patterns cannot: names in prose, indirect references, personal details embedded in a sentence. Neither layer redacts anything on its own.
Does it redact automatically?#
No. Review is mandatory. Nothing appears redacted in the output that a person has not approved, and nothing is approved in bulk unless you choose to do so. The system's job is to ensure nothing reaches you unflagged, not to make the disclosure decision for you.
How do we keep the requester's own details visible?#
Enter them as the case subject and in the aliases box. The preserve model keeps the data subject's own identifiers unredacted so the disclosure stays legible to the person who asked for it. Specific identifiers such as email addresses, ID numbers and full names are preserved wherever they appear. Ambiguous ones such as a bare first name are preserved where context anchors them to the subject and flagged for your judgement elsewhere. The same applies to associated individuals you decide should remain visible.
Are redactions removed, or just covered over?#
Removed. Redactions are burned into the page, so the underlying text cannot be recovered from the delivered file by selecting, copying or extracting.
Files and formats
What can we upload?#
PDF, Word, Excel, CSV, JSON, text and HTML documents, email as PST, MSG or EML, and zip archives including nested ones, which are unpacked recursively. A Microsoft Purview export loads as it arrives. Two things to know: for mailbox content export as PST rather than .msg, because a PST is expanded message by message with every attachment scanned as a document in its own right whereas content inside a .msg attachment is not read; and standalone image files such as .jpg or .png are not currently processed, so review those separately. Scanned pages inside a PDF are OCR'd normally.
What about scanned documents and image-only PDFs?#
OCR runs on them locally in your browser. Pages read with low confidence are flagged for manual review rather than passed off as clean, and the export warns you if any are still unresolved when you download. Two limits are worth knowing up front: handwriting is not machine-readable and will not be detected, and where a copier has already embedded a text layer, that layer is taken as given.
Is there a size limit?#
Individual files are subject to the browser's own ceiling of 2 GB per file, which is why we ask for the Purview PST package size to be set to 1 GB or 2 GB rather than the 5 GB default. Purview splits large mailboxes across several packages at that setting, and you load them all to the same case. Total case volume is agreed with you before the case starts.
What happens to a file that cannot be processed?#
A file that fails during processing is reported, not skipped quietly: per-file status is visible during the run and the case ends with a manifest of what failed and why. One thing the manifest does not cover is file types outside the supported list, such as standalone images, which are passed over without an entry. Reconciling the processed count against the item report from your Purview export is what catches those.
Running a case
How long does processing take?#
About an hour and a half per 2,500 documents, so a 10,000-document case is a working day and 20,000 is an overnight run. The timeline table gives the figure for each case size, alongside the export and download times that usually matter more.
Does the browser tab have to stay open?#
Yes, for the duration of processing. That is what keeps your files on your own machine. Before a long run, disable sleep, keep the machine on power, and check that device management will not force a restart overnight, which is the most common way a long case gets interrupted. Use a desktop browser on a machine with 16 GB of memory for cases above roughly 10,000 documents. The full list is under what the machine needs.
What happens if the tab closes or the machine restarts?#
Work is checkpointed file by file as it completes. Reopen the case under the same name, re-attach the same source files, and processing resumes where it stopped instead of starting over. Because the files themselves are never stored, re-attaching them is the one manual step. Renaming the case orphans the checkpoint, so keep the name.
Can several people work the same case at once?#
Not today. A case runs in one person's browser, which is the same design decision that keeps your documents off our servers, and there is no shared queue to hand a case between reviewers mid-run. Multiple people can hold accounts under your organisation and run separate cases in parallel. If a case needs more than one reviewer, split it by custodian.
Can we run more than one case?#
Yes. Cases are listed and can be reopened individually.
Should we load the whole collection at once?#
Start with one custodian. Processing a single mailbox first shows you what detection looks like against your own material and lets you correct the subject and alias details before committing the full volume. Catching that early is the difference between a clean run and reviewing the same mistake twenty thousand times.
Review and output
How long does review take?#
Less than the document count suggests. Review scales with the number of distinct people and identifiers in the collection rather than the number of documents, because a decision about a person applies everywhere that person appears. A large collection concerning few people reviews quickly.
What do we get at the end?#
The redacted documents, a manifest of what was processed and what failed, and an audit trail of the decisions taken. All of it belongs to your organisation.
Can we evidence what was done?#
The audit trail records the decisions and the export events, including any case where you chose to download with items still flagged. It is exportable and intended to be shown.
Can we reconcile against what we exported from Purview?#
Yes, and you should. Purview's own item report gives an independent count of what it exported, including items it could not retrieve at all. Compare that against the processed count and the failure manifest at the end of the case. The two failure layers are separate, and a complete reconciliation uses both.
Commercial
How is it priced?#
Per case. Not per seat, not per year. The price for a case is set with you and shown before you download. There is nothing to hold between requests, which suits organisations whose DSAR volume is occasional and unpredictable.
When do we pay?#
At download, after processing and review. You see what the case produced before you pay for it.
Is there an annual option?#
Available where volume justifies it, quoted from observed usage rather than guessed at in advance. Most organisations are better served per case.
Is a trial run charged separately?#
No, provided it stays in the same case. The recommended pattern is to process one custodian first, confirm the export settings and the review workflow came through cleanly, then load the rest into that same case. A case is priced once regardless of how many times you add to it. Creating a second case is what creates a second price.
Still unanswered
Anything we have not covered: support@saferedact.app, answered within 24 hours, Monday to Friday. Support is US-based and not staffed at weekends, so from UK or European hours a morning message usually gets a reply the same day, a late-afternoon one the following morning, and anything after Friday lunchtime on Monday. On a case with a deadline, raise blocking questions before you start a run rather than during it, and start long runs early in the week.
For a first case we are happy to spend 30 minutes on a screen share to set up the subject details and walk the review workflow. It is the fastest way to get the aliases box right, and that field determines how good the suggestions are.
Working through your first case now? The Getting Started guide covers the whole sequence, from the Purview export to the download.