SafeRedact

Documentation

SafeRedact Enterprise documentation.

The same knowledge the in-app help assistant answers from, published in full. Knowledge pack HK-2026.08.25.1. Nothing here is a capability that has not been verified in the product.

What SafeRedact Is

SafeRedact is DSAR/PII redaction software. It finds third-party personal data across a whole document collection, puts every detection in front of a human reviewer for a decision, and burns the approved redactions into the disclosed output. Detection uses two layers: pattern matching for structured identifiers, and an AI layer that reads context for names in prose, indirect references, and personal details embedded in sentences. Neither layer redacts anything on its own; review is mandatory, and nothing appears redacted in the output that a person has not approved. Redactions are removed from the text, not covered over: the underlying text cannot be recovered from the delivered file by selecting, copying or extracting.

Data Handling

EU Processing Route

For organizations that have it enabled, detection runs on AWS infrastructure within the UK and EU (European regions including London) instead of the standard US-operated AI API. The route is set per organization, is fixed for each case at the moment the case is created, and every routing change is recorded in an audit trail. - What leaves the machine on the EU route is the same as on the standard route: extracted text only. Not the files, not images, not attachments. On the EU route, detection runs on AWS infrastructure that does not store or log the text, does not use it for training, and does not make it available to the model provider. The sub-processors for each route are named in the disclosure and in the DPA. - The EU route runs interactively only; background (batch) mode is not available on this route by design. The browser tab must stay open and the machine awake for the duration of every processing run. The close-the-tab-after-submission behavior described for background mode does not apply on this route. - Complete-set loading, per-file checkpointing, and resume rules apply unchanged: files already completed are skipped when re-added, and a closed tab resumes by reopening the case under the same name and re-attaching the same files. - Purview export guidance is unchanged on this route: set the PST package size to 1-2 GB, and extract archives locally before adding. - Processing durations: no timings are published for the EU route. Duration questions on this route go to support@saferedact.app. - Route enablement and route pricing questions go to support@saferedact.app.

== SUPPORTED FILES AND FORMATS (11 file types) == PST, PDF, DOCX, XLSX, EML, MSG, HTML, TXT, CSV, JSON, ZIP. - PDF, DOCX, XLSX, TXT, HTML, CSV, JSON process directly: text extracted, detections run, redacted output rendered to PDF. - EML: processed as full MIME messages, headers, recipients, body and attachments, each attachment routed through the matching extractor. Deepest coverage. - PST is the recommended mailbox export format. Each PST is unpacked in the browser: every message becomes an EML, attachments are carried across and scanned, including messages inside messages. Limits: up to 20 attachments per message, 25MB per attachment; anything beyond a limit is marked with an explicit placeholder, never dropped without trace. - MSG: message bodies, headers and recipients are detected, but attachment content inside a .msg is NOT examined. If a mailbox was already exported as .msg, treat attachments as unreviewed; re-export as PST or collect the attachments separately. - ZIP, including Purview SharePoint/OneDrive collections and nested archives, is unpacked recursively with contents routed by type. - NOT processed: standalone image files (JPG, PNG, TIF), whether uploaded directly or inside a container. Find them by filtering the Purview Items report by extension and handle them outside SafeRedact. Video and audio are out of scope entirely. - Scanned or image-only PDF pages ARE processed: OCR runs on them locally in the browser. Pages read with low confidence are flagged for manual review, and export warns if any are unresolved at download. Handwriting is never machine-readable and will not be detected. Where a copier already embedded a text layer, that layer is taken as given. - Size: browsers cannot open a file above 2 GB, which is why Purview PST package size must be set to 1 GB or 2 GB rather than the 5 GB default. Large mailboxes split across packages automatically; load them all into the same case.

Purview Export Settings

Two settings decide whether an export is usable: format and package size. Choose "Create PSTs for messages" and set maximum PST package size to 1 or 2 GB. The full recommended settings: - Export format: PSTs for messages (the only mailbox format where attachment content is read and scanned in its own right). - Maximum PST package size: 1 or 2 GB (browser ceiling is 2 GB per file). - Maximum .zip package size: 2 GB, must equal or exceed the PST size. - Select items to include: Indexed and partially indexed. Partially indexed is where scanned and image-only documents sit; the default leaves them out, and on a DSAR those are exactly the documents that cannot go quietly missing. - Access links (cloud attachments): On, latest version. Note the original search must have been scoped to include cloud attachments; if it was not, the collection itself may need re-running. - Separate folders or PSTs per location: On (keeps custodians separable). - Include folder and path of source: On. - Give each item a friendly name: On. - Condense paths to 259 characters: On. - Export type: Items with items report (report-only truncates strings at 255 characters). - Organize conversation into HTML transcript: On for Teams. Timing facts: mailbox content exports at roughly 2 GB per hour per mailbox and a single mailbox cannot be parallelised; export processes cancel automatically after 7 days; a single export is capped at 500,000 items; files over 5 GB are not exported at all; search exports are deleted 14 days after creation (review-set exports allow 30 days). Extract with 7-Zip rather than the Windows built-in extractor. Including Teams conversations widens each responsive message by a twelve-hour window either side.

Running A Case

Machine requirements: a desktop or laptop, not a phone or tablet; a current Chromium browser recommended (Edge or Chrome; Firefox and Safari also work); 16 GB of memory for cases above roughly 10,000 documents; sleep disabled and power connected; no forced restart or patching window overnight; not a remote/VDI session that times out; free disk space for the export. For interactive runs, the browser tab must stay open for the duration of processing; that is what keeps the files on the user's own machine. For runs submitted in background (batch) mode, the tab can be closed once the activity log confirms the batch was submitted; results are collected when the case is reopened. Extraction always happens locally before submission either way. Processing time, STANDARD ROUTE ONLY: about 1.5 hours per 2,500 documents. 5,000 is about 2.5 hours; 10,000 about 5 hours (a working day); 20,000 about 10 hours (overnight); 25,000 about 12 hours. Processing runs unattended once started. These figures are measured on the standard route and DO NOT APPLY to the EU processing route; no per-volume timings are published for the EU route. No per-volume timings are published for the EU route; send duration questions to support@saferedact.app. Interruption and resume: work is checkpointed file by file. If the tab closes or the machine restarts, reopen the case UNDER THE SAME NAME, re-attach the same source files, and processing resumes where it stopped. Renaming the case orphans the checkpoint. Because files are never stored server-side, re-attaching them is the one manual step. Background (batch) mode and loading the complete set: detection for large sets can run in background mode. It engages automatically when a processing run starts with 1,000 or more files attached; some organizations have it enabled from 50 files. The activity log states when a run has gone through in background mode. Loading the complete file set in one click is the way to benefit from it; adding one folder at a time can keep each run below the level where it engages, in which case very large files process one at a time in the open tab and can take hours. Files already completed are skipped automatically when re-added, so loading the complete set never redoes finished work or adds cost. Organizations using the EU processing route run interactively only. Session counter and progress: the file counter shows the files ATTACHED IN THE CURRENT SESSION, not the case total. Opening a new session and attaching only a new folder makes the counter show just that folder; nothing has been lost. Completed work is saved file by file on that machine and comes back into view when the full set is re-attached. IMPORTANT: the export contains what is attached at the moment of export, so the complete set must be attached before exporting. Where enabled for an organization, a Progress page also shows recent processing totals for the whole team across machines; it counts all submissions including trials and spot checks and does not state a case total. Cloud-synced folders (OneDrive, SharePoint sync): before processing, make sure the export folders are fully downloaded to the machine. In OneDrive, right-click the folder, choose "Always keep on this device", and let the sync finish. A file that is not fully present on disk can be read incompletely; the app checks each file as it reads it and names any file it cannot read fully, at the moment it happens, with instructions to re-add it. If a large file shows a warning that very little text was extracted, re-add it from a fully synced local copy and process again. Browser-local case model, the team rules: a case runs in one person's browser and stays with the browser and machine it started on. Use the same browser on the same computer from first upload through sign-off. NEVER clear that browser's browsing data while a case is open; the case record is exactly what would be cleared. Once the pack is downloaded, the browser copy no longer matters. Cases cannot be handed between reviewers mid-run and there is no shared queue. Teams and large cases: split by custodian at case creation, so each custodian's documents form a separately reviewable case run by one person on one machine. Start with one custodian first to confirm subject and alias details before committing the full volume. A trial custodian is not charged separately provided it stays in the same case; a case is priced once regardless of how many times material is added to it. Creating a second case is what creates a second price. Subject preservation: enter the requester's details as the case subject and in the aliases box. The preserve model keeps the data subject's own identifiers visible so the disclosure stays legible to the person who asked for it. The aliases box determines suggestion quality.

Review

Two views over the same detections. By file: split pane, document text beside its detections, for reading context. By entity: every detection of the same value grouped across the case with a count, for consistency; an entity-level decision records on every instance and the audit trail notes the cascade. Most reviews use entity view to clear high-volume repeated identifiers and file view where judgment is needed. Detection intents: Redact (third-party personal data, removed from disclosure) and Preserve (the data subject's own information, deliberately kept visible). Each detection can be approved, switched between redact and preserve, or dismissed as a false positive. Detections can be added manually by selecting text; manual additions are marked reviewer-added in the audit trail. The confidence dashboard summarizes per-file detection counts and confidence, and flags files with NO detections at all, which deserve a manual glance because silence can mean a scanned image. "Approve all" exists for the end of review, not instead of it. Review decisions save as you work, in the browser's own storage. Closing the tab and returning later resumes in place. Sign-off: export begins with an attestation, a named confirmation that review was performed and the output approved for disclosure. If files were never opened or carry unreviewed detections, the screen says so before attesting; proceeding is an explicit recorded choice.

The Export Pack

Download delivers a single ZIP: each source document in redacted form named to match its original (emails as redacted EML with paired PDF; documents and spreadsheets rendered to redacted PDF), plus: - Detection report (CSV): one row per detection across the case. - Audit trail (CSV): one row per file with detection counts, confidence, review status. The file-level accountability record. - Review decision log: per-decision history including entity cascades and manual additions. - Subject preservation summary: what was preserved as the data subject's own information, grouped by identifier. - Failure manifest: files that failed processing, with reasons. File types outside the supported list (e.g. standalone images) are NOT in the manifest; reconciling against Purview's own Items report is what catches those. The two failure layers are separate and a complete reconciliation uses both. - Attestation record: who attested, when, to what case state. Completeness argument: Purview's Items report says N items were exported; the audit trail plus failure manifest account for all N; the detection report and decision log account for every redaction; the preservation summary accounts for what was kept; the attestation record puts a name on the judgment.

== PRICING MODEL AND PAYMENT STATE == Enterprise DSAR work is priced per case, not per seat or per year. The price for a case is set with the customer and shown before download. Payment is taken at download, after processing and review; the customer sees what the case produced before paying. Processing and review themselves happen before any payment step. The download step is where the commercial gate sits. What a given user sees there depends on how their organization's account is set up: organizations with an agreement in place may have download access recorded against the organization, in which case download proceeds without a card; otherwise a checkout or a message about licensing appears at download. If a user believes their organization has an agreement in place but download is still asking for payment, or the state at download looks wrong in any way: this is an account-state question. Email support@saferedact.app with the organization name and the build stamp from the header, and the account team will confirm what is set against your organization.

Troubleshooting

== OPERATIONAL PLANNING (derive, using only these figures) == - Processing scales at roughly 1.5 hours per 2,500 documents, per machine, running unattended with the tab open. Machines process in parallel: three reviewers on three machines each process their own case at that rate simultaneously. - Purview mailbox export runs at roughly 2 GB per hour PER MAILBOX; one large mailbox cannot be parallelised, so adding custodians parallelises better than adding years. Export processes auto-cancel after 7 days; a single export caps at 500,000 items; files over 5 GB are not exported. - Download time is bounded by the customer's network and is measured in hours for multi-package exports. - Review time scales with the number of DISTINCT people and identifiers in the collection, not the document count, because entity-level decisions apply everywhere a person appears. It cannot be estimated from document count alone; processing one custodian first yields a real review rate to plan from. - Statutory clock under UK GDPR: one month from receipt, extendable by two further months where a request is complex or several arrive together; the user's DPO confirms which applies. Plan backwards from the deadline; the Purview export is usually the longest and least predictable stage. - When asked to plan a case schedule, combine these figures with the user's stated numbers, show the arithmetic, state assumptions, and flag review time as the step that needs a measured rate from their first custodian.

Support

support@saferedact.app, answered within 24 hours, Monday to Friday. Support is US-based and not staffed at weekends; from UK or European hours a morning message usually gets a reply the same day. On a case with a deadline, raise blocking questions before starting a run rather than during it, and start long runs early in the week. For a first case, SafeRedact offers a 30-minute screen share to set up subject details and walk the review workflow. Documentation pages: /enterprise/getting-started (full case walkthrough), /enterprise/purview-export-guide (export settings), /enterprise/formats (supported files), /enterprise/review-guide (review screen), /enterprise/export-pack (pack contents), /enterprise/faq (thirty common questions), /enterprise/dsar-guide (overview).

Run your next subject access request through it.

One real case under a signed data processing agreement, with support through the first week.