Sensitive data removal · redactor.cohostly.cloud

Take the names and numbers out of your files.

Chat exports, screenshots, emails and spreadsheets go in. The same files come back with every name, phone number and amount replaced, and the conversation still readable.

See how it works
  • MB per batch
  • left today
  • files max
  • No account
  • No generative AI
  • Passphrase encryption
  • Deleted in min
Before support-chat.txt
09:12 Priya Nair: My order 4471 never arrived.
09:13 Agent: Which number is on the account?
09:14 Priya Nair: +91 98765 43210. I paid ₹2,499.
After support-chat.txt
09:12 [PERSON_01]: My order [NUMBER_01] never arrived.
09:13 Agent: Which number is on the account?
09:14 [PERSON_01]: [PHONE_01]. I paid [AMOUNT_01].

Same person, same label, every file in the batch.

Drag them here, or paste a screenshot with ⌘V. Nothing is uploaded until you press Redact.

ready

over limit, remove files too many files more than you have left today

The key is derived in your browser and never stored. Your redacted ZIP is encrypted while it waits on the server, so nobody with access to that machine can open it. Lose the passphrase and the file cannot be recovered.

Uploaded over this connection, redacted, then deleted from the server.

Done. of files were changed.

download expires in minutes

Processed
Redacted
Unchanged
Set aside

no matching files

counts only. The values themselves are never shown or stored

nothing detected

job failed

Coverage

What gets taken out

Google Cloud DLP handles the standard identifiers. Built-in rules on top of it catch the things a generic detector misses: deal terms, handles, campaign IDs, and loose numbers.

People & brands

Creators, managers, brand contacts. The same person keeps the same label across every file in the batch, so threads still make sense.

Aarav Mehta → [PERSON_03]

Phone numbers

Indian and international formats, including numbers sitting inside a screenshot.

+91 98765 43210 → [PHONE_01]

Money & terms

Amounts, percentages, payment schedules, credit periods and contract values.

50% upfront → [PAYMENT_PERCENTAGE]

Every other number

Timestamps, dates, follower counts, IDs. If it reads as a number, it goes.

22:47 → [NUMBER]:[NUMBER]

Contact points

Email addresses, URLs, IP addresses, @handles and campaign IDs.

@aaravbuilds → [HANDLE_01]

Screenshots

Text inside images is read with OCR and covered with a black box. Only the matches, not the whole bubble.

chat screenshot → black bars

accepted formats
Chats.txt · .json · .csv · .md
Email.eml, including attachments
Images.png · .jpg · .jpeg · .webp
Documents.pdf (text and scanned) · .docx · .xlsx
Anything elsecopied to _unprocessed/ untouched and listed in the report

How it works

Three steps, no sign-up

  1. 01

    Put the files in

    A ZIP, a handful of files, or an entire folder with its subfolders. Preview anything before it leaves your machine. The batch limit is MB.

  2. 02

    Detection and masking run server-side

    Each file is scanned, matches are replaced with labels, and your folder layout is kept exactly as it was. Anything unsupported is set aside rather than dropped.

  3. 03

    Check it, then take it

    Open the redacted files in the browser, look at the per-type counts, and download one ZIP. Everything is erased from the server after minutes.

You upload 4 files
exports/
  support-chat.txt
  invoice-april.pdf
  screenshot-01.png
  contacts.csv

Redact

You download redacted.zip
exports/
  support-chat.txt   12 replaced
  invoice-april.pdf  7 replaced
  screenshot-01.png  5 boxed out
  contacts.csv       96 replaced

Data handling

What happens to your files

Plainly, with no marketing around it.

No generative AI is involved at any step. No LLM, no chatbot, no model that could memorise your text or repeat it to someone else. Detection is regular expressions and word lists from this application's config, plus Google Cloud DLP, a deterministic classifier that returns the position of things it recognises. Nothing you upload is used as training data. The same input always produces the same output.

ItemWhere it livesHow long
Uploaded fileTemp directory on this serverDeleted the moment the job ends
Extracted originalsTemp directory on this serverDeleted the moment the job ends
Name → label mappingServer memory only, never diskCleared when the job ends
Redacted ZIPTemp directory, AES-256 encrypted when you set a passphrase minutes, then deleted automatically
Processing reportInside your ZIPYours. Counts per type, never values
AccountsNonen/a
Product analyticsPage views and feature usage (PostHog)Never includes file names or file content
Your passphraseNever sent. The key derived from it lives in server memory for the length of one requestDiscarded when the request ends
Training data, model fine-tuningNone. No generative model is involved at any stepn/a

One honest caveat. Detection is automated and not perfect. A nickname in a screenshot header, an unusual spelling, or heavily mixed-language text can slip through, and ordinary words occasionally get masked by mistake. Read the output before you pass it on.

Hosting

Who runs this

This siteOperated by Cohostly at .
File processingRuns on Cohostly's own servers. Files are never handed to a storage bucket, queue or CDN.
Detection engine
What leaves the server
Front-end assetsFonts and UI components load from Google Fonts and jsDelivr. They see your browser's request for those files; they never see your data.

FAQ

Questions people actually ask

Privacy & security

No. There is no LLM, no chatbot and no generative model anywhere in the pipeline, so there is nothing that can memorise your text or reproduce it for someone else. Detection is two fixed mechanisms: regular expressions and word lists defined in this application's config, and Google Cloud DLP, a deterministic classifier that finds patterns such as phone numbers and names and returns their positions. Neither one writes your content anywhere for reuse, and Google Cloud DLP does not train on the content you send it. The same input always produces the same output, which is the point: you can check the rules rather than trust a model. Set a passphrase and the answer is: only during the seconds detection actually runs. Your browser turns the passphrase into a key with PBKDF2 and sends only that key; the passphrase itself never leaves your machine. The redacted ZIP is encrypted with AES-256-GCM before it touches the disk, so for the whole retention window the file sitting on our server is ciphertext that we cannot open. The key is held in memory for one request and then dropped.

The honest limit: detection has to read your text to find a name in it, so while the job runs the content exists in plaintext in memory. We are not going to claim otherwise. What the passphrase removes is the far larger window, the file at rest afterwards. If you lose the passphrase, nobody can recover the file, including us.
Your files are used for one thing: producing the redacted copy. The upload and its extracted contents are deleted as soon as the job finishes, the mapping from real values to labels is held in memory and never written to disk, and the report inside your ZIP carries counts rather than values. The redacted ZIP stays for minutes so you can download it, encrypted if you set a passphrase, then it is deleted on a timer whether you collected it or not. The server that receives your upload processes it, and Google Cloud DLP sees the text and images during detection (see below). There are no user accounts and no shared workspace. Basic, non-identifying product analytics (page views, feature usage) are collected separately from file processing, and file contents are never sent to analytics or written to the application log. Cohostly's operations staff hold administrator access to the machine and could in principle reach files during the short window they exist, which is true of any hosted service, and it is the reason the retention window is deliberately small. When Cloud DLP is in use, Google's handling of that content falls under your own Google Cloud agreement and data-processing terms, not ours. Not from the output. Labels like [PERSON_01] stay consistent inside a single job so conversations remain coherent, but the table linking them to real values is destroyed when the job ends and is never included in the download. Two separate jobs will not agree on numbering. No. No sign-up, no login, no email address. Yes. redactor.cohostly.cloud is served over HTTPS, so your upload and the redacted ZIP you download are encrypted in transit. The text sent to Google Cloud DLP for detection travels over HTTPS as well.

Using it

Good, not perfect. No automated detector is. Expect occasional misses on nicknames, contact names in screenshot headers, unusual spellings and mixed-language text, and occasional over-masking of ordinary words that look like identifiers. Preview the output and read it before sharing. MB and files per batch, as a ZIP, loose files or a folder. Uploaded ZIPs are checked for unsafe paths and decompression-bomb behaviour before anything is extracted. Nothing is dropped silently. Unsupported or failed files are copied into an _unprocessed/ folder inside your download and listed by name and reason in the report, so you can deal with them by hand. Not in the same batch. A ZIP goes up on its own. To mix things, select the files or the folder directly and they will be packed for you. Yes. Subfolders come back exactly as they went in, with the report and any _unprocessed/ folder added at the top level.

preview truncated