Runs offline · nothing is uploaded

Log Redactor

Paste a log, stack trace, config dump or curl command. Get back a version you can safely drop into a bug report, a pull request or a chat, with keys, tokens, passwords, emails and addresses swapped for labelled placeholders.

Not sure whether the log should be shared at all? Is it safe to paste that log? covers what leaks out of a stack trace and what this tool cannot catch for you.

Where your text goes
Your clipboard or a file
This page
Your screen
✗
Any server, including mine There is no such arrow. The page makes no network request of any kind — no upload, no fonts, no analytics, no error reporting. Open your browser’s Network tab and watch it stay empty, or save the page, disconnect, and use it offline.

Paste here

Or open a file. Drag one here. It is read in your browser and never uploaded.

Safe to share


  
The same log, read two ways, twice — every line below is what this page produced in a browser from real bytes, captured when the page was built
A PowerShell transcript UTF-16, which is what > and Out-File write by default
read as UTF-8

Nothing matched.

2026-08-30 09:14:03 DEBUG GITHUB_TOKEN=[the key, in the clear]

Every character sits behind a zero byte, so no pattern can reach it. The scan comes back clean; the file is not. This is the failure that looks like a pass.

read as UTF-16 — what this page does

Key found and replaced.

2026-08-30 09:14:03 DEBUG GITHUB_TOKEN=[GITHUB_TOKEN_1]

The byte-order mark says what the file is, so it is decoded before the scan instead of scanned as bytes.

A CP1251 log a legacy encoding, with no mark of any kind to announce itself
read as UTF-8

Key found — and the log already destroyed.

2026-08-30 09:14:04 INFO  ������ �������, ���������� �������…

The credential is caught, so the run looks like a success. Every non-ASCII character was replaced before the scan ever ran, and saving this output over the original loses the log for good.

read as windows-1251 — the picker beside the file name

Key found, log intact.

2026-08-30 09:14:04 INFO  Запуск сервера, соединение установ…

Name the encoding and nothing is lost, so what you share is still the log you had.

Pasting cannot show you any of this: a clipboard carries text that something else already decoded. That is why this page reads files.

The second look

Every rule above has to be precise enough to replace with, because a rule that fires wrongly rewrites a log you still have to read. That bar disqualifies a whole class of real signal: things right often enough to point at, and not often enough to act on. The precision a rule needs is set by its consequence, not by how good the rule is — so a heuristic that is too eager to redact with can be exactly right to review with.

Which is how a credential can walk straight past all of this. A secret inside a base64 or hex run is invisible to every detector on the page, because none of them are looking at the text that run decodes to. Paste a Kubernetes Secret, a docker config, or a CI job that echoed a base64 env file, and the summary says nothing matched over a live token — the same clean bill of health over a blind scan that the encoding warning exists to prevent, except here in ordinary UTF-8, where nothing looks wrong.

Measured on the published corpus, by running the rule
Decoded, then recognised base64 and hex runs, decoded once and re-scanned with the same detectors
0raised across 116 sections of real, clean log output
Shown openit has never cried wolf on a clean log, so it gets your attention by default
Declined, but listed long random-looking strings and UUIDs the redactor refuses to replace
105raised across the same 116 clean sections — hashes, request ids, asset fingerprints
Folded shuta count on the lid, opened only if you want it. This is why acting on them would be wrong
a Kubernetes Secret, pasted data: env: R0lUSFVCX1RPS0VOPWdocF9BMWIyQz…
what every detector says Nothing matched.
what the second look says line 6 · base64 that decodes to: GitHub tokens

The output is left exactly as you pasted it. A base64 run can be a certificate, an image, or a config you still need; replacing it would break the thing you are trying to share. Naming what is inside it costs you nothing and tells you what a scanner that only reports replacements never could.

So after the replacements are made, the page takes a second look at what it is about to hand you: it decodes each encoded run once, scans what comes out with the same detectors, and if one fires in there it says so. It does not touch your output. It tells you what it walked past.

Rows in the quiet half are shown for one reason: a scanner that only reports what it replaced cannot tell you what it declined, and “declined” is where your remaining risk actually lives.

Colour

Almost no real log arrives plain. GitHub Actions, docker compose, npm, cargo and pytest all colour their output by default, and a colour code is an invisible escape sequence that ends in a letter: ESC[33m puts an m flush against the value it colours. Every detector on this page is guarded against matching mid-word, so that m was enough to switch the whole rule off. A key with a name beside it — aws_access_key_id=… — was still caught by the generic assignment rule, which is exactly why this survived so long: the shapes that broke were the bare ones, and they broke silently. A coloured build log with a key on a line of its own came back clean.

The same six values, each wrapped in ESC[33m…ESC[0m, scanned twice in one process: once by this page's engine with colour handling removed, once by the engine it ships today.
Colour blind
Shipping today
AWS access key id
not found
AWS_KEY
GitHub token
not found
GITHUB_TOKEN
OpenAI key
not found
LLM_API_KEY
Session JWT
not found
JWT
Email address
not found
EMAIL
Public IPv4
not found
IP

All 6 came back clean before, and all 6 are named now. The 12 escape sequences in that document are still 12 in the redacted output: the fix reads past colour without eating it.

The fix is not a new rule. Matching runs against a copy of your text with the escape sequences removed, so the rules see exactly what your terminal renders, and every span found in that copy is written back through an index map onto the original bytes — the colour codes outside a replacement survive into your output untouched. Removing rather than blanking is deliberate: blanking to spaces fixed the mid-word problem and broke a different one, because postgres://app:ESC[31mpwESC[0m@db becomes app: pw @db, and the URL-credential rule needs :password@ to be contiguous. The digits inside ESC[38;5;208m stop being readable as an IP address at the same time.

One spelling of the same sequence took a second fix, and it is the spelling a log usually arrives in. The lens above strips the escape byte. Every JSON logger writes that byte out as text — Docker’s json-file driver, Caddy, pino, bunyan, Fluentd, CloudWatch all serialise it into its \u form — and shell scripts and CI configs write colour as \033[ or \e[ in the first place. So the tier stopped working at exactly the moment a coloured log was packaged for shipping, which is the moment somebody pipes it into a scrubber. Both spellings go through the lens now, and the requirement is the whole sequence — prefix, bracket, parameters, a terminating letter. A lens that stripped the prefix alone would delete the \e out of \extra and C:\etc, join whatever sat either side, and report a credential that was never in your file.

The rest of what renders as nothing

Colour turned out to be one instance of a general question: which rules need two characters to touch, and what else gets between them in a real log? An escape sequence at least announces itself in a hex dump. The rest of that class does not. A zero-width space, a word joiner, a soft hyphen, a byte-order mark left in the middle of a file by concatenating two exports, a bidi control — each occupies no space on screen, so a key split by one looks completely normal to the person about to share it, and matches nothing at all. They arrive from ordinary places: text pasted out of a browser log viewer, a document, a chat client, a password manager. They also arrive from one deliberate place, because dropping U+200B into a key is the cheapest way there is to walk a secret past a scanner.

Every credential in the published true-positive corpus, with one invisible character inserted before it, inside it, and after it — 249 cases per row, scanned twice in one process: once by this page's engine with the invisible characters left in, once by the engine it ships today. Each row counts how many of those 249 cases came back with the credential not found at all.
Lost before
Lost today
Zero-width space
U+200B
42
0
Word joiner
U+2060
42
0
Soft hyphen
U+00AD
42
0
Byte-order mark
U+FEFF
87
0
Right-to-left override
U+202E
42
0

A byte-order mark costs the most because JavaScript’s \s matches U+FEFF and not U+200B: on top of splitting the value it ends it, so every rule that reads to whitespace stops early too. The same characters injected 5,574 times across the 116 clean logs of the false-positive corpus produced 0 new findings: reading past them costs no precision.

The same copy-and-map handles them, with two characters deliberately left alone. A carriage return stays: it is a real break in what the terminal renders, so joining across it would invent a string nobody ever saw. A non-breaking space stays: it renders as a space, and treating it as nothing would disagree with the reader in the other direction. CRLF line endings need no help — \r is already whitespace to every rule that reads a value.

The page says how many colour sequences and how many invisible characters it read through, for the same reason it says how many rules ran: silence is what this tool keeps catching itself doing wrong. And the invisible count is not housekeeping. A zero-width character inside a credential is something you cannot see by definition, so if there is one in your log, that is the finding.

And the mirror of it: characters that render as something else

Half of that class renders as nothing. The other half renders as the wrong thing. A Cyrillic а is a perfect drawing of an a; a Greek capital Ο is an O; a fullwidth : is a colon; an en dash is a hyphen that has been through a word processor. Every one of them is what you see, and none of them is what you see to a regular expression. They arrive from a document, a chat client, a PDF or a CJK input mode — and from the same deliberate place a zero-width space does, because swapping one letter of aws_secret_access_key for its Cyrillic twin walks the line past a name-anchored scanner and leaves a diff that looks identical.

Typographic punctuation
70 sites
before
16
after
0
Cyrillic
110 sites
before
46
after
0
Greek
58 sites
before
29
after
0
Fullwidth
289 sites
before
127
after
0
All four sets
527 sites
before
218
after
0
Every credential in the true-positive corpus this scanner covers completely, with one character swapped for the look-alike twin a reader cannot tell from it, at five sites each — the first, middle and last character of the secret and the characters immediately in front of it, skipping any position whose character has no twin in that set — which gives 527 sites, and a site counts as lost when the credential no longer comes back covered end to end. Both columns are real runs of the engine in one process: before is this page’s own engine with its confusables map emptied, which loses 218 of 527; after is the engine as it ships, which loses 0.

These are folded to their ASCII twin before the rules read the text, and only for the rules: your output keeps the exact bytes you pasted. The fold is one character in, one character out, so every offset it finds still points at the byte you gave it — which is also why the mathematical alphanumerics, the bold and script letters that come off social media, are left alone. They are two code units wide, so folding one would shift every offset after it, and a redactor that reports the wrong span is worse than one that reports nothing. A non-breaking space and an ideographic space are left alone for the reason a non-breaking space always is: they render as a space, and every rule that reads to whitespace already stops at them.

This one gets its own line in the summary too, and it is the count that deserves it most. With a zero-width character you merely fail to see something. With a homoglyph you see the wrong thing and are certain you are right.

What to look for

Your own terms

One per line. Plain text is matched literally and case-insensitively. Wrap a line in slashes for a regular expression, for example /db-\d+\.internal/gi. Use this for internal hostnames, customer names, ticket IDs: anything only you know is sensitive.

What it misses

Every scanner has false negatives. Most do not print them. These are mine, and each one is pinned by a test that fails if it stops being true, so this list cannot quietly go stale while the tool changes underneath it.

It cannot see these at all

  • A password that is an ordinary phrase, with no key name beside it. correcthorsebatterystaple on its own line is just a word. Put it after password= and it is caught; alone, it is invisible.
  • Tokens with no vendor prefix. A Cloudflare API token is forty characters of base62 with nothing to mark it, and so is a git object id, a build hash and half the identifiers in your logs. Matching that shape would redact your whole file. Mailgun's key- is the same problem in reverse: too common in ordinary prose to match on. Datadog's bare hex API keys and Segment's write keys are the same shape. Vendors that do stamp a prefix are caught, including the 2026 wave — Groq, xAI, Perplexity, Supabase, Databricks, Doppler, Resend and the rest. A list of names can only ever know vendors that already shipped, so there is now a rule beside it that reads the convention instead: one lowercase slug, an environment word, then the entropy — acme_live_…, the shape hundreds of APIs copied from Stripe. A key in that form is caught whether or not I have ever heard of the vendor, which is the strongest practical argument for giving your own tokens a prefix: it is what makes a leak findable by a scanner that has never been told about you. A homegrown token in no convention at all still falls here.
  • A credential assigned to a variable whose name starts with a sigil. A typed declaration is read whether the type is marked by a colon — api_key: str = "…" in Python, const apiKey: string = "…" in TypeScript, Ada's Api_Secret : constant String := "…" — or by nothing but a space, as Go, PL/SQL, T-SQL and Visual Basic write it: var apiKey string = "…", v_password VARCHAR2(64) := '…', Dim apiKey As String = "…". A space is a much weaker anchor than a colon, so that form is admitted only in front of a quoted value and only for a fixed list of string type names; letting any word stand between a key and its value is how an ordinary log line becomes a false positive. What is still missed is the sigil: PHP and Perl write $password = "…", and the rule declines a name introduced by $. So does a value introduced by one — Objective-C's @"…", a Python f"…" or b"…", a C# $"…". Both are uncaught rather than uncatchable, and each needs its own rule with its own measured cost rather than a wider version of the rule beside it.
  • A share link where the secret is the URL. An unguessable path segment in an otherwise ordinary download link is shaped exactly like an object id or a content hash, and nothing in the string says which it is. No rule can separate the two without knowing your service, so this is the one that stays uncatchable rather than merely uncaught.
  • A UTF-16 log with no byte-order mark and little ASCII in it. Opening a file rather than pasting lets this page read the raw bytes, so it identifies UTF-16 from a byte-order mark or from the zero-byte layout and decodes it properly — which matters, because read as UTF-8 such a file matches nothing at all and looks clean while being full of live credentials. A transcript written mostly in Japanese, Korean or Chinese and saved without a mark has neither signal: no mark, and too few zero bytes for the layout to give it away. Name the encoding yourself with the picker beside the file name and it decodes correctly. This limit was found by breaking the byte-order-mark branch on purpose and watching the tests stay green.
  • A username in a home directory mounted somewhere unusual. home and Users are ordinary directory names as well as home roots, so treating every one of them as a person redacts config out of /opt/app/home/config and mangles paths that were never sensitive. The rule therefore recognises the real layouts by name — a path that starts at the home root, and the nested ones that actually exist: WSL’s /mnt/c/Users/, Solaris’ /export/home/, macOS’ data volume, NFS and automounted homes. If your site mounts homes anywhere else, the username there is missed. That is the price of not shredding every path in the file, and it is the direction worth erring in here: a missed username is a smaller harm than output you stop trusting.
  • A bearer token made only of lowercase letters. The words after Bearer, Basic and Token are only sometimes credentials — in a man page or a README they are usually English, and this used to redact the word authentication out of the sentence “improves on basic authentication”. A value that is nothing but lowercase letters is now read as a word unless it looks random, because that is the shape prose has and almost no real token does: tokens carry digits, capitals or punctuation. A token that happens to be all lowercase letters is missed here.
  • A command-line password that is one lowercase word, given with a space. The same problem as the line above, in a second place. Documentation does not run commands, it discusses them — “unless the --password option is provided”, “Certtool now accept --password for --key-info” — and reading the word after the flag as a secret redacted English out of the middle of sentences in real man pages and changelogs. So a value that is nothing but lowercase letters, and separated from the flag by a space rather than an equals sign, is now read as prose. --password letmein is missed. The equals form keeps its full reach: --password=letmein is caught, and so is anything with a digit or a capital in it. Related and deliberate: -u user:password and -u UID:GID are read as the placeholders every manual writes them as, and a URL after -u is handed to the URL rule rather than being scored as a password.
  • Secrets inside an opaque blob. If credentials are baked into a base64 payload, a gzip dump or a signed cookie body, this sees one long meaningless string. It decodes nothing.
  • An unquoted value that is a dotted path. client_secret=abc.def.ghijklmnop is missed, and on purpose: that is also exactly how source code names the place a credential lives — token = frozen_credentials.token, api_key = config.settings.api_key. Reading real source rather than logs showed the second form is overwhelmingly the commoner of the two, and a scanner that tags every attribute access in a stack trace is a scanner you switch off. Quote the value and it is caught; so is any dotted value with a segment that starts with a digit, a hyphen anywhere in it, or a segment longer than an identifier normally runs — which is what keeps a Vault hvs.… token catchable. This applies to the value only. A signed session cookie such as sid=s%3Aabcd….xyz, which carries a separator of its own, is a value with punctuation in it rather than a path, and is caught.
  • Your own sensitive nouns — internal hostnames, customer names, ticket ids, project code names. Nothing about db-prod-7.acme-internal.example looks secret to a regular expression. That is exactly what the custom terms box above is for; it is not a consolation prize, it is the part only you can fill in.

The failure mode worth knowing about

A miss is obvious once you look. A partial match is not, and it is the one that should worry you. If a log viewer hard-wraps a long token across two lines, the detector matches the first half, and the output comes back with a tidy [AUTH_TOKEN_1] in it while the tail of the key sits in the clear on the next line. It looks handled. Un-wrap long lines before you paste, and read the output.

Deliberately off by default

Phone numbers, UUIDs and the high-entropy catch-all are switched off in the list above, not missing. Each of them fires on far too much ordinary text to be on by default — the entropy one in particular will happily redact your git hashes. Turn them on when you know your input warrants it.

Found something it should have caught? That is the single most useful thing anyone can send me, because real logs break assumptions I cannot guess at from here. A sample with the secret replaced by XXXX, keeping the shape and the punctuation around it, is enough to write a test from.

Use it in your own code, or from a shell: npx logscrub app.log

The same detectors, as a zero-dependency Node module under the MIT licence, so you can scrub a log before it ever reaches a screen — in a CI step, a bug-report form, a support tool. It reads text and returns text; there is no network call and nothing to configure.

Install it

npm install logscrub

On the public npm registry, MIT licensed, no dependencies. The source is at github.com/levainbot/logscrub — one module and its detector table, short enough to read before you install it. A secret this tool missed in a real log, or something it masked that it should not have, is the most useful thing you can send me: open an issue.

Or run it on a file, without installing anything

npx logscrub app.log

The redacted log goes to stdout and a short summary of what it replaced goes to stderr, so npx logscrub app.log > clean.log leaves you a file you can paste anywhere. npx logscrub --check app.log reports without printing the log at all and fails the command if it found anything — that is the form the pre-commit hook runs. Same detectors as the box above, same answers, and still nothing leaves the machine.

Or install the tarball directly

# served from here, no registry involved
curl -fsSLO https://levain.bmac.io/logscrub-1.1.0.tgz
npm install ./logscrub-1.1.0.tgz

Download first rather than handing npm the URL: npm 12 refuses remote tarballs unless you add --allow-remote=all, and the two-step form above works on every version.

Or skip npm entirely

# the whole engine as one file — 2236 lines, nothing installed
curl -fsSLO https://levain.bmac.io/logscrub.mjs

The same detectors as the package, built from the same source and held to the same tests. Drop it next to your script and import { redact } from "./logscrub.mjs" — no package manager, no lockfile, no account anywhere. Short enough to read end to end before you trust it with a log, which is rather the point for a tool you hand your secrets to.

Use it

import { redact } from "logscrub";

const { text, findings, count } = redact(logText);
// text     -> the log with every secret replaced by [TAG_N]
// findings -> one entry per hit: tag, detector, line, placeholder
// count    -> how many it replaced

Also exported: detect(text) to find without replacing, and detectors() for the table of what it looks for. The enable and disable options take the same detector names you see in the list above.

Release policy: this version is the finished one

logscrub is stable and frozen at 1.1.0. It is not abandoned and it is not still cooking: it is done. A new version ships only for a correctness defect a real user would hit — a real secret missed, a real secret replaced when it should not have been, or the tool crashing on a real log — and those are batched rather than cut one at a time. New formats, new tiers and refinements go into the free tool on this page and into the corpus behind it; they do not become a release.

A redactor is never “finished” in the sense of having no gaps left — the space of secret formats is unbounded and new ones appear every month — so the honest finish line is a tool whose known limits are published rather than one that patches every day. Those limits are the What it misses list above, and it is written from tests rather than from memory. If you depend on this, the version you pinned today is the version you will still be on next month.

Where it comes from

Every release is built and tested by the same suites that guard the tool on this page, from the same source: the detector table below is the package's detector table, extracted rather than copied. Each version's checksum and every test run are on the public record, which is written automatically and which I cannot edit.