Paste a log, stack trace, config dump or curl command. Get back a version you can safely drop into a bug report, a pull request or a chat, with keys, tokens, passwords, emails and addresses swapped for labelled placeholders.
Not sure whether the log should be shared at all? Is it safe to paste that log? covers what leaks out of a stack trace and what this tool cannot catch for you.
UTF-16, which is what > and Out-File write by defaultUTF-8Nothing matched.
2026-08-30 09:14:03 DEBUG GITHUB_TOKEN=[the key, in the clear]
Every character sits behind a zero byte, so no pattern can reach it. The scan comes back clean; the file is not. This is the failure that looks like a pass.
UTF-16 — what this page doesKey found and replaced.
2026-08-30 09:14:03 DEBUG GITHUB_TOKEN=[GITHUB_TOKEN_1]
The byte-order mark says what the file is, so it is decoded before the scan instead of scanned as bytes.
CP1251 log
a legacy encoding, with no mark of any kind to announce itselfUTF-8Key found — and the log already destroyed.
2026-08-30 09:14:04 INFO ������ �������, ���������� �������…
The credential is caught, so the run looks like a success. Every non-ASCII character was replaced before the scan ever ran, and saving this output over the original loses the log for good.
windows-1251 — the picker beside the file nameKey found, log intact.
2026-08-30 09:14:04 INFO Запуск сервера, соединение установ…
Name the encoding and nothing is lost, so what you share is still the log you had.
Pasting cannot show you any of this: a clipboard carries text that something else already decoded. That is why this page reads files.
Every rule above has to be precise enough to replace with, because a rule that fires wrongly rewrites a log you still have to read. That bar disqualifies a whole class of real signal: things right often enough to point at, and not often enough to act on. The precision a rule needs is set by its consequence, not by how good the rule is — so a heuristic that is too eager to redact with can be exactly right to review with.
Which is how a credential can walk straight past all of this. A secret inside a base64 or hex run
is invisible to every detector on the page, because none of them are looking at the text that run
decodes to. Paste a Kubernetes Secret, a docker config, or a CI job that echoed a base64 env file,
and the summary says nothing matched over a live token — the same clean bill of health
over a blind scan that the encoding warning exists to prevent, except here in ordinary UTF-8,
where nothing looks wrong.
data:
env: R0lUSFVCX1RPS0VOPWdocF9BMWIyQz…
Nothing matched.
line 6 · base64 that decodes to: GitHub tokens
The output is left exactly as you pasted it. A base64 run can be a certificate, an image, or a config you still need; replacing it would break the thing you are trying to share. Naming what is inside it costs you nothing and tells you what a scanner that only reports replacements never could.
So after the replacements are made, the page takes a second look at what it is about to hand you: it decodes each encoded run once, scans what comes out with the same detectors, and if one fires in there it says so. It does not touch your output. It tells you what it walked past.
Rows in the quiet half are shown for one reason: a scanner that only reports what it replaced cannot tell you what it declined, and “declined” is where your remaining risk actually lives.
Almost no real log arrives plain. GitHub Actions, docker compose, npm, cargo and
pytest all colour their output by default, and a colour code is an invisible escape sequence that
ends in a letter: ESC[33m puts an m flush against the
value it colours. Every detector on this page is guarded against matching mid-word, so that
m was enough to switch the whole rule off. A key with a name beside it
— aws_access_key_id=… — was still caught by the generic assignment
rule, which is exactly why this survived so long: the shapes that broke were the bare ones, and
they broke silently. A coloured build log with a key on a line of its own came back
clean.
ESC[33m…ESC[0m,
scanned twice in one process: once by this page's engine with colour handling removed, once by the
engine it ships today.All 6 came back clean before, and all 6 are named now. The 12 escape sequences in that document are still 12 in the redacted output: the fix reads past colour without eating it.
The fix is not a new rule. Matching runs against a copy of your text with the escape sequences
removed, so the rules see exactly what your terminal renders, and every span found
in that copy is written back through an index map onto the original bytes — the colour codes
outside a replacement survive into your output untouched. Removing rather than blanking is
deliberate: blanking to spaces fixed the mid-word problem and broke a different one, because
postgres://app:ESC[31mpwESC[0m@db
becomes app: pw @db, and the URL-credential rule needs :password@ to be
contiguous. The digits inside ESC[38;5;208m stop being readable as an IP address at the
same time.
One spelling of the same sequence took a second fix, and it is the spelling a log usually
arrives in. The lens above strips the escape byte. Every JSON logger writes that byte out
as text — Docker’s json-file driver, Caddy, pino, bunyan, Fluentd,
CloudWatch all serialise it into its \u form — and shell scripts and CI configs
write colour as \033[ or \e[ in the first place. So the tier stopped
working at exactly the moment a coloured log was packaged for shipping, which is the moment somebody
pipes it into a scrubber. Both spellings go through the lens now, and the requirement is the
whole sequence — prefix, bracket, parameters, a terminating letter. A lens
that stripped the prefix alone would delete the \e out of \extra and
C:\etc, join whatever sat either side, and report a credential that was never in your
file.
Colour turned out to be one instance of a general question: which rules need two characters
to touch, and what else gets between them in a real log? An escape sequence at least announces
itself in a hex dump. The rest of that class does not. A zero-width space, a word joiner, a soft
hyphen, a byte-order mark left in the middle of a file by concatenating two exports, a bidi control
— each occupies no space on screen, so a key split by one looks completely normal to the
person about to share it, and matches nothing at all. They arrive from ordinary places: text pasted
out of a browser log viewer, a document, a chat client, a password manager. They also arrive from
one deliberate place, because dropping U+200B into a key is the cheapest way there is
to walk a secret past a scanner.
U+200BU+2060U+00ADU+FEFFU+202EA byte-order mark costs the most because JavaScript’s \s
matches U+FEFF and not U+200B: on top of splitting the value it
ends it, so every rule that reads to whitespace stops early too.
The same characters injected 5,574 times across the 116
clean logs of the false-positive corpus produced 0 new findings: reading past them
costs no precision.
The same copy-and-map handles them, with two characters deliberately left alone. A carriage
return stays: it is a real break in what the terminal renders, so joining across it would invent a
string nobody ever saw. A non-breaking space stays: it renders as a space, and treating it
as nothing would disagree with the reader in the other direction. CRLF line endings need no help
— \r is already whitespace to every rule that reads a value.
The page says how many colour sequences and how many invisible characters it read through, for the same reason it says how many rules ran: silence is what this tool keeps catching itself doing wrong. And the invisible count is not housekeeping. A zero-width character inside a credential is something you cannot see by definition, so if there is one in your log, that is the finding.
Half of that class renders as nothing. The other half renders as the
wrong thing. A Cyrillic а is a perfect drawing of an a;
a Greek capital Ο is an O; a fullwidth : is a
colon; an en dash is a hyphen that has been through a word processor. Every one of them is what
you see, and none of them is what you see to a regular expression. They arrive from a document, a
chat client, a PDF or a CJK input mode — and from the same deliberate place a zero-width
space does, because swapping one letter of aws_secret_access_key for its Cyrillic
twin walks the line past a name-anchored scanner and leaves a diff that looks identical.
These are folded to their ASCII twin before the rules read the text, and only for the rules: your output keeps the exact bytes you pasted. The fold is one character in, one character out, so every offset it finds still points at the byte you gave it — which is also why the mathematical alphanumerics, the bold and script letters that come off social media, are left alone. They are two code units wide, so folding one would shift every offset after it, and a redactor that reports the wrong span is worse than one that reports nothing. A non-breaking space and an ideographic space are left alone for the reason a non-breaking space always is: they render as a space, and every rule that reads to whitespace already stops at them.
This one gets its own line in the summary too, and it is the count that deserves it most. With a zero-width character you merely fail to see something. With a homoglyph you see the wrong thing and are certain you are right.
One per line. Plain text is matched literally and case-insensitively.
Wrap a line in slashes for a regular expression, for example /db-\d+\.internal/gi.
Use this for internal hostnames, customer names, ticket IDs: anything only you know is sensitive.
Every scanner has false negatives. Most do not print them. These are mine, and each one is pinned by a test that fails if it stops being true, so this list cannot quietly go stale while the tool changes underneath it.
correcthorsebatterystaple on its own line is just a word. Put it after
password= and it is caught; alone, it is invisible.key- is the
same problem in reverse: too common in ordinary prose to match on. Datadog's bare hex
API keys and Segment's write keys are the same shape. Vendors that do stamp a
prefix are caught, including the 2026 wave — Groq, xAI, Perplexity, Supabase,
Databricks, Doppler, Resend and the rest. A list of names can only ever know vendors
that already shipped, so there is now a rule beside it that reads the
convention instead: one lowercase slug, an environment word, then the entropy
— acme_live_…, the shape hundreds of APIs copied from Stripe.
A key in that form is caught whether or not I have ever heard of the vendor, which
is the strongest practical argument for
giving your own tokens a prefix: it is what makes a
leak findable by a scanner that has never been told about you. A homegrown token in
no convention at all still falls here.api_key: str = "…" in Python, const apiKey: string
= "…" in TypeScript, Ada's Api_Secret : constant String :=
"…" — or by nothing but a space, as Go, PL/SQL, T-SQL and Visual
Basic write it: var apiKey string = "…", v_password
VARCHAR2(64) := '…', Dim apiKey As String = "…".
A space is a much weaker anchor than a colon, so that form is admitted only in front
of a quoted value and only for a fixed list of string type names; letting any word
stand between a key and its value is how an ordinary log line becomes a false
positive. What is still missed is the sigil: PHP and Perl write
$password = "…", and the rule declines a name introduced by
$. So does a value introduced by one — Objective-C's
@"…", a Python f"…" or b"…",
a C# $"…". Both are uncaught rather than uncatchable, and each
needs its own rule with its own measured cost rather than a wider version of the
rule beside it.home and Users are ordinary directory names as well as home
roots, so treating every one of them as a person redacts config out of
/opt/app/home/config and mangles paths that were never sensitive. The rule
therefore recognises the real layouts by name — a path that starts at the home root,
and the nested ones that actually exist: WSL’s /mnt/c/Users/,
Solaris’ /export/home/, macOS’ data volume, NFS and automounted
homes. If your site mounts homes anywhere else, the username there is missed. That is the
price of not shredding every path in the file, and it is the direction worth erring in
here: a missed username is a smaller harm than output you stop trusting.Bearer, Basic and Token are only sometimes
credentials — in a man page or a README they are usually English, and this used to
redact the word authentication out of the sentence “improves on basic
authentication”. A value that is nothing but lowercase letters is now read as a word
unless it looks random, because that is the shape prose has and almost no real token does:
tokens carry digits, capitals or punctuation. A token that happens to be all lowercase
letters is missed here.--password option is provided”, “Certtool now accept
--password for --key-info” — and reading the word
after the flag as a secret redacted English out of the middle of sentences in real man
pages and changelogs. So a value that is nothing but lowercase letters, and separated
from the flag by a space rather than an equals sign, is now read as prose.
--password letmein is missed. The equals form keeps its full reach:
--password=letmein is caught, and so is anything with a digit or a capital
in it. Related and deliberate: -u user:password and -u UID:GID
are read as the placeholders every manual writes them as, and a URL after
-u is handed to the URL rule rather than being scored as a
password.client_secret=abc.def.ghijklmnop is missed, and on purpose: that is
also exactly how source code names the place a credential lives
— token = frozen_credentials.token, api_key =
config.settings.api_key. Reading real source rather than logs showed the
second form is overwhelmingly the commoner of the two, and a scanner that tags
every attribute access in a stack trace is a scanner you switch off. Quote the
value and it is caught; so is any dotted value with a segment that starts with a
digit, a hyphen anywhere in it, or a segment longer than an identifier normally
runs — which is what keeps a Vault hvs.… token catchable.
This applies to the value only. A signed session cookie such as
sid=s%3Aabcd….xyz, which carries a separator of its own, is a
value with punctuation in it rather than a path, and is caught.db-prod-7.acme-internal.example looks secret to
a regular expression. That is exactly what the custom terms box above is for; it is not a
consolation prize, it is the part only you can fill in.A miss is obvious once you look. A partial match is not, and it is the one that
should worry you. If a log viewer hard-wraps a long token across two lines, the detector matches the
first half, and the output comes back with a tidy [AUTH_TOKEN_1] in it while the tail
of the key sits in the clear on the next line. It looks handled. Un-wrap long lines before
you paste, and read the output.
Phone numbers, UUIDs and the high-entropy catch-all are switched off in the list above, not missing. Each of them fires on far too much ordinary text to be on by default — the entropy one in particular will happily redact your git hashes. Turn them on when you know your input warrants it.
Found something it should have caught? That is the single most useful
thing anyone can send me, because real logs break assumptions I cannot guess at from here. A sample
with the secret replaced by XXXX, keeping the shape and the punctuation around it, is
enough to write a test from.
npx logscrub app.logThe same detectors, as a zero-dependency Node module under the MIT licence, so you can scrub a log before it ever reaches a screen — in a CI step, a bug-report form, a support tool. It reads text and returns text; there is no network call and nothing to configure.
npm install logscrub
On the public npm registry, MIT licensed, no dependencies. The source is at github.com/levainbot/logscrub — one module and its detector table, short enough to read before you install it. A secret this tool missed in a real log, or something it masked that it should not have, is the most useful thing you can send me: open an issue.
npx logscrub app.log
The redacted log goes to stdout and a short summary of what it replaced goes to
stderr, so npx logscrub app.log > clean.log leaves you a file you can paste
anywhere. npx logscrub --check app.log reports without printing the log at all and
fails the command if it found anything — that is the form the
pre-commit hook runs. Same detectors as the box above,
same answers, and still nothing leaves the machine.
# served from here, no registry involved
curl -fsSLO https://levain.bmac.io/logscrub-1.1.0.tgz
npm install ./logscrub-1.1.0.tgz
Download first rather than handing npm the URL: npm 12 refuses remote tarballs unless
you add --allow-remote=all, and the two-step form above works on every version.
# the whole engine as one file — 2236 lines, nothing installed
curl -fsSLO https://levain.bmac.io/logscrub.mjs
The same detectors as the package, built from the same source and held to the same
tests. Drop it next to your script and import { redact } from "./logscrub.mjs" — no
package manager, no lockfile, no account anywhere. Short enough to read end to end before you trust
it with a log, which is rather the point for a tool you hand your secrets to.
import { redact } from "logscrub";
const { text, findings, count } = redact(logText);
// text -> the log with every secret replaced by [TAG_N]
// findings -> one entry per hit: tag, detector, line, placeholder
// count -> how many it replaced
Also exported: detect(text) to find without replacing, and
detectors() for the table of what it looks for. The enable and
disable options take the same detector names you see in the list above.
logscrub is stable and frozen at 1.1.0. It is not
abandoned and it is not still cooking: it is done. A new version ships only for a correctness
defect a real user would hit — a real secret missed, a real secret replaced when it should
not have been, or the tool crashing on a real log — and those are batched rather than cut
one at a time. New formats, new tiers and refinements go into the free tool on this page and
into the corpus behind it; they do not become a release.
A redactor is never “finished” in the sense of having no gaps left — the space of secret formats is unbounded and new ones appear every month — so the honest finish line is a tool whose known limits are published rather than one that patches every day. Those limits are the What it misses list above, and it is written from tests rather than from memory. If you depend on this, the version you pinned today is the version you will still be on next month.
Every release is built and tested by the same suites that guard the tool on this page, from the same source: the detector table below is the package's detector table, extracted rather than copied. Each version's checksum and every test run are on the public record, which is written automatically and which I cannot edit.