Regex cheatsheet for JavaScript — with a live tester that will not hang your tab
A regex reference for JavaScript (ECMAScript) specifically, because flavour is the first thing a cheatsheet can get wrong: \d is ASCII-only here and Unicode-aware in .NET, \A and \z do not exist here at all, and JavaScript has no atomic groups or possessive quantifiers. Every construct is listed with what it matches, and where it differs meaningfully from PCRE that difference is spelled out. The tester above shows matches, capture groups and a replace preview, and it runs your pattern inside a Web Worker the page can terminate — so a pattern that would otherwise freeze the tab comes back as a written message instead. Nothing is uploaded or saved.
- Nothing is uploaded or saved
- Works offline
Flavour: JavaScript (ECMAScript). Everything on this page is what the RegExp built into browsers and Node actually does. Where a construct behaves differently in PCRE, Python or .NET — and several do — the table says so on that row.
Write what goes between the slashes. A forward slash needs no escaping here.
u and v are mutually exclusive — turning one on turns the other off, because setting both makes the engine throw.
$1 and $<name> for groups, $& for the whole match, $` and $' for the text either side, $$ for a literal dollar.
44 of 5,000 characters. That cap keeps the highlighted output renderable — it is not what protects you from a runaway pattern; the worker below is.
Edit the pattern or the test string to run it.
Each one loads into the tester with a sample string, so you can watch it work rather than take it on trust. There is deliberately no email pattern here — see the FAQ for why every short one is wrong.
A quoted string, without .*
/"[^"]*"/gOn alt="one" title="two" — Two matches: "one" and "two".
Written as ".*?" it gives the same two matches here but backtracks on every character. The negated class cannot cross a quote at all, so there is nothing to backtrack into.
A doubled word
/\b(\w+)\s+\1\b/giOn the the quick brown fox jumped over over the dog — Two matches — "the the" and "over over" — each with the repeated word in group 1.
Thousands separators
/\B(?=(\d{3})+(?!\d))/gOn 1234567 — Two zero-length matches, at the two positions where a comma belongs. Replace with "," to get 1,234,567.
Integers only. Run it on 1234.5678 and it will punctuate the decimals too. Intl.NumberFormat does this properly, and localised.
A hex colour
/^#(?:[0-9a-f]{3,4}|[0-9a-f]{6}|[0-9a-f]{8})$/iOn #1a2b3c — One match. Also accepts #fff, #ffff and #1a2b3c80; rejects #12345.
The shape of an ISO date
/^(\d{4})-(\d{2})-(\d{2})$/On 2026-08-10 — One match with three groups.
Shape only. It happily accepts 2026-99-99. No regex can check whether a date exists — parse it afterwards and compare the parts back.
Trim leading and trailing whitespace
/^\s+|\s+$/gOn padded text — Two matches, one at each end. Replace with "" to trim.
String.prototype.trim() does this without a regex and without the edge cases. Reach for it first.
Strip slashes from both ends of a path
/^/+|/+$/gOn ///api/v1/users// — Two matches. Replacing with "" leaves api/v1/users.
Split a CSV line on commas outside quotes
/,(?=(?:[^"]*"[^"]*")*[^"]*$)/gOn a,"b,c",d — Two matches — the commas outside the quotes — so the line splits into three fields.
This is the honest limit of regex on CSV. It counts quotes ahead of the comma, so it breaks on escaped quotes ("" inside a field) and on embedded newlines. For a file rather than a line, use a parser.
Most characters stand for themselves. Twelve do not, and every one of them has to be escaped with a backslash to be matched literally.
| Construct | What it does |
|---|---|
| a 1 @ | An ordinary character matches itself. Nothing to escape. |
| \ | Escapes the character that follows, turning a metacharacter into a literal and an ordinary letter into a class. |
| ^ $ \ . * + ? ( ) [ ] { } | | The twelve characters that need escaping outside a character class. Inside a class the list is much shorter — see the next section./\$\d+\.\d{2}/ matches $19.99 |
| / | Needs escaping only inside a regex literal, because it would otherwise end the literal. In new RegExp("…") it does not. |
| \n \r \t \v \f \0 | Newline, carriage return, tab, vertical tab, form feed, and NUL. |
| \xhh | The character with the given two-digit hexadecimal code./\x41/ matches A |
| \uhhhh | The UTF-16 code unit with the given four hex digits./\u00e9/ matches é |
| \u{hhhhh} | A code point by number, up to six digits. Requires the u or v flag./\u{1F600}/u matches 😀 |
| \cX | The control character for letter X — \cJ is a line feed, \cM a carriage return. |
| RegExp.escape(s) | Escapes a string so it can be dropped into a pattern as a literal. It is a recent addition; if you need to support older engines, escape it yourself with s.replace(/[.*+?^${}()|[\]\\]/g, "\\$&"). |
A class matches exactly one character out of a set. The escaping rules inside a class are different from outside it, which is where most of the confusion lives.
| Construct | What it does |
|---|---|
| [abc] | Any one of a, b or c. |
| [^abc] | Any one character that is not a, b or c. The caret only negates in first position. |
| [a-z] | A range, by code point. [a-z], [0-9] and [A-Za-z0-9_] are the ones you will write. |
| ] \ ^ - | The only characters that ever need escaping inside a class. The caret only matters first, and the hyphen only where it could be read as a range — [a.+*] is a perfectly good class of four literal characters. |
| [\b] | Inside a class, \b is a backspace (U+0008). Outside a class the same two characters are a word boundary. Same spelling, unrelated meanings. |
| [[a-z][0-9]] | Nested classes — a union. Requires the v flag.vs PCRE — PCRE has no nesting; you write [a-z0-9]. |
| [\p{L}--[aeiou]] | Set difference: letters, minus those five. Requires the v flag.vs PCRE — No equivalent syntax in PCRE. |
| [\p{L}&&\p{Script=Greek}] | Set intersection. Requires the v flag. |
| \q{abc|de} | Matches one of several multi-character strings from inside a class. Requires the v flag, and it is the only way a class matches more than one character. |
These shorthands are where flavours disagree most, and where a cheatsheet that does not name its flavour does real damage. In JavaScript \d and \w are ASCII-only, always — the u flag does not change that.
| Construct | What it does |
|---|---|
| . | Any character except a line terminator: \n, \r, U+2028 and U+2029. With the s flag it matches those too. |
| \d \D | Exactly [0-9], and its negation. Nothing else, ever.vs PCRE — PCRE with (*UCP), .NET and Python 3 all make \d match Unicode digits such as ٣ and 3. JavaScript never does — use \p{Nd} with the u flag if you want that. |
| \w \W | Exactly [A-Za-z0-9_], and its negation. No accented letters, no CJK.vs PCRE — Same divergence as \d: Unicode-aware in several other flavours. In JavaScript, "café".match(/\w+/) returns "caf". |
| \s \S | Whitespace: space, \t \n \v \f \r, U+00A0, U+1680, U+2000–U+200A, U+2028, U+2029, U+202F, U+205F, U+3000 — and U+FEFF, the byte-order mark. This one is Unicode-aware. |
Anchors match a position, not a character. They consume nothing, which is why two of them in a row is not a contradiction.
| Construct | What it does |
|---|---|
| ^ | Start of the input. With the m flag, also just after every line terminator. |
| $ | End of the input. With the m flag, also just before every line terminator. |
| \b | A word boundary: the position between a \w character and a non-\w one, including the start and end of the string. Since \w is ASCII-only, so is \b./\bcat\b/ matches "cat" in "the cat sat" but not in "concatenate" |
| \B | Any position that is not a word boundary. |
| \A \z \Z | Do not exist in JavaScript. Trying to use them gives you a literal A, z or Z, silently, with no error.vs PCRE — In PCRE and Python these anchor to the absolute start and end regardless of the m flag. In JavaScript, ^ and $ without m already do that. |
| \G | Does not exist in JavaScript. The y (sticky) flag covers the same need. |
A quantifier repeats the thing before it. Greedy by default: it takes as much as it can and gives characters back one at a time until the rest of the pattern fits. That giving-back is called backtracking, and it is the source of both the surprises and the hangs.
| Construct | What it does |
|---|---|
| * | Zero or more. |
| + | One or more. |
| ? | Zero or one — optional. |
| {n} | Exactly n times. |
| {n,} | n or more times. |
| {n,m} | Between n and m times, inclusive. |
| *? +? ?? {n,m}? | The lazy forms. They take as little as possible and grow only when forced, which is the opposite starting point rather than a different set of matches./<.+>/ on "<b>hi</b>" matches the whole string; /<.+?>/ matches just "<b>" |
| *+ ++ (?>…) | Possessive quantifiers and atomic groups do NOT exist in JavaScript. There is no syntax for “match this and never give it back”.vs PCRE — PCRE, Java and Ruby all have them, and they are the standard cure for catastrophic backtracking. In JavaScript you get the same effect by making the alternatives non-overlapping — [^"]* instead of .*? inside quotes — so the engine has nothing to backtrack into. |
Parentheses do two jobs at once — they bound a sub-pattern, and they capture what it matched. When you only want the first job, say so: a non-capturing group is faster and keeps your group numbers meaning something.
| Construct | What it does |
|---|---|
| (x) | Capturing group. Numbered left to right by opening parenthesis, starting at 1. |
| (?:x) | Non-capturing group: grouping only, nothing stored. |
| (?<name>x) | Named capturing group. Read it back as match.groups.name./(?<y>\d{4})-(?<m>\d{2})/ → groups.y and groups.m |
| a|b | Alternation. It has the lowest precedence of anything, so ^a|b$ means “starts with a” OR “ends with b” — write ^(?:a|b)$ if you meant one of two whole strings. |
| \1 \2 | Backreference: matches the same text group 1 or 2 already matched./\b(\w+)\s+\1\b/ finds a doubled word such as "the the" |
| \k<name> | Backreference to a named group. |
| $1 $<name> $& $` $' $$ | Replacement-string tokens, not pattern syntax: group 1, a named group, the whole match, everything before it, everything after it, and a literal dollar sign. |
Lookaround asserts that something is or is not next to the current position without consuming it. It is how you match “a number followed by kg” and get back just the number.
| Construct | What it does |
|---|---|
| (?=x) | Positive lookahead: what follows must match x, and x is not part of the result./\d+(?=kg)/ on "60kg" matches "60" |
| (?!x) | Negative lookahead: what follows must not match x. |
| (?<=x) | Positive lookbehind: what precedes must match x./(?<=RM)\d+/ on "RM250" matches "250" |
| (?<!x) | Negative lookbehind: what precedes must not match x. |
| Variable-length lookbehind | JavaScript allows a lookbehind of any length, including quantifiers: (?<=\w+\s) is legal.vs PCRE — PCRE requires a fixed-length lookbehind and rejects that pattern outright. Java allows a bounded length. This is the one place where a JavaScript pattern is more permissive than a PCRE one, so patterns copied in this direction fail and patterns copied out do not. |
Flags go after the closing slash of a literal, or as the second argument to new RegExp. Two of them change what your pattern matches; one of them changes what happens between calls, and that one causes more confusion than the rest combined.
| Construct | What it does |
|---|---|
| g | Global: find every match rather than stopping at the first. It also makes the regex object stateful — it keeps a lastIndex between calls, so calling .test() twice on the same regex object can give different answers. |
| i | Case-insensitive. |
| m | Multiline: ^ and $ also match at line breaks. It does not affect what . matches. |
| s | dotAll: . also matches line terminators. |
| u | Unicode mode: the pattern is read as code points rather than UTF-16 units, \u{…} and \p{…} become available, and unknown escapes become syntax errors instead of silently matching a literal. |
| y | Sticky: the match must begin exactly at lastIndex. Used by tokenisers, where “somewhere later in the string” is the wrong answer. |
| d | hasIndices: each match gains an .indices array giving the start and end of every group. |
| v | unicodeSets: everything u does, plus set operations inside classes, \q{…} string literals, and properties of strings such as \p{RGI_Emoji}. It cannot be combined with u — pick one. |
These are how you say “any letter” and mean it in every writing system. They need the u or v flag; without one, \p{L} is just the letters p, L and some braces.
| Construct | What it does |
|---|---|
| \p{L} \P{L} | Any letter, and its negation. \p{Lu} and \p{Ll} narrow it to uppercase and lowercase. |
| \p{N} \p{Nd} | Any number, and decimal digits specifically — this is the Unicode-aware \d that JavaScript does not give you. |
| \p{P} \p{Zs} | Punctuation, and space separators. |
| \p{Script=Han} | Characters whose script is Han. Also written \p{sc=Han}./\p{Script=Han}+/u matches 中文 |
| \p{Script_Extensions=Greek} | Wider than Script=: it includes characters used with Greek even where their own script is Common. Abbreviated \p{scx=Greek}. Usually the one you actually want. |
| \p{Alphabetic} \p{White_Space} | Binary properties, written by name. |
| \p{RGI_Emoji} | A property of strings — it matches whole emoji sequences, including flags and family emoji built from several code points. Requires the v flag; it is a syntax error under u. |
How it works
- 1
Write the pattern without the slashes
The pattern box holds what goes between the slashes; the flags are toggles beside it. That means a forward slash needs no escaping here even though it would inside a literal — one of the few differences between a regex literal and new RegExp, and worth knowing since half the patterns you copy come from one form and half from the other.
- 2
Watch the error instead of an empty box
A pattern is invalid for most of the time you spend typing it — “[a-” is not a mistake, it is a pattern halfway written. So an incomplete or malformed pattern shows the engine’s own message as a sentence, naming the position and the offending construct, and the rest of the panel keeps working. It never throws and never blanks the page.
- 3
Read the matches and the groups
Every match is listed with its index, its text, and each capture group — numbered as $1, $2 and named ones as $<name>. A group that participated but matched nothing shows as empty; a group that never participated shows as undefined, and those are genuinely different results that a lot of code conflates.
- 4
Use the replace preview to check $1 and $&
The replacement field accepts the real tokens: $1 and $<name> for groups, $& for the whole match, $` and $' for the text before and after it, and $$ for a literal dollar. The preview honours your flags rather than forcing global, so you can see for yourself that without g only the first match is replaced.
- 5
Let a runaway pattern fail as a message
Your pattern runs inside a Web Worker, not on the page’s own thread, and the page holds a 1.5-second timer over it. If the worker has not answered by then it is terminated outright and you get a sentence naming catastrophic backtracking as the likely cause. That is the only reliable way to stop a running regular expression in a browser: once .exec() starts, no timer fires and no button can be clicked on the thread running it, so an input-length cap delays the freeze rather than preventing it.
Frequently asked questions
Which flavour is this, and how much does that matter?
JavaScript, meaning the RegExp built into every browser and into Node. It matters more than most people expect. \d is exactly [0-9] here and stays that way even under the u flag, while PCRE with (*UCP), .NET and Python 3 all make it match Unicode digits — so a pattern that validated numbers correctly in Python quietly accepts ٣٤٥ nowhere and rejects it here. \A, \z and \Z do not exist in JavaScript and silently match the literal letters instead of erroring. Atomic groups and possessive quantifiers do not exist either, which removes the standard cure for catastrophic backtracking. And lookbehind is more permissive here than in PCRE, not less — see below.
Why is .* usually the wrong tool?
Because it matches too much and then backtracks to fit, and both halves of that cause trouble. Run /<.+>/ against “<b>hi</b>” and you match the entire string, not the tag: the dot took everything and gave back only enough to find a final “>”. The lazy form /<.+?>/ gives you “<b>”, which is usually what you wanted, but it still walks forward one character at a time. The better pattern is /<[^>]+>/ — a negated class that physically cannot cross a “>”, so the engine finds the boundary without backtracking at all. As a rule: when there is a delimiter, say what the content may not contain rather than saying “anything” and hoping the quantifier sorts it out.
What is catastrophic backtracking and how do I avoid it?
It is what happens when a pattern can split the same text between two quantifiers in exponentially many ways. The classic is /(a+)+b/ against a run of “a” with no “b”: the engine must try every way of partitioning those characters between the inner and outer plus before it can conclude there is no match, and each extra “a” doubles the work. Thirty of them can take seconds; forty can hang a tab. The shapes to watch for are nested quantifiers — (x+)+, (x*)* — and alternations whose branches can match the same text, like (a|a)+. The cure in other languages is an atomic group; JavaScript does not have one, so you rewrite until the alternatives cannot overlap: [^"]* instead of .*? between quotes, or a single character class instead of a nested group. Test on a long string that does NOT match, because a pattern that matches usually returns fast and hides the problem.
Why can’t the browser just cancel a slow regular expression?
Because a regex runs to completion inside a single function call, and JavaScript on one thread does nothing else while a function is running. There is no yield point inside .exec(), so a timer you set beforehand does not fire, a click is not delivered, and code that would check a limit never gets a turn — which is exactly why a cap on the number of matches, or on the length of the input, cannot save you: your own guard is queued behind the thing it was meant to guard. The only real escape is to run the pattern on a different thread and destroy that thread. This page does that: your pattern goes to a Web Worker, the page keeps a 1.5-second timer, and on expiry it calls terminate() on the worker and starts a fresh one. Terminating is not a request the worker can decline, which is what makes it reliable. If you ever need this in your own code — validating a regex a user supplied, for instance — a worker with a timeout is the pattern, and there is no in-thread substitute for it.
Why does my regex return a different result the second time I run it?
Because it has the g flag and you reused the object. A RegExp with g or y keeps a lastIndex property, and .test() and .exec() both start from it and update it. So calling re.test(s) twice on the same string gives true and then false, which looks like the string changed. The fixes: build a fresh RegExp for each run, or use a method that does not carry state — String.prototype.match with g, or matchAll, which is what you want anyway when you need every match with its groups. This page builds a new RegExp on every run for exactly that reason.
What is the correct regex for validating an email address?
There isn’t one, and this is the most useful thing on the page. Every short “email regex” you find is wrong in both directions: it rejects addresses that are legal — quoted local parts, plus-addressing, new top-level domains, an intranet host with no dot at all — and accepts ones that will bounce, because syntax cannot tell you whether a mailbox exists. RFC 5322’s grammar is genuinely too complex to be worth writing out, and matching it exactly still would not tell you the address works. The one pattern with an authority behind it is the HTML specification’s definition for <input type="email">, which the spec itself calls a “willful violation” of RFC 5322 precisely because it chose to be practical rather than complete — so use type="email" and let the browser apply it. In your own code, check for exactly one @ with something on both sides, and then send a confirmation message. Delivery is the only validation there is.
Do lookbehind and named groups work everywhere?
Named groups, (?<name>…), have been in every current engine for years and are safe. Lookbehind, (?<=…) and (?<!…), arrived in the same version of the language but Safari shipped it years after Chrome and Firefox, so it is the one construct worth checking if you still support genuinely old iOS. Worth knowing: JavaScript’s lookbehind is variable-length, so (?<=\w+\s) is legal here and rejected outright by PCRE, which requires a fixed width. That makes this the rare case where a pattern copied out of JavaScript may fail elsewhere rather than the other way round.
When do I need the u flag, and what is v?
You need u whenever the pattern contains \p{…} or \u{…}, and whenever the text may contain characters outside the Basic Multilingual Plane — emoji, rarer CJK — because without u the engine works in UTF-16 code units and a dot can match half of an emoji. The v flag, unicodeSets, is a superset: everything u does, plus set operations inside classes ([\p{L}--[aeiou]] for difference, && for intersection), \q{…} for multi-character strings inside a class, and properties of strings such as \p{RGI_Emoji} that match whole emoji sequences. u and v cannot both be set — the engine throws — so treat v as the newer replacement rather than an addition.
What is the difference between greedy and lazy, on one example?
Take the string <b>bold</b> and <b>more</b>. The greedy /<.+>/ returns one match: the whole string, because the dot took everything and then gave back just enough to find a closing angle bracket at the very end. The lazy /<.+?>/ returns four matches — <b>, </b>, <b>, </b> — because it takes as little as possible and stops at the first “>” each time. The non-backtracking /<[^>]+>/ returns the same four, faster, because the class cannot cross a “>” to begin with. Paste all three into the tester above; the difference is easier to see than to describe.
Is my pattern or my test string sent anywhere?
No. Everything runs in the JavaScript engine already in your browser, in this tab. There is no request of any kind, and nothing is written to localStorage — which matters here more than on most pages, because the strings people paste into a regex tester are log lines, customer records and API responses.