Regex Cheat Sheet: Browser Syntax, Flags, and the Empty-Match Trap

Every regex cheat sheet lists the syntax. Almost none mention the pattern class that

freezes a browser tab. This reference covers the syntax that runs in every modern

engine, plus the failure modes we hit while building our own tester, whose defaults are

cited throughout so you can reproduce every example live.

The defaults you can reproduce in one minute

Our tester opens with the pattern matching one or more digits, the global flag on, and

this sample text: WebRecast 2024, sürüm 1.0.3, güncelleme 42. Running it returns five

matches, 2024, 1, 0, 3, and 42, and the counter reads 5 matches found. If you paste the

same text into any other tester and get a different count, the other tool changed the

flags, and the flags are the first thing to check, never the pattern.

Character classes


.        any character except a newline
[abc]    any one of a, b, c
[^abc]   any character except a, b, c
[a-z]    a range, lowercase letters
\d       a digit
\D       a non-digit
\w       a word character, letters, digits, underscore
\W       a non-word character
\s       whitespace, space, tab, newline
\S       a non-whitespace character

Escape a special character to match it literally: backslash dot for a period, backslash

star for an asterisk, backslash backslash for the backslash itself. One class expression

matches exactly one character, so [abc] matches the a in cat and nothing more.

Quantifiers


*        zero or more
+        one or more
?        zero or one
{3}      exactly three
{2,4}    between two and four
{3,}     three or more

Quantifiers are greedy by default and backtrack to satisfy the rest of the pattern. Add

a question mark after a quantifier to make it lazy: the pattern matching any characters

up to a quote behaves differently when lazy, because it stops at the first quote instead

of the last.

The empty-match trap

A star or a question mark can match zero characters. In JavaScript, a global exec loop

that receives an empty match does not advance its position, because the match end equals

the match start. A loop without a guard then repeats the same empty match forever and

freezes the page.

Our own tester demonstrates the danger in its counting loop, which lacks the standard

guard. Type a pattern that can match nothing, such as digits with a star quantifier, and

the loop spins without advancing. The guard, when you write such a loop yourself, is one

line: if the match index equals the loop position, increment the position before the

next iteration. String replace and matchAll advance past empty matches for you, which is

why the highlighted preview in the tester survives patterns that stall the counter.

Anchors and boundaries


^        start of the string, or of a line with the m flag
$        end of the string, or of a line with the m flag
\b       a word boundary
\B       not a word boundary

An anchor consumes no characters. The pattern for a line containing only digits is a

start anchor, one or more digits, and an end anchor, and it fails on any line holding a

comma. Word boundaries make substring matches behave: digits bounded on both sides match

the 42 in güncelleme 42 and skip past the digits inside longer tokens.

Groups and lookaround


(abc)    a capturing group
(?:abc)  a non-capturing group
(a|b)    alternation, a or b
(?=x)    lookahead, followed by x
(?!x)    negative lookahead
(?<=x)   lookbehind, preceded by x
(?<!x)   negative lookbehind

Capturing groups also define replacements: the first group reappears in the replacement

string as dollar-one, the second as dollar-two. Lookaround asserts a condition without

consuming, so a lookahead for a digit after a word character splits text at every

letter-to-digit transition. State this cautiously: lookbehind arrived in browsers later

than the rest of the list, so very old Safari builds reject patterns containing it,

while server engines built on RE2 reject lookarounds entirely because they forgo

backtracking. Test on the engine that will run your pattern, not on a convenient one.

Flags

Our tester exposes four flag buttons, g for all matches, i for case insensitivity, m for

multiline anchors, and s to let the dot match newlines. Internally it strips any

character outside gimsuy before constructing the expression, so a pasted flag set

containing anything else cannot throw from the flags alone. The u flag enables full

Unicode handling and y anchors each match at the current position, both valid in the

sanitizer set even without dedicated buttons.

Patterns we keep reusing


^\d+$                  a line of digits only
[\w.+-]+@[\w-]+\.[\w.]+  a permissive email shape for triage, not validation
https?://\S+           a bare URL token
\s{2,}                 runs of whitespace to collapse
(\d+)-(\d+)-(\d+)      a date with capturable parts

The email shape deserves its warning. Any short expression accepts addresses it should

reject and rejects some it should accept. Real validation sends a confirmation message,

and the regex only catches obvious garbage first.

Checklist before you trust a pattern

  • You tested it on the exact engine that will run it.
  • It cannot match the empty string, or your loop has the zero-width guard.
  • Anchors match your line expectations, with the m flag when needed.
  • The g flag is on for counting, off for a single boolean check.
  • Error output is displayed, since the tester surfaces engine messages for invalid

syntax instead of failing silently.

  • You saved one canonical test string, like our sürüm 1.0.3 sample, as a regression

fixture.

Syntax is the easy half of regex. The failure modes above, empty matches, greedy

backtracking, engine differences, are what break production, and each one is reproducible

in a live tester in under a minute.

Which pattern froze or fooled you last? Paste it with your sample text into the

[regex tester](https://webrecast.com/en/regex-tester) and check it against the empty-match

trap before it reaches your code.