A segfault in a Node script feels like a broken runtime. Ours was not. A
single regex, executed hundreds of times against a large file, overflowed
the regular expression engine's stack and took the whole process down with
exit code 139. The same code ran fine on smaller inputs for months.
If a Node build step dies with a segmentation fault and no stack trace,
suspect your regexes before your binaries. Here is the full diagnosis so you
can skip the detours we took.
The failure signature
Three observations, in order of how misleading they were:
1. No stack trace. The process printed nothing and died with 139, the shell
code for SIGSEGV.
2. Different crash points each run. Page 120 in one run, 1,934 in another.
Nondeterminism pointed at memory corruption or a runtime bug.
3. Version independent. We reproduced it on Node 20 LTS and Node 25.
The nondeterminism was the clue that mattered. The script iterated pages in
the same order every run, so a data-dependent bug should crash at the same
place. A crash that moves between runs under identical input means the
failure depends on accumulated state or stack depth, not on the input value.
The actual bug
The renderer built an HTML shell for every page. For blog pages, it needed
the article body, which lives in a 1.4 MB TypeScript file. The code did
this per page:
const src = readFileSync(blogDataPath, 'utf8');
const m = src.match(new RegExp(`slug: <Q>SLUG'[\\s\\S]*?content: \`([\\s\\S]*?)\``));
Two problems, one visible and one silent.
The visible problem: this ran for every article, in every language, hundreds
of times per build. Each call re-read the file and re-scanned it from the
slug position to the content field with a lazy quantifier. V8's Irregexp
engine recurses on backtracking patterns like this, and the recursion
consumes the native stack. Enough repeated deep scans tip it over. The
overflow is not a caught JavaScript exception. It corrupts the process, and
you get SIGSEGV.
The silent problem: the lazy match to the first backtick ended the article
body at the first escaped inline-code backtick in the source. Most of our
technical articles were rendering 13 to 39 words instead of 700 to 1,000.
No error, no warning. The build passed.
How we confirmed it
Two steps, both cheap:
1. Raise the V8 stack limit with `--stack-size=4000` and run again. The
crash moved from page 120 to page 1,934. A stack-dependent failure
shifting with the limit is regex backtracking, not corrupt memory.
2. Instrument the loop with a counter printed per page. The crash always
landed immediately after a specific class of page: blog articles with
long bodies and inline code.
The fix
Stop scanning the file repeatedly. Parse once, then serve from a map:
1. Split the source on slug boundaries, one chunk per article.
2. Within each small chunk, extract the content template literal with a
pattern that matches through escape sequences:
`content: \`((?:\\[\s\S]|[^`\\])*)\``
3. Store slug-to-content in a Map built once per run.
4. Unescape the extracted string: backtick, dollar, backslash.
The bracket-balance lesson generalizes. Any parser that matches to the first
closing bracket will break on nested structures, and it will break
silently, because a truncated match is still a match. Our FAQ parser had
the same disease: `faq: \[([\s\S]*?)\]` stopped at the first inner bracket,
so every FAQ list rendered empty and nobody noticed for months. Counting
bracket depth with a simple loop does not overflow and does not truncate.
Checklist for regex-driven build scripts
- Never run a lazy `\[\s\S\]*?` scan over a whole large file inside a loop.
Parse once, cache the result.
- Treat nested delimiters as a parser problem, not a regex problem. Use
bracket counting or a real tokenizer.
- Run build scripts against the largest realistic input before shipping
them, not the current smallest one.
- Watch exit code 139 with no stack trace as a regex stack overflow
signature, especially under `--stack-size` sensitivity.
- Assert on parsed output counts. Our silent FAQ truncation would have been
caught by one line: fail the build if an article known to have a FAQ
section renders without one.
After the fix, the renderer processes 20,930 routes in a single run with
exit code 0, on both Node versions, and the articles render their full
bodies.
---
Have you chased a segfault that turned out to be a regex? Describe the
pattern that caused it. If you test regular expressions against real input
before shipping them, our regex tester runs entirely in your browser, so
the patterns you paste never leave your machine:
https://webrecast.com/en/regex-tester