Converting between data formats is where assumptions die. The file looked
fine, the parser looked fine, the pipeline returned success, and the
output was wrong. Three bugs from one real migration, each a pattern that
will visit your CSV and JSON pipelines too.
Bug 1: line endings are data
The task: rewrite specific values in a 175,000-line data file, one line per
record. The script matched lines ending in a comma using `
in multilinemode. It matched nothing. The file used CRLF, and `
asserts before `\n`,not before `\r\n`.
The damage pattern is important: the regex was correct for LF files, the
substitution quietly succeeded on zero records, and the process exited
zero. Success-shaped failure.
The fix, in order of preference:
1. Normalize endings once at ingest: `content.replace(/\r\n/g, '\n')`.
2. Or make endings explicit in every pattern: `\r?
instead of `.3. Never trust a substitution count of zero on a file you did not inspect.
Validation that catches it: count matches before and after. We printed
zero and the pipeline still felt finished. Print the count and assert it
against an expectation.
Bug 2: the tool disagreed with itself
Filtering a URL list for rows containing a path segment, this pattern:
grep "/blog/" urls.txt
returned zero rows. The same filter with the slashes removed returned the
expected 190 rows. The tool, it turned out, was an alternative grep
implementation that treats a pattern wrapped in slashes as a delimited
regex with different semantics, not as a literal.
In a data pipeline this is worse than an error. Zero rows is a valid,
plausible answer for a filter. Nothing crashed. Downstream steps
processed an empty list and reported success.
The fixes:
1. Use fixed-string matching when you mean a literal: `grep -F`.
2. Or do the filtering in the same language as the rest of the pipeline.
Our final version extracted URLs with a proper regex in the script
itself, one tool, one semantics.
3. Treat an unexpectedly empty result set as an exception, not an answer.
Assert minimum counts at every stage.
Bug 3: quoting shifts everything after it
A generated config file held a list of 63 quoted strings. The header
comment above it contained an apostrophe, written naturally: "the blog's
value signal". A parser collecting quoted strings with a pairwise regex
matched that apostrophe against the opening quote of the first list entry.
Every subsequent match shifted by one boundary. The parser returned 63
entries and every one of them was garbage, including commas matched
between adjacent quotes.
The validation that saved us: a membership check. "Is this specific known
entry in the parsed result?" It was not. The count, 63, was coincidentally
correct and would have passed a length assertion.
The fixes:
1. Parse line-anchored: `^[ \t]*'([a-z0-9-]+)',
cannot match prose.2. Strip comments before parsing quoted content.
3. Validate with known members, not lengths. Counts lie when boundaries
shift. Named entries do not.
The general pattern: success is not validation
All three bugs shared a signature. The pipeline completed. Exit codes were
zero. The output existed and had plausible shape. Only the content was
wrong, in ways that counted as valid data to every later stage.
CSV and JSON conversion pipelines need three assertions that have nothing
to do with success:
1. Volume: the output record count matches an expected range derived from
the input. Zero is outside every sane range.
2. Membership: specific known records appear in the output. Pick three you
can quote from memory.
3. Round trip: convert the output back and compare against a normalized
form of the input. Field-level disagreement is a boundary bug.
Where these bugs live in your stack
- Line endings: any file crossing Windows and Unix systems, which is every
repository with mixed contributors.
- Tool semantics: any pipeline stitched from shell commands where grep,
sed, and awk implementations vary by platform.
- Quoting shifts: any parser over content that mixes comments with quoted
data, which includes most config formats and half of all CSV dialects.
Convert between formats with a real parser for the target format when one
exists. The regex route is for extraction, and extraction needs the
assertions above or it will one day hand you sixty-three pieces of
garbage with a success code.
---
What is the worst silent corruption a pipeline has handed you? If you are
converting between CSV and JSON right now, our converter runs entirely in
your browser, so the data never leaves your machine while you check the
output: https://webrecast.com/en/csv-json-converter