Three Data Pipeline Bugs That Only Bit During Conversion

Converting between data formats is where assumptions die. The file looked

fine, the parser looked fine, the pipeline returned success, and the

output was wrong. Three bugs from one real migration, each a pattern that

will visit your CSV and JSON pipelines too.

Bug 1: line endings are data

The task: rewrite specific values in a 175,000-line data file, one line per

record. The script matched lines ending in a comma using `

in multiline

mode. It matched nothing. The file used CRLF, and `

asserts before `\n`,

not before `\r\n`.

The damage pattern is important: the regex was correct for LF files, the

substitution quietly succeeded on zero records, and the process exited

zero. Success-shaped failure.

The fix, in order of preference:

1. Normalize endings once at ingest: `content.replace(/\r\n/g, '\n')`.

2. Or make endings explicit in every pattern: `\r?

instead of `.

3. Never trust a substitution count of zero on a file you did not inspect.

Validation that catches it: count matches before and after. We printed

zero and the pipeline still felt finished. Print the count and assert it

against an expectation.

Bug 2: the tool disagreed with itself

Filtering a URL list for rows containing a path segment, this pattern:

grep "/blog/" urls.txt

returned zero rows. The same filter with the slashes removed returned the

expected 190 rows. The tool, it turned out, was an alternative grep

implementation that treats a pattern wrapped in slashes as a delimited

regex with different semantics, not as a literal.

In a data pipeline this is worse than an error. Zero rows is a valid,

plausible answer for a filter. Nothing crashed. Downstream steps

processed an empty list and reported success.

The fixes:

1. Use fixed-string matching when you mean a literal: `grep -F`.

2. Or do the filtering in the same language as the rest of the pipeline.

Our final version extracted URLs with a proper regex in the script

itself, one tool, one semantics.

3. Treat an unexpectedly empty result set as an exception, not an answer.

Assert minimum counts at every stage.

Bug 3: quoting shifts everything after it

A generated config file held a list of 63 quoted strings. The header

comment above it contained an apostrophe, written naturally: "the blog's

value signal". A parser collecting quoted strings with a pairwise regex

matched that apostrophe against the opening quote of the first list entry.

Every subsequent match shifted by one boundary. The parser returned 63

entries and every one of them was garbage, including commas matched

between adjacent quotes.

The validation that saved us: a membership check. "Is this specific known

entry in the parsed result?" It was not. The count, 63, was coincidentally

correct and would have passed a length assertion.

The fixes:

1. Parse line-anchored: `^[ \t]*'([a-z0-9-]+)',

cannot match prose.

2. Strip comments before parsing quoted content.

3. Validate with known members, not lengths. Counts lie when boundaries

shift. Named entries do not.

The general pattern: success is not validation

All three bugs shared a signature. The pipeline completed. Exit codes were

zero. The output existed and had plausible shape. Only the content was

wrong, in ways that counted as valid data to every later stage.

CSV and JSON conversion pipelines need three assertions that have nothing

to do with success:

1. Volume: the output record count matches an expected range derived from

the input. Zero is outside every sane range.

2. Membership: specific known records appear in the output. Pick three you

can quote from memory.

3. Round trip: convert the output back and compare against a normalized

form of the input. Field-level disagreement is a boundary bug.

Where these bugs live in your stack

  • Line endings: any file crossing Windows and Unix systems, which is every

repository with mixed contributors.

  • Tool semantics: any pipeline stitched from shell commands where grep,

sed, and awk implementations vary by platform.

  • Quoting shifts: any parser over content that mixes comments with quoted

data, which includes most config formats and half of all CSV dialects.

Convert between formats with a real parser for the target format when one

exists. The regex route is for extraction, and extraction needs the

assertions above or it will one day hand you sixty-three pieces of

garbage with a success code.

---

What is the worst silent corruption a pipeline has handed you? If you are

converting between CSV and JSON right now, our converter runs entirely in

your browser, so the data never leaves your machine while you check the

output: https://webrecast.com/en/csv-json-converter