Handling Complex CSV Files with Multiline Breaks and Escaped Quotes
Comma-Separated Values (CSV) appears deceivingly simple until real-world enterprise exports encounter embedded commas, CRLF line terminators inside text cells, and escaped quotation marks. Understanding RFC 4180 compliance ensures your data pipeline remains robust.
1. The Grammar of Escaped Quotes (RFC 4180 §2.7)
According to RFC 4180, if fields are not enclosed with double quotes, commas and line breaks cannot exist inside the field values. However, when a field contains commas, line breaks, or quotation marks, the entire field MUST be enclosed in double quotes.
If double-quotes are used to enclose fields, then a double-quote appearing inside a field must be escaped by preceding it with another double quote (""). Naive splitting by comma inevitably breaks these records.
"ID","Description","Quantity"
101,"Widget, Type ""A"" (High precision)",42
102,"Multi-line
comment with
line breaks",152. State Machine Parsing in Client Memory
Instead of relying on simplistic RegExp expressions, robust parsers utilize a deterministic finite automaton (DFA) with four distinct states: FieldStart, InsideQuotedField, AfterQuoteInQuotedField, and InsideUnquotedField.
By processing characters sequentially through an ArrayBuffer or streaming chunks, the parser correctly identifies when a newline belongs to cell content rather than indicating a row termination.
3. Safe Conversion to JSON and Schema Invariants
When converting CSV rows into structured JSON, numbers with leading zeros (such as ZIP codes or account identifiers like '00421') must preserve string representations unless explicit numeric coercion is requested.
Always enforce UTF-8 normalization and handle BOM (Byte Order Mark, \uFEFF) gracefully at the stream boundary to prevent phantom column keys.
Summary & Best Practices
Robust CSV processing requires full state-machine lexing to guarantee that quoted carriage returns and nested quotation marks do not corrupt data downstream.
