• mission
  • work
  • contact
  • blog
  • Fast CSV parsing

    Max Henger | 18th of January, 2026

    In many of our projects we are processing large amounts of data for our clients. Both on servers, and their and our machines. This data is often supplied in CSV format, where we have to process hundreds of CSV files measuring around 50 MiB in size. After profiling our applications we saw that we were spending a significant amount of time parsing CSVs, and decided to optimize our loading routine. We thought it would be interested to share our findings.

    The problem

    CSV files are a notoriously messy format, each CSV writer has their own quirks. In our case we're mostly dealing with files that adhere to RFC4180 . Summarizing the spec: rows are separated by newlines ( \n , or \r\n ), columns are separated by a configurable separator, and fields containing characters that might be parsed ambiguously are enclosed in double quotes.

    Our own requirements are:

    • We never sacrifice correctness for speed. Input must be checked to be UTF-8 and fields must be properly escaped with quotes.

    • Files may be arbitrarily large, and we do not want to waste memory, so we're going to parse the data from buffers.

    • On local machines data is present on disk, but on servers it may come in through sockets, so we don't want to rely on seek ing through the data.

    • The parser will be single threaded. On servers we want to leave processing power for other requests, and on local machines the processing is usually embarrassingly parallel .

    Decomposing the problem

    We use a combination of cargo flamegraph , cargo asm , and our own profiler (that uses perf_event for hardware statistics, and performs instrumented profiling) to investigate and optimize our code. Once we've identified a bottleneck then we write benchmarks and flip back and forth between our profiler and the assembly to iteratively improve the performance.

    To bring structure to our optimizations, and to gauge the upper speed limits, we decompose the CSV reader into the following consituent subproblems:

    • Reading from a stream into buffers (as we've already decided we do not want to load complete CSV files into memory).

    • Iterating through rows and columns of the CSV.

    • Composing everything into a CSV parser. Taking care to design our benchmarks to reflect real-world use.

    Optimizing

    Reading from a stream

    We wrote a set of benchmarks that reads a file of 50 MiB in chunks of powers of two from 1 KiB up to 1024 KiB . We'll have a read -based implementation for filesystem-cached and filesystem-cleared reading, and a mmap implementation (that performs munmap to explicitly release memory). The results that show the general trends are as following:

    filesystem chunk speed cached speed cleared speed mmap
    cached 4 KiB 4.98 GiB/s 2.69 GiB/s 0.70 GiB/s
    cached 8 KiB 7.48 GiB/s 2.75 GiB/s 1.45 GiB/s
    cached 16 KiB 9.92 GiB/s 2.90 GiB/s 2.35 GiB/s
    cached 32 KiB 12.0 GiB/s 2.55 GiB/s 2.97 GiB/s
    cached 64 KiB 13.2 GiB/s 2.91 GiB/s 3.06 GiB/s
    cached 128 KiB 3.44 GiB/s 2.24 GiB/s 3.10 GiB/s
    cached 256 KiB 3.52 GiB/s 2.30 GiB/s 3.23 GiB/s

    Note that the SSD we're testing on has a buffered speed limit of approximately 3.10 GiB/s . So we find that on servers and when reading from sockets a buffer size of 16 KiB is sufficient while keeping the memory footprint low. On local machines memmap can be used when reading files. Note the reduced speeds of the two read ing variants at 128 KiB . This is approximately half of the L1d cache on our workstation CPU. So faster buffers will ameliorate the overhead of syscall s, but too large and you'll be trashing your cache. This finding is repeatable on machines with different cache sizes.

    Iterating through rows and columns

    We know that our solution at least needs to be able to find CSV row/column separators. So to gauge performance limits we'll first limit ourselves to counting the number of separators in a file.

    An initial implementation might iterate through UTF-8 characters and increase a counter whenever a newline/delimiter is encountered. The code is intentionally kept "branchy".

    fn count_chars(data: &str) -> (usize, usize) {
        let mut num_newline = 0;
        let mut num_delimiter = 0;
    
        for c in data.chars() {
            if c == '\n' {
                num_newline += 1;
            } else if c == ',' {
                num_delimiter += 1;
            }
        }
    
        return (num_newline, num_delimiter);
    }

    But we know that we're only seeking ASCII characters. Due to the design of UTF-8 we can get away with checking that the input is a UTF-8 string, interpreting it as bytes, and then to find the newline and delimiter bytes, i.e.

    for &c in bytes {
        if c == b'\n' {
            num_newline += 1;
        } else if c == b',' {
            num_delimiter += 1;
        }
    }

    We then turn to a single instruction, multiple data (SIMD) implementation. Given that our client's machines will generally not have support for AVX512 , we'll use a (widely supported) AVX2 implementation.

    Targeting AVX2 (and the SSE2 , PCLMULQDQ CPU features) we have the following initial code. Looking ahead at our later implementation we will process 64 bytes (512 bits) here already, instead of the 32-byte AVX2 registers. Our data reader will ensure the 64-byte alignment of the beginning and end of the buffer.

    use std::arch::x86_64 as a;
    
    impl ChunkIter {
        pub(crate) fn next(&mut self) -> Option {
            // We'll ensure these properties elsewhere
            debug_assert!(self.data.as_ptr().align_offset(64) == 0);
            debug_assert!(self.data.len() % 64 == 0);
    
            if self.pos >= self.data.len() { return None; }
    
            // Compare the 64 characters and turn them into a bitmask
            fn cmpeq_mask(bytes_lo: a::__m256i, bytes_hi: a::__m256i, cmp_byte: u8) -> u64 {
                unsafe {
                    let cmp = a::_mm256_set1_epi8(cmp_byte.cast_signed());
    
                    let cmp_lo  = a::_mm256_cmpeq_epi8(bytes_lo, cmp);
                    let mask_lo = a::_mm256_movemask_epi8(cmp_lo).cast_unsigned();
    
                    let cmp_hi  = a::_mm256_cmpeq_epi8(bytes_hi, cmp);
                    let mask_hi = a::_mm256_movemask_epi8(cmp_hi).cast_unsigned();
    
                    return (mask_hi as u64) << 32 | (mask_lo as u64);
                }
            }
            
            unsafe {
                // Load bytes
                let mut ptr = self.data.as_ptr().add(self.pos);
                let bytes_lo = a::_mm256_load_si256(ptr as *const a::__m256i);
                
                ptr = ptr.add(32);
                let bytes_hi = a::_mm256_load_si256(ptr as *const a::__m256i);
    
                // Compute newline/delimiter bitsets
                let bits_newline   = cmpeq_mask(bytes_lo, bytes_hi, b'\n');
                let bits_delimiter = cmpeq_mask(bytes_lo, bytes_hi, self.delimiter);
                        
                // Update for next iteration
                let pos = self.pos;
                self.pos += 64;
                
                return Some(BlockIter{
                    base:         pos,
                    is_separator: bits_newline | bits_delimiter,
                    is_newline:   bits_newline,
                });
            }
        }
    }
    
    impl BlockIter {
        pub(crate) fn next(&mut self) -> Option<(usize, bool)> {
            if self.is_separator == 0 { return None };
    
            let bit_index = self.is_separator.trailing_zeros() as usize;
            let byte_index = self.base + bit_index;
            let is_newline = (self.is_newline >> bit_index) & 0x01;
            let is_newline = if is_newline == 1 { true } else { false };
    
            self.is_separator &= self.is_separator - 1;
    
            return Some((byte_index, is_newline));
        }
    }

    After making sure the assembly looks as expected (using cargo asm and #[inline(never)] where appropriate) we run our benchmarks. For comparison we also include a 32-byte (256-bit) aligned version:

    benchmark speed cycles insns cache misses cache access branch misses branch insns
    UTF-8 iter 2.05 GiB/s 111 M 485 M 13.5 k 1.53 M 180 k 193 M
    byte iter 2.36 GiB/s 96.4 M 271 M 13.6 k 1.53 M 221 k 103 M
    AVX2 256-bit 9.67 GiB/s 23.5 M 67.0 M 13.3 k 1.53 M 120 k 7.98 M
    AVX2 512-bit 10.0 GiB/s 22.7 M 65.1 M 13.6 k 1.53 M 87.8 k 6.47 M

    We see that the AVX2 versions are roughly 3.9 times faster than byte iteration, and 6.8 times faster than UTF-8 character iteration.

    These initial findings provide an upper boundary to our processing speed.

    CSV Parsing

    We'll extend our 512-bit AVX2 variant with actual CSV parsing. We'll have to account for quote pairs, misplaced quote marks, newlines that can be \n as well as \r\n . We'll show our code first, then explain the applied tricks.

    impl ChunkIter {
        pub(crate) fn next(&mut self, data: &[u8]) -> Option {
            // Additional checks that are upheld in the data reader
            debug_assert!(data.len() <= u32::MAX as usize);
            debug_assert!(data.as_ptr().align_offset(ALIGN) == 0);
            debug_assert!(data.len() % ALIGN == 0);
            debug_assert_eq!(64, ALIGN);
    
            if self.pos >= data.len() as u32 {
                return None;
            }
            
            unsafe {
                // Load CSV data in two 32-byte chunks.
                let mut ptr = data.as_ptr().add(self.pos as usize);
                let bytes_lo = a::_mm256_load_si256(ptr.cast::());
                ptr = ptr.add(32);
                let bytes_hi = a::_mm256_load_si256(ptr.cast::());
    
                // Identifying quotes, and turning this into a mask that is 1
                // when between a quote pair
                let bits_quote = cmpeq_mask(bytes_lo, bytes_hi, b'"');
                let mask_quote = a::_mm_cvtsi128_si64(a::_mm_clmulepi64_si128::<0>(
                    a::_mm_set_epi64x(0, bits_quote.cast_signed()),
                    a::_mm_set1_epi8(0xFFu8.cast_signed())
                ));
                let mask_quote = mask_quote.cast_unsigned() ^ self.last_quote_mask;
    
                // Detecting carriage feed, line feeds, delimiters
                let bits_feed        = cmpeq_mask(bytes_lo, bytes_hi, b'\r');
                let mut bits_newline = cmpeq_mask(bytes_lo, bytes_hi, b'\n');
                let mut bits_delim   = cmpeq_mask(bytes_lo, bytes_hi, self.delimiter);
    
                // Account for separators inside quotes
                bits_newline &= !mask_quote;
                bits_delim   &= !mask_quote;
    
                let bits_separator = bits_newline | bits_delim;
    
                // Detect '\r\n' patterns
                let bits_feed_newline = ((bits_feed << 1) | self.last_feed) & bits_newline;
    
                // Detect quoting errors
                let cur_full_mask = mask_quote | (mask_quote << 1) | self.last_quote_switch >> 63;
                let prev_full_mask = cur_full_mask << 1 | self.last_full_quote_mask;
    
                let leading_quote = !prev_full_mask & cur_full_mask;
                let bits_quote_err = leading_quote & !((bits_separator << 1) | self.last_bits_separator);
    
                // Prepare for next iteration
                let base_pos = self.pos;
    
                self.pos += ALIGN as u32;
                self.last_feed = bits_feed >> 63;
                self.last_bits_separator = bits_separator >> 63;
                self.last_quote_switch = (mask_quote.cast_signed() >> 63).cast_unsigned();
                self.last_full_quote_mask = cur_full_mask >> 63;
    
                return Some(BlockIter{
                    base_pos,
                    bits_feed_newline,
                    bits_newline,
                    bits_separator: bits_separator | bits_quote_err,
                    bits_quote_err,
                })
            }
        }
    }
    
    #[derive(PartialEq, Eq, Clone, Copy)]
    pub(crate) enum SeparatorKind {
        Delimiter = 0,
        Newline = 1,
        QuoteErr = 2,
    }
    
    impl BlockIter {
        pub(crate) fn next(&mut self) -> Option<(u32, u32, SeparatorKind)> {
            if self.bits_separator == 0 { return None };
    
            // Extract next separator
            let bit_index = self.bits_separator.trailing_zeros() as u32;
            let separator_index = self.base_pos + bit_index;
    
            // Account for the larger length of '\r\n'
            let feed_offset = ((self.bits_feed_newline >> bit_index) & 0x0000_0001) as u32;
    
            let kind = match (
                ((self.bits_quote_err >> bit_index) & 0x01),
                ((self.bits_newline >> bit_index) & 0x01),
            ) {
                (0, 0) => SeparatorKind::Delimiter,
                (0, 1) => SeparatorKind::Newline,
                (1, _) => SeparatorKind::QuoteErr,
                _ => unreachable!(),
            };
    
            // Returning bounding indices into buffer, and the kind of separator
            self.bits_separator &= self.bits_separator - 1; // clear lowest set bit
            return Some((separator_index - feed_offset, separator_index + 1, kind));
        }
    }

    Going through the design:

    • Accounting for quote marks uses a trick described in the blog of Geoff Langdale in use by (among others) simdjson . We first detect quote marks as usual using cmpeq_mask , but then apply a running XOR to the bytes using a carryless product (i.e. XOR multiplication). As an example in ASCII art:

      input:      this is "an input string", where "" chars are "between quotes"
      bits_quote: 00000000100000000000000010000000011000000000001000000000000001
      mask_quote: 00000000111111111111111100000000010000000000001111111111111110

      So if we mask our separators using !mask_quote , then we hide the separators that are actually between quotes.

      This is the reason we're processing in 64-byte chunks: it is significantly faster to do the _mm_clmulepi64_si128 and bit twiddling once per 64-byte chunk than every 32 bytes.

    • Accounting for carriage feeds (as in \r\n ) is done by shifting a mask for carriage feed up by 1, and then AND ing with the mask for newlines. This way we can account for \r using the bits_feed_newline bitset in BlockIter . BlockIter returns the last column's (exclusive) end index, and the next column's (inclusive) starting index. If we find a \r\n , then we'll subtract 1 from last column's end index.

    • Detecting quote errors requires some extra explanation. We wanted to optimize for the most common kind of data that we encounter in the CSVs we receive. Those are unquoted columns containing UTF-8 (to be more specific: ASCII) data.

      If a field is quoted, then we must decode that field. So we'll always have to check for a leading quote. But we prefer not to look through the entire field to check if there are any disallowed quotes.

      So we apply some bit twiddling using our masks to detect leading quotes ( !prev_full_mask & cur_full_mask ), and then detect the ones that are not placed after delimiters ( & !((bits_separator << 1) | self.last_bits_separator) ). And emit them from the BlockIter as errors. Once we hit such an error we immediately return from our processing loop.

      This way: if the leading character in a field is not a quote, then we can be certain that the field doesn't contain any quotes at all.

    • To make sure our data is always aligned and the size is a multiple of ALIGN we apply two additional tricks.

      Firstly we make sure that we're always reading to an aligned address. This may be done by:

      let expanded_size = requested_size + ALIGN - 1;
      let mut buffer = Vec::new();
      buffer.resize(expanded_size, 0u8);
      
      let begin = buffer.as_ptr().align_offset(ALIGN);
      let aligned_buffer = &mut buffer[begin..begin + requested_size];

      But our ChunkIter additionally requires the length to be a multiple of ALIGN . At the end of the file we might not be reading an exact multiple of ALIGN bytes.

      So we apply three tricks to achieve this: we write our parser to not respond to the end of the file, only to actual newlines. This way we're free to append bytes to the read buffer as long as they don't influencing the parser. Secondly we ensure that our buffer always has room for two more bytes. And lastly: if we end the file without a newline, we append a \r\n sequence, and fill the buffer up with NULL bytes up to the next ALIGN boundary.

      Combining all these tricks we properly read CSVs whether they end with a newline or not, and we ensure that our AVX2 iterator can always read the next ALIGN number of bytes.

    • Finally, the match expression producing the SeparatorKind might stand out as being "unoptimized". As expected, and confirmed by checking the assembly, LLVM is capable of figuring out the unreachable!() can be removed. Once the BlockIter::next() call is inlined in the CSV parsing code then LLVM is capable of producing more optimized assembly by removing the transformation into the SeparatorKind enum.

    Putting it all together

    Our final implementation performs UTF-8 validation, invalid quote detection, quoted field parsing, and delimiter detection. Choosing between AVX2 or the regular fallback implementation is done at runtime. The following benchmarks are reading actual CSV files using a high-level read_csv() function. Only the overhead of reading from disk is removed.

    To make sure we have realistic benchmarks that can be interpreted for their effects on real-world code (we found that several libraries advertising very high read speeds do not reach those speeds in practical use-cases), we have the following benchmark variations:

    • Using the AVX2 versus the regular non-SIMD iterator implemenation. This indicates the practical speedup we achieved.

    • Reading raw columns versus checking for a leading quote mark. Note that our test data, like our real-world data, rarely contains quoted fields. Hence this models the cost of checking for the leading quote mark.

    • Explicitly read ing from a buffer in chunks versus loading and prealigning the entire file. This models the memcpy overhead.

    • As pseudo workload on the results: computing the string length of the column versus summing the bytes of the column strings. The former will not touch the memory of the file's contents anymore, as the length is precomputed in the &str slice. Hence this estimates the raw speed of our parsing method. The latter will touch the memory of the file's contents, hence reflect the speed of real world usage.

    The benchmarks are ran on a AMD Zen3 5900x, using a buffer of 16 KiB in the buffering benchmarks. The actual speed in GiB/s is hardware dependent, but the speedup is consistent across multiple different machines. Displaying the results of the AVX2 and the fallback code side by side produces:

    read method column access workload AVX2 speed fallback speed speedup
    chunked read raw access string length 3.18 GiB/s 1.08 GiB/s 2.94x
    chunked read raw access byte sum 1.92 GiB/s 0.90 GiB/s 2.13x
    chunked read decoding quotes string length 2.84 GiB/s 1.10 GiB/s 2.58x
    chunked read decoding quotes byte sum 1.78 GiB/s 0.93 GiB/s 1.91x
    prealigned in-memory raw access string length 3.45 GiB/s 1.25 GiB/s 2.76x
    prealigned in-memory raw access byte sum 2.01 GiB/s 0.97 GiB/s 2.07x
    prealigned in-memory decoding quotes string length 2.92 GiB/s 1.17 GiB/s 2.50x
    prealigned in-memory decoding quotes byte sum 1.91 GiB/s 0.97 GiB/s 1.97x

    Hence spending a day or two optimizing CSV parsing resulted in a parser that is 2.5 to 3.0 times faster, and will result in our clients being able to compute their results from real-world data 2.0 times faster.

    CONTACT

    We are Mario Verhagen and Max Henger, two engineers with over 25 years of experience in aerospace engineering and software development. Celus was founded by our drive to fundamentally enhance the value of your software.

    Ready to unlock your competitive advantage? Contact Celus for a free technical brainstorm on building secure, high-performance applications and simulation software that transforms your workflows and existing tools into intuitive and performant applications.

    ADDRESS

    Barnsteenhorst 180
    2592 EN Den Haag
    The Netherlands

    E-MAIL

    email