cargo / base64 / audit
cargo : base64 @ 0.22.1
PE Patrick Elsen signed 2026-05-27 published 2026-05-27

Claims

algorithm-impl-boundsalgorithm-impl-correctalgorithm-impl-safealgorithm-impl-testedhas-binarieshas-build-exechas-fuzz-testshas-install-exechas-integration-testshas-property-testshas-unit-testsimpl-algorithmimpl-concurrencyimpl-cryptoimpl-datastructureimpl-interpreterimpl-jitimpl-parserimpl-protocolis-benignparser-impl-correctparser-impl-safeparser-impl-testeduses-concurrencyuses-cryptouses-environmentuses-execuses-filesystemuses-interpreteruses-jituses-networkuses-unsafe

Summary

base64 0.22.1 is the de-facto RFC 4648 codec for Rust: #![forbid(unsafe_code)], no runtime deps, table-driven parser with checked-arithmetic length computations and cross-validated tests plus four cargo-fuzz harnesses. No findings; safe to deploy.

Report

Subject

base64 is a no-std-friendly Rust library for encoding and decoding the base64 transfer encoding (RFC 4648). It exposes an Engine trait and a table-driven GeneralPurpose implementation parameterised by Alphabet and GeneralPurposeConfig (padding behaviour, trailing-bits tolerance, padding mode). The crate ships in-memory APIs (encode/decode to String/Vec, slice-in/slice-out variants), streaming wrappers (read::DecoderReader, write::EncoderWriter, EncoderStringWriter), and a display::Base64Display fmt::Display adapter. Six pre-defined alphabets (STANDARD, URL_SAFE, CRYPT, BCRYPT, IMAP_MUTF7, BIN_HEX) are provided as constants.

Methodology

The published crate contents were compared against the upstream Git repository at the commit recorded in .cargo_vcs_info.json using diff. All 21 source files in src/ (~6500 lines including the templated test harness) were read in full, along with the integration tests in tests/, the example in examples/, and the fuzzer entry points kept in the VCS-only fuzz/ directory. Manifest metadata, the clippy.toml, the .gitignore, and the crate-level deny/forbid attributes were inspected to bound the crate's runtime surface.

The review focused on parser/encoder soundness (panic-freedom on adversarial input, correct error reporting, no out-of-bounds access), adherence to RFC 4648, malleability handling (non-canonical padding and trailing bits), and the streaming decoder's behaviour at arbitrary buffer boundaries.

Results

The diff between the published crate and the upstream commit shows only cargo's standard Cargo.toml normalisation, the auto-generated .cargo_vcs_info.json, and a Cargo.lock cargo adds on publish; the upstream fuzz/ workspace is excluded from the published crate, which is normal. No source files differ.

The crate sets #![forbid(unsafe_code)] at the crate root, which rules out unsafe blocks, FFI, and raw-pointer code throughout the library and supports uses-unsafe. The manifest declares no build.rs, no [lib] proc-macro = true, no [build-dependencies], and no runtime dependencies, justifying has-build-exec and has-install-exec, and has-binaries (the crate ships only Rust source, two licence files, the SVG sponsor icon, and the standard markdown documentation). The codebase performs no network, file, process, or environment access (justifying uses-network, uses-filesystem, uses-exec, uses-environment), spawns no threads and uses no async runtime (justifying uses-concurrency); there is no dynamic code execution, JIT, or embedded interpreter (justifying uses-jit and uses-interpreter), and does not implement or call into any cryptographic primitive (justifying uses-crypto and impl-crypto). The GeneralPurpose engine documents that it is not constant-time and steers cryptographic-key use cases to a forthcoming constant-time engine.

The decoder is a table-driven parser that validates each input byte against a 256-entry decode_table and reports the offending offset on the first invalid byte; the final-quad handler in decode_suffix.rs enumerates the four classes of malformed padding and uses a bit mask to detect non-canonical trailing bits per DecodePaddingMode. All buffer accesses use safe slice indexing, and encoded_len performs its arithmetic with checked_mul / checked_add, returning None on usize overflow (the crate-level docs document the corresponding panic). The streaming DecoderReader re-bases error offsets onto the cumulative input_consumed_len and remembers the first padding byte seen so that errors discovered in a later chunk are reported at the position users would expect from non-streaming decode. Together these support parser-impl-safe, parser-impl-correct, algorithm-impl-safe, algorithm-impl-correct, and algorithm-impl-bounds (the crate implements the RFC 4648 base64 codec, justifying impl-parser and impl-algorithm, and does not implement a data structure, interpreter, JIT, protocol, or any concurrency primitive, justifying impl-datastructure, impl-interpreter, impl-jit, impl-protocol, and impl-concurrency) (the encoder and decoder are strictly linear in input length and allocate at most a single output buffer).

Testing is comprehensive. The engine test suite uses rstest_reuse to fan every test across three EngineWrapper implementations (production GeneralPurpose, a deliberately-simple Naive reference, and DecoderReader), giving a strong cross-validation property: each behavioural test runs against the naive implementation as a reference oracle. The suite covers RFC 4648 vectors, randomised roundtrip (~10000 inputs of up to 1000 bytes), exhaustive last-symbol validity for the two-and three-symbol-suffix cases, malleability across all padding modes, slice sizing edge cases, and the streaming reader's behaviour across all input-split positions. The crate ships in-module #[cfg(test)] blocks throughout src/ and an extensive tests/ directory exercising the public API, justifying has-unit-tests and has-integration-tests. Four cargo-fuzz fuzzers in the upstream repo (roundtrip, roundtrip_no_pad, roundtrip_random_config, decode_random) provide ongoing randomised testing, justifying has-fuzz-tests and supporting parser-impl-tested and algorithm-impl-tested. The crate has no dedicated proptest/quickcheck property tests, hence has-property-tests is false; the rstest-driven and randomised tests cover similar ground in practice.

No findings were raised. No malicious patterns, obfuscation, or unexpected side effects were observed, justifying is-benign.

Conclusion

base64 is a small, focused, audit-friendly library that does only what its name says. The crate has no runtime dependencies, no unsafe code, no build-time or install-time code execution, no IO surface, and a thorough cross-validated test suite. It is safe to deploy.

Findings

No findings.

Annotations(10)

Cargo.toml

Manifest declares no build.rs, no [lib] proc-macro = true, no [build-dependencies], and only pure dev-dependencies (criterion, rand, clap, strum, rstest, rstest_reuse, once_cell). The crate's runtime dependency set is empty, justifying has-build-exec and has-install-exec.

src/alphabet.rs

src/alphabet.rs, line 79-125

    pub const fn new(alphabet: &str) -> Result<Self, ParseAlphabetError> {
        let bytes = alphabet.as_bytes();
        if bytes.len() != ALPHABET_SIZE {
            return Err(ParseAlphabetError::InvalidLength);
        }

        {
            let mut index = 0;
            while index < ALPHABET_SIZE {
                let byte = bytes[index];

                // must be ascii printable. 127 (DEL) is commonly considered printable
                // for some reason but clearly unsuitable for base64.
                if !(byte >= 32_u8 && byte <= 126_u8) {
                    return Err(ParseAlphabetError::UnprintableByte(byte));
                }
                // = is assumed to be padding, so cannot be used as a symbol
                if byte == PAD_BYTE {
                    return Err(ParseAlphabetError::ReservedByte(byte));
                }

                // Check for duplicates while staying within what const allows.
                // It's n^2, but only over 64 hot bytes, and only once, so it's likely in the single digit
                // microsecond range.

                let mut probe_index = 0;
                while probe_index < ALPHABET_SIZE {
                    if probe_index == index {
                        probe_index += 1;
                        continue;
                    }

                    let probe_byte = bytes[probe_index];

                    if byte == probe_byte {
                        return Err(ParseAlphabetError::DuplicatedByte(byte));
                    }

                    probe_index += 1;
                }

                index += 1;
            }
        }

        Ok(Self::from_str_unchecked(alphabet))
    }

Alphabet::new is a const fn that validates: exactly 64 bytes, every byte printable ASCII (32-126), no =, and no duplicates (O(n^2) within const-fn limits over a 64-byte alphabet). Misuse therefore fails at compile time for const alphabets.

src/encode.rs

src/encode.rs, line 97-125

pub const fn encoded_len(bytes_len: usize, padding: bool) -> Option<usize> {
    let rem = bytes_len % 3;

    let complete_input_chunks = bytes_len / 3;
    // `?` is disallowed in const, and `let Some(_) = _ else` requires 1.65.0, whereas this
    // messier syntax works on 1.48
    let complete_chunk_output =
        if let Some(complete_chunk_output) = complete_input_chunks.checked_mul(4) {
            complete_chunk_output
        } else {
            return None;
        };

    if rem > 0 {
        if padding {
            complete_chunk_output.checked_add(4)
        } else {
            let encoded_rem = match rem {
                1 => 2,
                // only other possible remainder is 2
                // can't use a separate _ => unreachable!() in const fns in ancient rust versions
                _ => 3,
            };
            complete_chunk_output.checked_add(encoded_rem)
        }
    } else {
        Some(complete_chunk_output)
    }
}

encoded_len performs all arithmetic with checked_mul / checked_add and returns None on usize overflow rather than wrapping silently. Callers (Engine::encode, Engine::encode_slice, encode_with_padding) expect non-None and expect() if overflow occurs, producing a documented panic per the lib.rs "Panics" section.

src/engine/general_purpose/decode.rs

src/engine/general_purpose/decode.rs, line 35-121

pub(crate) fn decode_helper(
    input: &[u8],
    estimate: GeneralPurposeEstimate,
    output: &mut [u8],
    decode_table: &[u8; 256],
    decode_allow_trailing_bits: bool,
    padding_mode: DecodePaddingMode,
) -> Result<DecodeMetadata, DecodeSliceError> {
    let input_complete_nonterminal_quads_len =
        complete_quads_len(input, estimate.rem, output.len(), decode_table)?;

    const UNROLLED_INPUT_CHUNK_SIZE: usize = 32;
    const UNROLLED_OUTPUT_CHUNK_SIZE: usize = UNROLLED_INPUT_CHUNK_SIZE / 4 * 3;

    let input_complete_quads_after_unrolled_chunks_len =
        input_complete_nonterminal_quads_len % UNROLLED_INPUT_CHUNK_SIZE;

    let input_unrolled_loop_len =
        input_complete_nonterminal_quads_len - input_complete_quads_after_unrolled_chunks_len;

    // chunks of 32 bytes
    for (chunk_index, chunk) in input[..input_unrolled_loop_len]
        .chunks_exact(UNROLLED_INPUT_CHUNK_SIZE)
        .enumerate()
    {
        let input_index = chunk_index * UNROLLED_INPUT_CHUNK_SIZE;
        let chunk_output = &mut output[chunk_index * UNROLLED_OUTPUT_CHUNK_SIZE
            ..(chunk_index + 1) * UNROLLED_OUTPUT_CHUNK_SIZE];

        decode_chunk_8(
            &chunk[0..8],
            input_index,
            decode_table,
            &mut chunk_output[0..6],
        )?;
        decode_chunk_8(
            &chunk[8..16],
            input_index + 8,
            decode_table,
            &mut chunk_output[6..12],
        )?;
        decode_chunk_8(
            &chunk[16..24],
            input_index + 16,
            decode_table,
            &mut chunk_output[12..18],
        )?;
        decode_chunk_8(
            &chunk[24..32],
            input_index + 24,
            decode_table,
            &mut chunk_output[18..24],
        )?;
    }

    // remaining quads, except for the last possibly partial one, as it may have padding
    let output_unrolled_loop_len = input_unrolled_loop_len / 4 * 3;
    let output_complete_quad_len = input_complete_nonterminal_quads_len / 4 * 3;
    {
        let output_after_unroll = &mut output[output_unrolled_loop_len..output_complete_quad_len];

        for (chunk_index, chunk) in input
            [input_unrolled_loop_len..input_complete_nonterminal_quads_len]
            .chunks_exact(4)
            .enumerate()
        {
            let chunk_output = &mut output_after_unroll[chunk_index * 3..chunk_index * 3 + 3];

            decode_chunk_4(
                chunk,
                input_unrolled_loop_len + chunk_index * 4,
                decode_table,
                chunk_output,
            )?;
        }
    }

    super::decode_suffix::decode_suffix(
        input,
        input_complete_nonterminal_quads_len,
        output,
        output_complete_quad_len,
        decode_table,
        decode_allow_trailing_bits,
        padding_mode,
    )
}

Decode is a table-driven loop with explicit bounds-derived slicing (chunks_exact) and per-byte validity checks against a 256-entry decode_table; invalid alphabet bytes are reported as DecodeError::InvalidByte with the offending offset. All buffer access is via safe slice indexing, and #![forbid(unsafe_code)] rules out any pointer-level shortcuts. Supports parser-impl-safe.

src/engine/general_purpose/decode_suffix.rs

src/engine/general_purpose/decode_suffix.rs, line 31-109

    for (leftover_index, &b) in input[input_index..].iter().enumerate() {
        // '=' padding
        if b == PAD_BYTE {
            // There can be bad padding bytes in a few ways:
            // 1 - Padding with non-padding characters after it
            // 2 - Padding after zero or one characters in the current quad (should only
            //     be after 2 or 3 chars)
            // 3 - More than two characters of padding. If 3 or 4 padding chars
            //     are in the same quad, that implies it will be caught by #2.
            //     If it spreads from one quad to another, it will be an invalid byte
            //     in the first quad.
            // 4 - Non-canonical padding -- 1 byte when it should be 2, etc.
            //     Per config, non-canonical but still functional non- or partially-padded base64
            //     may be treated as an error condition.

            if leftover_index < 2 {
                // Check for error #2.
                // Either the previous byte was padding, in which case we would have already hit
                // this case, or it wasn't, in which case this is the first such error.
                debug_assert!(
                    leftover_index == 0 || (leftover_index == 1 && padding_bytes_count == 0)
                );
                let bad_padding_index = input_index + leftover_index;
                return Err(DecodeError::InvalidByte(bad_padding_index, b).into());
            }

            if padding_bytes_count == 0 {
                first_padding_offset = leftover_index;
            }

            padding_bytes_count += 1;
            continue;
        }

        // Check for case #1.
        // To make '=' handling consistent with the main loop, don't allow
        // non-suffix '=' in trailing chunk either. Report error as first
        // erroneous padding.
        if padding_bytes_count > 0 {
            return Err(
                DecodeError::InvalidByte(input_index + first_padding_offset, PAD_BYTE).into(),
            );
        }

        last_symbol = b;

        // can use up to 8 * 6 = 48 bits of the u64, if last chunk has no padding.
        // Pack the leftovers from left to right.
        let morsel = decode_table[b as usize];
        if morsel == INVALID_VALUE {
            return Err(DecodeError::InvalidByte(input_index + leftover_index, b).into());
        }

        morsels[morsels_in_leftover] = morsel;
        morsels_in_leftover += 1;
    }

    // If there was 1 trailing byte, and it was valid, and we got to this point without hitting
    // an invalid byte, now we can report invalid length
    if !input.is_empty() && morsels_in_leftover < 2 {
        return Err(DecodeError::InvalidLength(input_index + morsels_in_leftover).into());
    }

    match padding_mode {
        DecodePaddingMode::Indifferent => { /* everything we care about was already checked */ }
        DecodePaddingMode::RequireCanonical => {
            // allow empty input
            if (padding_bytes_count + morsels_in_leftover) % 4 != 0 {
                return Err(DecodeError::InvalidPadding.into());
            }
        }
        DecodePaddingMode::RequireNone => {
            if padding_bytes_count > 0 {
                // check at the end to make sure we let the cases of padding that should be InvalidByte
                // get hit
                return Err(DecodeError::InvalidPadding.into());
            }
        }
    }

Final-quad handling explicitly enumerates the four classes of malformed padding (non-padding after padding, padding too early, excess padding, non-canonical padding) and uses leftover_num & mask (lines 133-141) to detect non-canonical trailing bits per the encoded RFC 4648 reference. DecodePaddingMode selects between Indifferent / RequireCanonical / RequireNone. Supports parser-impl-correct.

src/engine/general_purpose/mod.rs

src/engine/general_purpose/mod.rs, line 17-22

/// A general-purpose base64 engine.
///
/// - It uses no vector CPU instructions, so it will work on any system.
/// - It is reasonably fast (~2-3GiB/s).
/// - It is not constant-time, though, so it is vulnerable to timing side-channel attacks. For loading cryptographic keys, etc, it is suggested to use the forthcoming constant-time implementation.

Documentation explicitly states the engine is not constant-time and is vulnerable to timing side-channels, directing users with secret-dependent inputs (cryptographic keys) to await a planned constant-time engine. The crate does not itself implement or use cryptography, justifying uses-crypto and impl-crypto.

src/engine/tests.rs

src/engine/tests.rs, line 26-42

#[template]
#[rstest(engine_wrapper,
case::general_purpose(GeneralPurposeWrapper {}),
case::naive(NaiveWrapper {}),
case::decoder_reader(DecoderReaderEngineWrapper {}),
)]
fn all_engines<E: EngineWrapper>(engine_wrapper: E) {}

/// Some decode tests don't make sense for use with `DecoderReader` as they are difficult to
/// reason about or otherwise inapplicable given how DecoderReader slice up its input along
/// chunk boundaries.
#[template]
#[rstest(engine_wrapper,
case::general_purpose(GeneralPurposeWrapper {}),
case::naive(NaiveWrapper {}),
)]
fn all_engines_except_decoder_reader<E: EngineWrapper>(engine_wrapper: E) {}

Engine tests use rstest_reuse to fan every test out across three EngineWrapper implementations (GeneralPurpose, Naive, DecoderReader). Roundtrip, malleability, padding-mode, invalid-byte, invalid-last-symbol, and slice-fits tests therefore exercise both production paths and the streaming reader against a deliberately-simple reference. Supports parser-impl-tested and algorithm-impl-tested.

src/lib.rs

src/lib.rs, line 233-233

#![forbid(unsafe_code)]

Crate-level #![forbid(unsafe_code)] rules out any unsafe block, FFI, or raw-pointer code in the published library, justifying uses-unsafe.

src/read/decoder.rs

src/read/decoder.rs, line 132-206

    /// Decode the requested number of bytes from the b64 buffer into the provided buffer. It's the
    /// caller's responsibility to choose the number of b64 bytes to decode correctly.
    ///
    /// Returns a Result with the number of decoded bytes written to `buf`.
    ///
    /// # Panics
    ///
    /// panics if `buf` is too small
    fn decode_to_buf(&mut self, b64_len_to_decode: usize, buf: &mut [u8]) -> io::Result<usize> {
        debug_assert!(self.b64_len >= b64_len_to_decode);
        debug_assert!(self.b64_offset + self.b64_len <= BUF_SIZE);
        debug_assert!(!buf.is_empty());

        let b64_to_decode = &self.b64_buffer[self.b64_offset..self.b64_offset + b64_len_to_decode];
        let decode_metadata = self
            .engine
            .internal_decode(
                b64_to_decode,
                buf,
                self.engine.internal_decoded_len_estimate(b64_len_to_decode),
            )
            .map_err(|dse| match dse {
                DecodeSliceError::DecodeError(de) => {
                    match de {
                        DecodeError::InvalidByte(offset, byte) => {
                            match (byte, self.padding_offset) {
                                // if there was padding in a previous block of decoding that happened to
                                // be correct, and we now find more padding that happens to be incorrect,
                                // to be consistent with non-reader decodes, record the error at the first
                                // padding
                                (PAD_BYTE, Some(first_pad_offset)) => {
                                    DecodeError::InvalidByte(first_pad_offset, PAD_BYTE)
                                }
                                _ => {
                                    DecodeError::InvalidByte(self.input_consumed_len + offset, byte)
                                }
                            }
                        }
                        DecodeError::InvalidLength(len) => {
                            DecodeError::InvalidLength(self.input_consumed_len + len)
                        }
                        DecodeError::InvalidLastSymbol(offset, byte) => {
                            DecodeError::InvalidLastSymbol(self.input_consumed_len + offset, byte)
                        }
                        DecodeError::InvalidPadding => DecodeError::InvalidPadding,
                    }
                }
                DecodeSliceError::OutputSliceTooSmall => {
                    unreachable!("buf is sized correctly in calling code")
                }
            })
            .map_err(|e| io::Error::new(io::ErrorKind::InvalidData, e))?;

        if let Some(offset) = self.padding_offset {
            // we've already seen padding
            if decode_metadata.decoded_len > 0 {
                // we read more after already finding padding; report error at first padding byte
                return Err(io::Error::new(
                    io::ErrorKind::InvalidData,
                    DecodeError::InvalidByte(offset, PAD_BYTE),
                ));
            }
        }

        self.padding_offset = self.padding_offset.or(decode_metadata
            .padding_offset
            .map(|offset| self.input_consumed_len + offset));
        self.input_consumed_len += b64_len_to_decode;
        self.b64_offset += b64_len_to_decode;
        self.b64_len -= b64_len_to_decode;

        debug_assert!(self.b64_offset + self.b64_len <= BUF_SIZE);

        Ok(decode_metadata.decoded_len)
    }

DecoderReader::decode_to_buf re-bases reported error offsets onto the global input_consumed_len and tracks the first observed padding so that padding errors from a later chunk are reported at the first padding byte, matching non-streaming decode semantics.

src/write/encoder.rs

src/write/encoder.rs, line 400-407

impl<'e, E: Engine, W: io::Write> Drop for EncoderWriter<'e, E, W> {
    fn drop(&mut self) {
        if !self.panicked {
            // like `BufWriter`, ignore errors during drop
            let _ = self.write_final_leftovers();
        }
    }
}

Drop impl for EncoderWriter silently swallows errors during finalization (matching BufWriter), and finish() must be called manually if those errors need to be observed. Documented at lines 18-19 and 47-50.