cargo / base64ct / audit
cargo : base64ct @ 1.8.3
PE Patrick Elsen signed 2026-05-27 published 2026-05-27

Claims

algorithm-impl-boundsalgorithm-impl-correctalgorithm-impl-safealgorithm-impl-testedhas-binarieshas-build-exechas-fuzz-testshas-install-exechas-integration-testshas-property-testshas-unit-testsimpl-algorithmimpl-concurrencyimpl-cryptoimpl-datastructureimpl-interpreterimpl-jitimpl-parserimpl-protocolis-benignparser-impl-correctparser-impl-safeparser-impl-testedunsafe-documentedunsafe-minimalunsafe-safeunsafe-testeduses-concurrencyuses-cryptouses-environmentuses-execuses-filesystemuses-interpreteruses-jituses-networkuses-unsafe

Summary

Pure-Rust constant-time Base64 codec for PEM-formatted cryptographic material. Branchless 6-bit decode/encode and deferred error accumulation preserve the data-independent timing claim. Two low-severity findings: decode_in_place diverges from decode on non-canonical trailing-bit inputs, and encoded_len silently returns 0 on usize overflow.

Report

Subject

base64ct is a pure-Rust, no_std, allocation-optional Base64 encoder/decoder published by the RustCrypto project. It implements the standard, URL-safe, bcrypt, crypt(3) (SHA-crypt and a deprecated big-endian variant), and PBKDF2 alphabets in both padded and unpadded forms, and provides streaming Encoder / Decoder types with optional line wrapping. The crate is targeted at parsing PEM-format cryptographic material (used by pem-rfc7468, pkcs8, ssh-key, etc.), and uses branchless integer arithmetic on every per-symbol code path so that decode timing depends only on input length, not on input bytes. The crate does not itself perform cryptography (no hashing, signing, or encryption), justifying uses-crypto and impl-crypto; the constant-time design exists to protect callers that handle secrets immediately after decoding.

Methodology

The published .crate archive was compared against the upstream Git checkout at the commit recorded in .cargo_vcs_info.json. All source under src/ (~1500 lines across the alphabet definitions, the shared Alphabet trait with the branchless encode/decode core, the in-place and streaming codecs, the line-ending helpers, the buffered Decoder line reader, and the error types) was read in full, and all integration tests under tests/ (alphabet-specific test vectors plus property-based equivalence tests against the base64 reference crate) were surveyed. Cargo.toml was reviewed for build-time hooks, proc macros, default features, and dependency surface. The .github/workflows/base64ct.yml CI workflow in the parent repository was inspected. No code was compiled or executed.

Results

The comparison between the published crate contents and the upstream Git checkout shows that all source files match byte-for-byte; the only differences are the added Cargo.lock and .cargo_vcs_info.json, plus cargo's standard Cargo.toml normalisation.

The crate ships no binary artefacts, no build.rs, and is not a proc-macro crate (Cargo.toml declares only [features] alloc/std, no [lib] proc-macro = true), justifying has-binaries, has-build-exec, and has-install-exec. Runtime capabilities are limited to in-memory buffer manipulation: the source contains no std::net, std::fs, std::env, std::process, or threading code, justifying uses-network, uses-filesystem, uses-environment, uses-exec, uses-jit, uses-interpreter, and uses-concurrency. The crate has only two dev-dependencies (base64 for equivalence testing and proptest); both are non-shipping. The code is benign in intent (RustCrypto upstream, no obfuscation, no telemetry), justifying is-benign.

The crate implements both a Base64 parser (decoder for several alphabets, with documented constant-time properties) and a non-trivial branchless encoding algorithm (justifying impl-parser and impl-algorithm); it does not implement a data structure, interpreter, JIT, protocol, or concurrency primitive, justifying impl-datastructure, impl-interpreter, impl-jit, impl-protocol, and impl-concurrency. The branchless decode-step combiner in src/alphabet.rs is the security-relevant primitive: each candidate range or equality match is folded into a 6-bit lane via signed-mask arithmetic, the per-symbol output is OR-accumulated into an error word that is only branched on after all data-dependent work is done, and validation of the last block round-trips through the encoder using a non-short-circuiting fold-XOR comparison. The test suite covers each alphabet with hand-curated vectors plus property-based equivalence against the upstream base64 crate for the standard and unpadded standard alphabets (256-byte inputs, incremental encode/decode, line-wrapped decode at variable widths). In-module #[cfg(test)] blocks cover each alphabet's encode/decode paths and the streaming Encoder/Decoder edge cases, with integration tests in tests/ providing the alphabet-specific vectors and property-based equivalence against the upstream base64 crate, justifying has-unit-tests and has-integration-tests. The branchless decode-step combiner, validated lengths feeding into decode_in_place, and OR-accumulated error word together justify parser-impl-safe. There are no fuzz tests (has-fuzz-tests = false), but the proptest coverage and the cited C++ reference implementation give reasonable assurance of correctness, justifying parser-impl-correct, parser-impl-tested, algorithm-impl-correct, and algorithm-impl-tested. Per-symbol cost is constant and per-input cost is linear in the input length, with no recursion or unbounded loops, justifying algorithm-impl-bounds.

Three unsafe blocks were reviewed, all in src/encoding.rs: aliased raw-pointer reads/writes in decode_in_place (sound because the 4-byte read window and 3-byte write window for each chunk are sequenced and never concurrent, with bounds guarded by debug assertions and prior length math), a get_unchecked on the resulting prefix (bounds derived from the validated dlen), and two from_utf8_unchecked calls that wrap output produced exclusively by the alphabet's encode function (whose output is by construction printable ASCII). All have SAFETY comments; usage is minimal and matches the pattern needed to avoid bounds-check overhead on the hot decode path. Round-trip property tests exercise these paths, justifying uses-unsafe, unsafe-safe, unsafe-documented, unsafe-minimal, unsafe-tested. There is no FFI, no Miri job in the upstream CI workflow inspected.

Two low-severity findings were filed. FINDING-1 documents that Encoding::decode_in_place does not call validate_last_block and therefore diverges from Encoding::decode by accepting non-canonical unpadded inputs (e.g. "Mi"); the surrounding doc comment notes a narrower padding-validation TODO but not this broader divergence. FINDING-2 notes that Encoding::encoded_len silently returns 0 on usize overflow rather than signalling an error.

Conclusion

The crate is small, focused, well-documented, and built around a security goal (sidechannel resistance for Base64-encoded cryptographic material) that is faithfully reflected in the implementation. The published source matches the upstream Git tree, ships no binary artefacts or build-time code, and exercises no I/O or OS-facing capabilities. The two findings are minor and do not affect the documented security properties.

Findings(2)

FINDING-1 correctness low

decode_in_place skips round-trip last-block validation

The non-streaming Encoding::decode path calls validate_last_block to round-trip the final block and reject non-canonical encodings where the trailing bits of the last Base64 character do not decode to zero (e.g. unpadded input "Mi" is rejected). Encoding::decode_in_place (src/encoding.rs:106) does not call validate_last_block and therefore accepts non-canonical encodings that the non-in-place path rejects.

The doc comment on decode_in_place calls out a narrower limitation ("this method does not (yet) validate that padding is well-formed") but does not mention the broader divergence in canonical-encoding handling. The associated tests (tests/standard.rs reject_non_canonical_encoding, tests/pbkdf2.rs reject_non_canonical_encoding) only exercise decode, not decode_in_place.

Impact is limited: callers that rely on canonical-encoding rejection for security (e.g. signature-malleability mitigations) should use decode. For PEM parsing the divergence is benign.

FINDING-2 quality low

encoded_len silently returns 0 on usize overflow

Encoding::encoded_len (src/encoding.rs:250) returns 0 when encoded_len_inner overflows usize (input length greater than usize::MAX/4). The trait doc comment notes this with a "WARNING". A caller using the returned length to size an output buffer would then write 0 bytes for very large inputs, which can mask programming errors.

On 32-bit targets the threshold is reachable in principle (~1 GiB inputs); on 64-bit targets it is not a realistic input. The accompanying encode_string panics in this case (expect("input is too big")), so the dangerous behaviour is only via the slice-based API. Returning a Result or Option would be preferable but is a breaking change.

Annotations(3)

src/alphabet.rs

src/alphabet.rs, line 36-104

    fn decode_3bytes(src: &[u8], dst: &mut [u8]) -> i16 {
        debug_assert_eq!(src.len(), 4);
        debug_assert!(dst.len() >= 3, "dst too short: {}", dst.len());

        let c0 = Self::decode_6bits(src[0]);
        let c1 = Self::decode_6bits(src[1]);
        let c2 = Self::decode_6bits(src[2]);
        let c3 = Self::decode_6bits(src[3]);

        dst[0] = ((c0 << 2) | (c1 >> 4)) as u8;
        dst[1] = ((c1 << 4) | (c2 >> 2)) as u8;
        dst[2] = ((c2 << 6) | c3) as u8;

        ((c0 | c1 | c2 | c3) >> 8) & 1
    }

    /// Decode 6-bits of a Base64 message.
    fn decode_6bits(src: u8) -> i16 {
        let mut ret: i16 = -1;

        for step in Self::DECODER {
            ret += match step {
                DecodeStep::Range(range, offset) => {
                    // Compute exclusive range from inclusive one
                    let start = *range.start() as i16 - 1;
                    let end = *range.end() as i16 + 1;
                    (((start - src as i16) & (src as i16 - end)) >> 8) & (src as i16 + *offset)
                }
                DecodeStep::Eq(value, offset) => {
                    let start = *value as i16 - 1;
                    let end = *value as i16 + 1;
                    (((start - src as i16) & (src as i16 - end)) >> 8) & *offset
                }
            };
        }

        ret
    }

    /// Encode 3-bytes of a Base64 message.
    #[inline(always)]
    fn encode_3bytes(src: &[u8], dst: &mut [u8]) {
        debug_assert_eq!(src.len(), 3);
        debug_assert!(dst.len() >= 4, "dst too short: {}", dst.len());

        let b0 = src[0] as i16;
        let b1 = src[1] as i16;
        let b2 = src[2] as i16;

        dst[0] = Self::encode_6bits(b0 >> 2);
        dst[1] = Self::encode_6bits(((b0 << 4) | (b1 >> 4)) & 63);
        dst[2] = Self::encode_6bits(((b1 << 2) | (b2 >> 6)) & 63);
        dst[3] = Self::encode_6bits(b2 & 63);
    }

    /// Encode 6-bits of a Base64 message.
    #[inline(always)]
    fn encode_6bits(src: i16) -> u8 {
        let mut diff = src + Self::BASE as i16;

        for &step in Self::ENCODER {
            diff += match step {
                EncodeStep::Apply(threshold, offset) => ((threshold as i16 - diff) >> 8) & offset,
                EncodeStep::Diff(threshold, offset) => ((threshold as i16 - src) >> 8) & offset,
            };
        }

        diff as u8
    }

Constant-time encode/decode core. decode_6bits ORs together a set of branchless range/equality predicates implemented via signed-arithmetic masks (((start - x) & (x - end)) >> 8), accumulating a per-alphabet offset; mismatching characters cause the high bit to propagate into the error accumulator ((c0 | c1 | c2 | c3) >> 8) & 1). encode_6bits mirrors this. No data-dependent branches, no lookup tables. This is the constant-time core that the crate's existence is justified by (cf. README and the cited Util::Lookup paper). Justifies impl-algorithm, algorithm-impl-safe, algorithm-impl-correct, algorithm-impl-bounds, impl-parser, parser-impl-safe, parser-impl-correct.

src/encoding.rs

src/encoding.rs, line 120-170

        for chunk in 0..full_chunks {
            // SAFETY: `p3` and `p4` point inside `buf`, while they may overlap,
            // read and write are clearly separated from each other and done via
            // raw pointers.
            #[allow(unsafe_code)]
            unsafe {
                debug_assert!(3 * chunk + 3 <= buf.len());
                debug_assert!(4 * chunk + 4 <= buf.len());

                let p3 = buf.as_mut_ptr().add(3 * chunk) as *mut [u8; 3];
                let p4 = buf.as_ptr().add(4 * chunk) as *const [u8; 4];

                let mut tmp_out = [0u8; 3];
                err |= Self::decode_3bytes(&*p4, &mut tmp_out);
                *p3 = tmp_out;
            }
        }

        let src_rem_pos = 4 * full_chunks;
        let src_rem_len = buf.len() - src_rem_pos;
        let dst_rem_pos = 3 * full_chunks;
        let dst_rem_len = dlen - dst_rem_pos;

        err |= !(src_rem_len == 0 || src_rem_len >= 2) as i16;
        let mut tmp_in = [b'A'; 4];
        tmp_in[..src_rem_len].copy_from_slice(&buf[src_rem_pos..]);
        let mut tmp_out = [0u8; 3];

        err |= Self::decode_3bytes(&tmp_in, &mut tmp_out);

        if err == 0 {
            // SAFETY: `dst_rem_len` is always smaller than 4, so we don't
            // read outside of `tmp_out`, write and the final slicing never go
            // outside of `buf`.
            #[allow(unsafe_code)]
            unsafe {
                debug_assert!(dst_rem_pos + dst_rem_len <= buf.len());
                debug_assert!(dst_rem_len <= tmp_out.len());
                debug_assert!(dlen <= buf.len());

                core::ptr::copy_nonoverlapping(
                    tmp_out.as_ptr(),
                    buf.as_mut_ptr().add(dst_rem_pos),
                    dst_rem_len,
                );
                Ok(buf.get_unchecked(..dlen))
            }
        } else {
            Err(InvalidEncodingError)
        }
    }

In-place decode loop uses three unsafe blocks: aliased mut-pointer arithmetic to overlay 3-byte writes onto a 4-byte read window (lines 124-135), and get_unchecked for the final slice (lines 154-166). Aliasing is sound because each chunk reads 4 bytes via *const and then writes 3 bytes via *mut to a non-overlapping prefix; reads and writes are sequenced and never concurrent. Bounds for add, copy_nonoverlapping, and get_unchecked are checked against buf.len() via debug_assert and are guaranteed by the prior dlen calculation and chunk loop bounds. Justifies uses-unsafe, unsafe-safe, unsafe-documented, unsafe-minimal. Supports FINDING-1.

src/encoding.rs, line 226-248


        debug_assert!(str::from_utf8(dst).is_ok());

        // SAFETY: values written by `encode_3bytes` are valid one-byte UTF-8 chars
        #[allow(unsafe_code)]
        Ok(unsafe { str::from_utf8_unchecked(dst) })
    }

    #[cfg(feature = "alloc")]
    fn encode_string(input: &[u8]) -> String {
        let elen = encoded_len_inner(input.len(), T::PADDED).expect("input is too big");
        let mut dst = vec![0u8; elen];
        let res = Self::encode(input, &mut dst).expect("encoding error");

        debug_assert_eq!(elen, res.len());
        debug_assert!(str::from_utf8(&dst).is_ok());

        // SAFETY: `dst` is fully written and contains only valid one-byte UTF-8 chars
        #[allow(unsafe_code)]
        unsafe {
            String::from_utf8_unchecked(dst)
        }
    }

Two from_utf8_unchecked uses: str::from_utf8_unchecked on the encoded slice and String::from_utf8_unchecked on the encoded vector. Both are sound because encode_3bytes only writes bytes produced by encode_6bits, which by construction always lies in the printable ASCII alphabet (one-byte UTF-8) of the active alphabet, plus the = padding constant. A debug_assert!(str::from_utf8(...).is_ok()) guards each call in debug builds. Justifies unsafe-safe.

src/encoding.rs, line 264-297

#[inline(always)]
pub(crate) fn decode_padding(input: &[u8]) -> Result<(usize, i16), InvalidEncodingError> {
    if input.len() % 4 != 0 {
        return Err(InvalidEncodingError);
    }

    let unpadded_len = match *input {
        [.., b0, b1] => is_pad_ct(b0)
            .checked_add(is_pad_ct(b1))
            .and_then(|len| len.try_into().ok())
            .and_then(|len| input.len().checked_sub(len))
            .ok_or(InvalidEncodingError)?,
        _ => input.len(),
    };

    let padding_len = input
        .len()
        .checked_sub(unpadded_len)
        .ok_or(InvalidEncodingError)?;

    let err = match *input {
        [.., b0] if padding_len == 1 => is_pad_ct(b0) ^ 1,
        [.., b0, b1] if padding_len == 2 => (is_pad_ct(b0) & is_pad_ct(b1)) ^ 1,
        _ => {
            if padding_len == 0 {
                0
            } else {
                return Err(InvalidEncodingError);
            }
        }
    };

    Ok((unpadded_len, err))
}

decode_padding reports padding-byte mismatches as a deferred i16 error to be OR-combined with the per-block decode error before any branch on the result. is_pad_ct (line 358) is a branchless equality test against =. This pattern, combined with the unconditional fold-based comparison in validate_last_block (lines 326-331), preserves the constant-time property when decoding rejects malformed input. The if input.len() % 4 != 0 length check is data-independent (length is public). Justifies the "data"-constant-time claim in the README.

src/encoding.rs, line 106-170

    fn decode_in_place(mut buf: &mut [u8]) -> Result<&[u8], InvalidEncodingError> {
        // TODO: eliminate unsafe code when LLVM12 is stable
        // See: https://github.com/rust-lang/rust/issues/80963
        let mut err = if T::PADDED {
            let (unpadded_len, e) = decode_padding(buf)?;
            buf = &mut buf[..unpadded_len];
            e
        } else {
            0
        };

        let dlen = decoded_len(buf.len());
        let full_chunks = buf.len() / 4;

        for chunk in 0..full_chunks {
            // SAFETY: `p3` and `p4` point inside `buf`, while they may overlap,
            // read and write are clearly separated from each other and done via
            // raw pointers.
            #[allow(unsafe_code)]
            unsafe {
                debug_assert!(3 * chunk + 3 <= buf.len());
                debug_assert!(4 * chunk + 4 <= buf.len());

                let p3 = buf.as_mut_ptr().add(3 * chunk) as *mut [u8; 3];
                let p4 = buf.as_ptr().add(4 * chunk) as *const [u8; 4];

                let mut tmp_out = [0u8; 3];
                err |= Self::decode_3bytes(&*p4, &mut tmp_out);
                *p3 = tmp_out;
            }
        }

        let src_rem_pos = 4 * full_chunks;
        let src_rem_len = buf.len() - src_rem_pos;
        let dst_rem_pos = 3 * full_chunks;
        let dst_rem_len = dlen - dst_rem_pos;

        err |= !(src_rem_len == 0 || src_rem_len >= 2) as i16;
        let mut tmp_in = [b'A'; 4];
        tmp_in[..src_rem_len].copy_from_slice(&buf[src_rem_pos..]);
        let mut tmp_out = [0u8; 3];

        err |= Self::decode_3bytes(&tmp_in, &mut tmp_out);

        if err == 0 {
            // SAFETY: `dst_rem_len` is always smaller than 4, so we don't
            // read outside of `tmp_out`, write and the final slicing never go
            // outside of `buf`.
            #[allow(unsafe_code)]
            unsafe {
                debug_assert!(dst_rem_pos + dst_rem_len <= buf.len());
                debug_assert!(dst_rem_len <= tmp_out.len());
                debug_assert!(dlen <= buf.len());

                core::ptr::copy_nonoverlapping(
                    tmp_out.as_ptr(),
                    buf.as_mut_ptr().add(dst_rem_pos),
                    dst_rem_len,
                );
                Ok(buf.get_unchecked(..dlen))
            }
        } else {
            Err(InvalidEncodingError)
        }
    }

decode_in_place validates padding length via decode_padding but does not call validate_last_block, so non-canonical encodings whose trailing bits do not decode to zero (e.g. unpadded "Mi") decode successfully here while decode rejects them. Documented narrowly as a padding-validation TODO; the broader divergence is reported in FINDING-1.

src/encoding.rs, line 250-252

    fn encoded_len(bytes: &[u8]) -> usize {
        encoded_len_inner(bytes.len(), T::PADDED).unwrap_or(0)
    }

encoded_len swallows overflow by returning 0. The trait doc comment warns about this. Supports FINDING-2.

tests/proptests.rs

Property-based equivalence suite that compares this crate's standard padded and unpadded encodings against the upstream base64 crate over 256-byte random inputs, covers incremental decoding and line-wrapping, and includes round-trip tests for the incremental encoder. Justifies has-property-tests, parser-impl-tested, algorithm-impl-tested, unsafe-tested.