Packages
Browse packages indexed by the registry.
Safe Rust wrappers for NVIDIA cuFile (GPUDirect Storage). Linux-only; scaffolding at v0.1.
Raw FFI bindings and dynamic loader for NVIDIA cuFile (GPUDirect Storage, Linux-only).
Safe Rust wrappers for CUPTI. Scaffolding at v0.1.
Raw FFI bindings and dynamic loader for CUPTI (CUDA Profiling Tools Interface).
Safe Rust wrappers for NVIDIA cuRAND (pseudo- and quasi-random number generation).
Raw FFI bindings and dynamic loader for NVIDIA cuRAND.
Safe Rust wrappers for NVIDIA cuSOLVER (dense LU factorization at v0.1).
Raw FFI bindings and dynamic loader for NVIDIA cuSOLVER (Dn subset).
Safe Rust wrappers for NVIDIA cuSPARSE (generic-API SpMV at v0.1).
Raw FFI bindings and dynamic loader for NVIDIA cuSPARSE.
Safe Rust wrappers for NVIDIA cuTENSOR. Scaffolding at v0.1.
Raw FFI bindings and dynamic loader for NVIDIA cuTENSOR (tensor contraction).
Safe Rust wrapper for compiled CUTLASS kernels: plan-based GEMM and grouped GEMM with caller-supplied workspace, typed device-buffer arguments, and capture-safe launch.
Compiled CUTLASS template instantiations for the baracuda ecosystem. Hosts curated .cu kernel sources, builds them via baracuda-forge, exposes extern "C" entry points for the safe baracuda-cutlass crate.
Header acquisition for NVIDIA CUTLASS as a baracuda workspace dependency. Sparse-checkout fetch with file-locked caching; emits cargo:include for downstream build.rs consumers.
Safe Rust wrappers for NVIDIA cuVS (GPU vector search / approximate nearest neighbours).
Raw FFI bindings and dynamic loader for NVIDIA cuVS (GPU vector search / ANN).
Safe Rust wrappers for NVIDIA CV-CUDA. Scaffolding at v0.1.
Raw FFI bindings and dynamic loader for NVIDIA CV-CUDA (computer-vision operators).
Safe Rust wrappers for the CUDA Driver API (devices, contexts, streams, events, memory, kernels, graphs).
Safe, typed Rust wrappers for NVIDIA FlashInfer's inference-serving kernels: batched paged-KV attention decode, decode-time KV-cache append, cascade / prefix-cache attention-state merge, and sort-free top-K / top-P / min-P sampling. The canonical vLLM-style serving surface for the baracuda CUDA stack. Apache-2.0 (FlashInfer upstream).
Raw C-ABI FFI surface for the vendored FlashInfer inference kernels (paged-KV decode, paged-KV append, cascade state merge, sort-free sampling). The launcher shims + vendored FlashInfer sources are compiled by `baracuda-kernels-sys`; this crate is a thin re-export facade so the FlashInfer C-ABI is reachable under its own crate name. Apache-2.0 (FlashInfer upstream) — see `crates/baracuda-kernels-sys/vendor/flashinfer/`.
Build-time CUDA kernel compiler for the baracuda ecosystem: nvcc-driven incremental builds, parallel compilation, GPU auto-detection, and CUTLASS / custom git dependency support.
Driver-free classifier vocabulary for the baracuda kernel facade: the pure-data kernel-dispatch tags (ElementKind, ArchSku, OpCategory, layout / op-family tags), the StructureKey / OperandDesc classifier key + `structure_key` derivation, the dispatch-table types, and the PlanPreference / PrecisionGuarantee descriptors. Carved out of baracuda-kernels-types so neutral consumers (the kernel generator, selectors, Fuel) depend on the vocabulary WITHOUT the CUDA driver (baracuda-driver -> baracuda-cuda-sys) transitive pull. The device-view half (MatrixRef / TensorRef / Workspace + the tensor->OperandDesc adapter) stays in baracuda-kernels-types, which re-exports this crate so its public API is unchanged. Seeds the ThinkersJournal KISS (Kernel Interface Standards Suite) classifier-vocabulary reference implementation.
Kernel generator + the live JIT synthesizer for the Fuel kernel seam. The SUPPORTED surface is the `seam` feature (BaracudaSynthesizer, the fuel-kernel-seam Synthesizer impl); the generator internals (IR, emitters, contracts, dispatch artifacts) are alpha-fluid — pin exact versions.
Unified ML op facade for the baracuda CUDA ecosystem. Exposes every primitive an ML framework would expect (union of PyTorch torch.* + nn.functional and JAX lax.* / numpy ops) through a single Plan-based Rust surface, internally dispatching to baracuda-cutlass, the baracuda-* NVIDIA-library wrappers, or bespoke baracuda-kernels-sys kernels.
Compiled bespoke .cu kernel template instantiations for the baracuda ML kernel facade plus C-ABI FFI facades for the library-backed plans (cuDNN conv/pool, cuSOLVER linalg, cuFFT/cuRAND, CUTLASS GEMM re-export). Hosts curated CUDA kernel sources (int8/FP8/int4/bin GEMM RRR, elementwise, reduce, norm, attention, …), builds them via baracuda-forge, exposes extern "C" entry points for the safe baracuda-kernels crate. CUTLASS template kernels live in the sibling baracuda-cutlass-kernels-sys crate and are re-exported here under the unified baracuda_kernels_gemm_* namespace.
Shared type vocabulary for the baracuda ML kernel facade: Element / IntElement / FpElement / BiasElement trait hierarchy, layout / epilogue / activation tags, MatrixRef / TensorRef views, PlanPreference, PrecisionGuarantee, and Workspace. Lifted from baracuda-cutlass so that baracuda-kernels and the per-library wrapper crates can share one vocabulary.
Megatron-LM-style tensor-parallel primitives (Column / Row Parallel Linear) for the baracuda CUDA stack. Pure-composition crate — local GEMM via baracuda-cublas + cross-rank collectives via baracuda-nccl. No new CUDA kernels. NEW in Phase 57; deliberate scope expansion (distributed-training-framework-adjacent). Off-by-default in baracuda-kernels via the `megatron_tp` cargo feature so non-distributed consumers don't pay the dep surface cost. Algorithmic reference: Shoeybi et al. arXiv:1909.08053 (NVIDIA Megatron-LM, Apache-2.0).
Safe Rust wrappers for NVIDIA NCCL (multi-GPU collective communication).
Raw FFI bindings and dynamic loader for NVIDIA NCCL (multi-GPU collective communication).
Safe Rust wrappers for NVIDIA NPP (Performance Primitives). Core + signal subset at v0.1.
Raw FFI bindings and dynamic loader for NVIDIA NPP (Performance Primitives).
Safe Rust wrappers for NVIDIA nvCOMP (GPU compression). Scaffolding at v0.1.
Raw FFI bindings and dynamic loader for NVIDIA nvCOMP (GPU compression).
Safe Rust wrappers for NVIDIA nvImageCodec (unified GPU image decode).
Raw FFI bindings and dynamic loader for NVIDIA nvImageCodec (unified GPU image codecs).
Safe Rust wrappers for NVIDIA nvJitLink (CUDA 12.0+ JIT linker).
Raw FFI bindings and dynamic loader for NVIDIA nvJitLink (CUDA 12.0+ JIT linker).
Safe Rust wrappers for NVIDIA nvJPEG (GPU JPEG decode).
Raw FFI bindings and dynamic loader for NVIDIA nvJPEG.
Safe Rust wrappers for the NVIDIA Management Library (NVML) — driver-bundled GPU monitoring.
Raw FFI bindings and dynamic loader for the NVIDIA Management Library (NVML).
Safe Rust wrappers for NVIDIA NVRTC (compile CUDA C++ to PTX at runtime).
Raw FFI bindings and dynamic loader for NVIDIA NVRTC (runtime CUDA-C++ to PTX compiler).
Safe Rust wrappers for NVIDIA NVSHMEM host API (symmetric-heap one-sided GPU communication).
Raw FFI bindings and dynamic loader for NVIDIA NVSHMEM host library (OpenSHMEM symmetric-heap one-sided comms on GPUs).
Optimizer kernels (Adam / LAMB / SGD) for the baracuda CUDA stack, built on the multi_tensor_apply idiom vendored from NVIDIA Apex (BSD-3-Clause). One launch over thousands of parameter tensors — critical for the optimizer step on large-model training stacks. NEW in Phase 49; deliberate scope expansion (training-framework-adjacent). Off-by-default in baracuda-kernels via the `optim` cargo feature so inference-only consumers don't pay the FFI surface cost.
Safe Rust wrapper for baracuda's clean-fork of Hiroyuki Ootomo's ozIMMU — Ozaki-scheme FP64 GEMM that synthesizes a DGEMM from S^2 int8 tensor-core matmuls. Provides an RAII handle + drop-in `dgemm` shim suitable for the `BackendKind::Ozaki` path of `baracuda-kernels`'s `GemmPlan`. Opt-in (NOT bit-equivalent to native DGEMM); the default FP64 path stays on CUTLASS / cuBLAS.
Build + raw FFI bindings to baracuda's clean-fork of Hiroyuki Ootomo's ozIMMU — the Ozaki-scheme FP64 GEMM library that synthesizes a DGEMM from S² int8 tensor-core matmuls. Phase 44b internalized the upstream sources under `cuda/` (no more `vendor/` subdir; cutf submodule eliminated). Linked statically into the baracuda CUDA stack; consumed by the safe wrapper crate `baracuda-ozimmu`. MIT-licensed (original ozIMMU MIT — see `ATTRIBUTION.md`).