Safe Rust wrappers for NVIDIA cuFile (GPUDirect Storage). Linux-only; scaffolding at v0.1.

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA cuFile (GPUDirect Storage, Linux-only).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for CUPTI. Scaffolding at v0.1.

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for CUPTI (CUDA Profiling Tools Interface).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA cuRAND (pseudo- and quasi-random number generation).

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA cuRAND.

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA cuSOLVER (dense LU factorization at v0.1).

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA cuSOLVER (Dn subset).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA cuSPARSE (generic-API SpMV at v0.1).

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA cuSPARSE.

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA cuTENSOR. Scaffolding at v0.1.

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA cuTENSOR (tensor contraction).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrapper for compiled CUTLASS kernels: plan-based GEMM and grouped GEMM with caller-supplied workspace, typed device-buffer arguments, and capture-safe launch.

0 audits · 45 versions · updated 2026-06-20

Compiled CUTLASS template instantiations for the baracuda ecosystem. Hosts curated .cu kernel sources, builds them via baracuda-forge, exposes extern "C" entry points for the safe baracuda-cutlass crate.

0 audits · 45 versions · updated 2026-06-20

Header acquisition for NVIDIA CUTLASS as a baracuda workspace dependency. Sparse-checkout fetch with file-locked caching; emits cargo:include for downstream build.rs consumers.

0 audits · 45 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA cuVS (GPU vector search / approximate nearest neighbours).

0 audits · 6 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA cuVS (GPU vector search / ANN).

0 audits · 6 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA CV-CUDA. Scaffolding at v0.1.

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA CV-CUDA (computer-vision operators).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for the CUDA Driver API (devices, contexts, streams, events, memory, kernels, graphs).

0 audits · 49 versions · updated 2026-06-20

Safe, typed Rust wrappers for NVIDIA FlashInfer's inference-serving kernels: batched paged-KV attention decode, decode-time KV-cache append, cascade / prefix-cache attention-state merge, and sort-free top-K / top-P / min-P sampling. The canonical vLLM-style serving surface for the baracuda CUDA stack. Apache-2.0 (FlashInfer upstream).

0 audits · 6 versions · updated 2026-06-20

Raw C-ABI FFI surface for the vendored FlashInfer inference kernels (paged-KV decode, paged-KV append, cascade state merge, sort-free sampling). The launcher shims + vendored FlashInfer sources are compiled by `baracuda-kernels-sys`; this crate is a thin re-export facade so the FlashInfer C-ABI is reachable under its own crate name. Apache-2.0 (FlashInfer upstream) — see `crates/baracuda-kernels-sys/vendor/flashinfer/`.

0 audits · 6 versions · updated 2026-06-20

Build-time CUDA kernel compiler for the baracuda ecosystem: nvcc-driven incremental builds, parallel compilation, GPU auto-detection, and CUTLASS / custom git dependency support.

0 audits · 45 versions · updated 2026-06-20

Driver-free classifier vocabulary for the baracuda kernel facade: the pure-data kernel-dispatch tags (ElementKind, ArchSku, OpCategory, layout / op-family tags), the StructureKey / OperandDesc classifier key + `structure_key` derivation, the dispatch-table types, and the PlanPreference / PrecisionGuarantee descriptors. Carved out of baracuda-kernels-types so neutral consumers (the kernel generator, selectors, Fuel) depend on the vocabulary WITHOUT the CUDA driver (baracuda-driver -> baracuda-cuda-sys) transitive pull. The device-view half (MatrixRef / TensorRef / Workspace + the tensor->OperandDesc adapter) stays in baracuda-kernels-types, which re-exports this crate so its public API is unchanged. Seeds the ThinkersJournal KISS (Kernel Interface Standards Suite) classifier-vocabulary reference implementation.

0 audits · 0 versions

Kernel generator + the live JIT synthesizer for the Fuel kernel seam. The SUPPORTED surface is the `seam` feature (BaracudaSynthesizer, the fuel-kernel-seam Synthesizer impl); the generator internals (IR, emitters, contracts, dispatch artifacts) are alpha-fluid — pin exact versions.

0 audits · 0 versions

Unified ML op facade for the baracuda CUDA ecosystem. Exposes every primitive an ML framework would expect (union of PyTorch torch.* + nn.functional and JAX lax.* / numpy ops) through a single Plan-based Rust surface, internally dispatching to baracuda-cutlass, the baracuda-* NVIDIA-library wrappers, or bespoke baracuda-kernels-sys kernels.

0 audits · 42 versions · updated 2026-06-20

Compiled bespoke .cu kernel template instantiations for the baracuda ML kernel facade plus C-ABI FFI facades for the library-backed plans (cuDNN conv/pool, cuSOLVER linalg, cuFFT/cuRAND, CUTLASS GEMM re-export). Hosts curated CUDA kernel sources (int8/FP8/int4/bin GEMM RRR, elementwise, reduce, norm, attention, …), builds them via baracuda-forge, exposes extern "C" entry points for the safe baracuda-kernels crate. CUTLASS template kernels live in the sibling baracuda-cutlass-kernels-sys crate and are re-exported here under the unified baracuda_kernels_gemm_* namespace.

0 audits · 42 versions · updated 2026-06-20

Shared type vocabulary for the baracuda ML kernel facade: Element / IntElement / FpElement / BiasElement trait hierarchy, layout / epilogue / activation tags, MatrixRef / TensorRef views, PlanPreference, PrecisionGuarantee, and Workspace. Lifted from baracuda-cutlass so that baracuda-kernels and the per-library wrapper crates can share one vocabulary.

0 audits · 42 versions · updated 2026-06-20

Megatron-LM-style tensor-parallel primitives (Column / Row Parallel Linear) for the baracuda CUDA stack. Pure-composition crate — local GEMM via baracuda-cublas + cross-rank collectives via baracuda-nccl. No new CUDA kernels. NEW in Phase 57; deliberate scope expansion (distributed-training-framework-adjacent). Off-by-default in baracuda-kernels via the `megatron_tp` cargo feature so non-distributed consumers don't pay the dep surface cost. Algorithmic reference: Shoeybi et al. arXiv:1909.08053 (NVIDIA Megatron-LM, Apache-2.0).

0 audits · 12 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA NCCL (multi-GPU collective communication).

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA NCCL (multi-GPU collective communication).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA NPP (Performance Primitives). Core + signal subset at v0.1.

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA NPP (Performance Primitives).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA nvCOMP (GPU compression). Scaffolding at v0.1.

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA nvCOMP (GPU compression).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA nvImageCodec (unified GPU image decode).

0 audits · 6 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA nvImageCodec (unified GPU image codecs).

0 audits · 6 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA nvJitLink (CUDA 12.0+ JIT linker).

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA nvJitLink (CUDA 12.0+ JIT linker).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA nvJPEG (GPU JPEG decode).

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA nvJPEG.

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for the NVIDIA Management Library (NVML) — driver-bundled GPU monitoring.

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for the NVIDIA Management Library (NVML).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA NVRTC (compile CUDA C++ to PTX at runtime).

0 audits · 49 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA NVRTC (runtime CUDA-C++ to PTX compiler).

0 audits · 49 versions · updated 2026-06-20

Safe Rust wrappers for NVIDIA NVSHMEM host API (symmetric-heap one-sided GPU communication).

0 audits · 6 versions · updated 2026-06-20

Raw FFI bindings and dynamic loader for NVIDIA NVSHMEM host library (OpenSHMEM symmetric-heap one-sided comms on GPUs).

0 audits · 6 versions · updated 2026-06-20

Optimizer kernels (Adam / LAMB / SGD) for the baracuda CUDA stack, built on the multi_tensor_apply idiom vendored from NVIDIA Apex (BSD-3-Clause). One launch over thousands of parameter tensors — critical for the optimizer step on large-model training stacks. NEW in Phase 49; deliberate scope expansion (training-framework-adjacent). Off-by-default in baracuda-kernels via the `optim` cargo feature so inference-only consumers don't pay the FFI surface cost.

0 audits · 12 versions · updated 2026-06-20

Safe Rust wrapper for baracuda's clean-fork of Hiroyuki Ootomo's ozIMMU — Ozaki-scheme FP64 GEMM that synthesizes a DGEMM from S^2 int8 tensor-core matmuls. Provides an RAII handle + drop-in `dgemm` shim suitable for the `BackendKind::Ozaki` path of `baracuda-kernels`'s `GemmPlan`. Opt-in (NOT bit-equivalent to native DGEMM); the default FP64 path stays on CUTLASS / cuBLAS.

0 audits · 12 versions · updated 2026-06-20

Build + raw FFI bindings to baracuda's clean-fork of Hiroyuki Ootomo's ozIMMU — the Ozaki-scheme FP64 GEMM library that synthesizes a DGEMM from S² int8 tensor-core matmuls. Phase 44b internalized the upstream sources under `cuda/` (no more `vendor/` subdir; cutf submodule eliminated). Linked statically into the baracuda CUDA stack; consumed by the safe wrapper crate `baracuda-ozimmu`. MIT-licensed (original ozIMMU MIT — see `ATTRIBUTION.md`).

0 audits · 12 versions · updated 2026-06-20