File size: 1,372 Bytes
35bf405
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
---
license: apache-2.0
tags:
- cuda
- kernel
- gpu-optimization
- hpc
---

# ada_tensor_core_bf16

Ada Tensor Core BF16 GEMM prototype with WMMA-backed 32x64x16 correctness and timing harness.

This repository contains the standalone CUDA source for the `ada_tensor_core_bf16` lane from
 the PyC kernel lab. It is a source artifact for inspection and benchmarking;
it is not a precompiled binary and the result below is not a universal ranking.

## Performance

| Kernel | GPU / architecture | Shape | Best recorded result | Evidence |
|---|---|---|---|---|
| `ada_tensor_core_bf16` | not recorded | not recorded | Not measured in the published campaign | No published performance receipt was found for this lane. |

![Performance plot](performance.svg)

The result is reported with the original campaign's timing and correctness
context. Compare kernels only when GPU, CUDA version, matrix shape, warmup,
repeats, and reference/correctness mode match.

## Source

- `kernel.cu` — copied from `kernels/prototypes/ada/tensor_core/kernel.cu`.
- Original lane tags: `cuda, matmul, ada, sm89, prototype, tensor-core, bf16`.

## Build/run contract

```text
{nvcc} -O3 -std=c++17 -lineinfo -DPYC_ADA_TENSOR_CORE_USE_BF16=1 -gencode arch=compute_89,code=sm_89 -gencode arch=compute_89,code=compute_89 {source} -o {build_dir}/{name}
{build_dir}/{name} 1024 1024 1024 10 50
```