Use Noeris's kernel performance database (110+ shape buckets, 3 GPUs) as a hardware cost model for NAS. What attention head_dim / FFN ratio / norm placement is fastest given real kernel measurements? Nobody has closed this loop with real Triton data. This flips the problem from "optimize kernels for models" to "design models for kernels."
Use Noeris's kernel performance database (110+ shape buckets, 3 GPUs) as a hardware cost model for NAS. What attention head_dim / FFN ratio / norm placement is fastest given real kernel measurements? Nobody has closed this loop with real Triton data. This flips the problem from "optimize kernels for models" to "design models for kernels."