Better kernels for real workloads.

Bring a GPU bottleneck. Set a reward and a deadline. Give engineers a clear problem to solve, and receive source you can put to the test.

A useful result needs more than a faster benchmark: correct outputs, reproducible measurements, and a fit for your production stack. Set those requirements before the work begins.

Looking to optimize? Browse open bounties.

Kernel optimization workflowA workload grid branches into source candidate lanes, moves through reproduction and inspection gates, then converges into a handoff stack.WORKLOADSOURCE CANDIDATESREPRODUCEINSPECTHANDOFF
  1. 01DefineFreeze the workload, environment, and scoring contract.
  2. 02OptimizeDevelop focused kernel candidates against one baseline.
  3. 03EvaluateReproduce results and inspect integration requirements.
  4. 04IntegrateHand off source, evidence, and implementation notes.

Illustrative workflow. Evaluation and payouts are arranged manually.

CUDATritonPyTorch

Bounty board

Open work

0 posted
No bounties posted yet
The board starts with a real problem and rules a sponsor is willing to stand behind.