RightNow
Popular repositories Loading
-
autokernel
autokernel PublicAutoresearch for GPU kernels. Give it any PyTorch model, go to sleep, wake up to optimized Triton kernels.
-
AutoMegaKernel
AutoMegaKernel PublicAn agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682
-
qwen3.5-triton
qwen3.5-triton PublicPure Triton kernels for Qwen3.5-27B inference on NVIDIA B200
Repositories
- runinfra-cli Public
- inkling-turbo Public
Faster attention kernels for serving TML's Inkling model on vLLM. 2.7x over the shipping path on H100, and the only implementation that runs on A100.
- sglang Public Forked from sgl-project/sglang
SGLang is a high-performance serving framework for large language models and multimodal models.
- Memoir Public
Weights that never freeze. A continual-learning architecture, and the pre-registered experiment that falsified its central claim. https://arxiv.org/abs/2607.20792
- autotree Public
Tree execution engine for LLM inference: fork, merge, prune KV cache at token granularity
- bonsai-turbo Public
Single-launch batch-1 decode engine for PrismML Bonsai 27B (ternary and 1-bit) on NVIDIA GPUs. 1.76x the vendor llama.cpp fork on H100, same outputs.
- auto Public
the agi compiler: records llm agent behavior, proves what repeats, and compiles it into verified, sandboxed wasm binaries that run for microdollars. nothing figured out twice, paper: https://arxiv.org/abs/2607.04542
- AutoMegaKernel Public
An agent harness that compiles a model into one provably-correct, self-retargeting CUDA megakernel and self-tunes it past cuBLAS at batch-1 LLM decode, paper: https://arxiv.org/abs/2606.09682
Top languages
Loading…
Most used topics
Loading…