Highlights
- Pro
Pinned Loading
-
Agentic-Kernel-Generation
Agentic-Kernel-Generation PublicLLM agents that generate, verify, and evolve Triton GPU kernels. Includes a reward-hack-resistant benchmarking harness with strict correctness verification and fresh-input evaluation. Achieves up t…
-
Parallel-Eagle
Parallel-Eagle PublicA lossless speculative decoder: a parallel multi-token drafter + dynamic tree verification.
Python
-
DiSpec
DiSpec PublicFrom-scratch LLM inference engine for Qwen2.5/Qwen3: paged attention with prefix caching, continuous batching, a Triton decode kernel, CUDA-graph decode, FlashAttention prefill, speculative decodin…
Python
-
kvxfer
kvxfer PublicReuse a small model's prefill in a larger same-family model by mapping its KV cache instead of re-prefilling. Closed-form per-layer maps plus an optional trained residual — 4.6–9.1× faster than vLL…
Python
-
Schema-Linking-with-SFT
Schema-Linking-with-SFT PublicLoRA fine-tuned Qwen2.5-Coder-1.5B for schema linking in NL-to-SQL: given a question and a database schema, predicts the referenced tables and columns as JSON. Uses self-consistency decoding (k=10)…
Python
-
rag-context-engineering
rag-context-engineering PublicRetrieval-augmented generation for documentation QA under a 2,000-token context budget — multi-resolution chunking, hybrid BM25+FAISS retrieval, cross-encoder reranking, and grounded generation, w…
Python
If the problem persists, check the GitHub status page or contact support.

