forked from ml-explore/mlx-swift-lm
-
Notifications
You must be signed in to change notification settings - Fork 7
Pull requests: Layr-Labs/mlx-swift-lm
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Qwen4Exp: incremental QSA indexer and gathered sparse attention past the budget
#151
opened Sep 16, 2026 by
davidtai
Loading…
Preserve optional function strict metadata in OpenAI tools
#147
opened Sep 12, 2026 by
jonathan308
Loading…
4 tasks done
Add packed paged KV formats with bounded workspace admission
#142
opened Sep 7, 2026 by
Gajesh2007
Member
•
Draft
Gemma 4 26B-A4B: base port + MTP speculative decoding (serial path)
#138
opened Sep 5, 2026 by
davidtai
Loading…
Merge upstream mlx-swift-lm while preserving Darkbloom runtime support
#124
opened Aug 26, 2026 by
glg2672
Loading…
perf(cbv2): gate Qwen D256 attention execution
#109
opened Aug 17, 2026 by
Gajesh2007
Member
Loading…
Laguna XS 2.1 DFlash: checkpoint conversion + laguna_xs draft support + Laguna target
#98
opened Jul 26, 2026 by
anupsv
Loading…
DFlash framework evolution: draft model, verify path, batched engines, multi-target support
#97
opened Jul 26, 2026 by
anupsv
Loading…
feat: pipeline-parallel model shards for GPT-OSS and Gemma 4
#53
opened Jun 25, 2026 by
crypt0fairy
•
Draft
feat(scheduler): emit a prefill-start admission RequestOutput at admit
#49
opened Jun 23, 2026 by
Gajesh2007
Member
Loading…
Add LlamaModelTP: tensor-parallel variant of LlamaModel
#25
opened May 21, 2026 by
anupsv
Loading…
3 of 4 tasks
Add Llama callPartial for pipeline-parallel inference
#24
opened May 21, 2026 by
anupsv
Loading…
2 of 3 tasks
perf: MLP fusion + Gemma/GPT-OSS inference optimizations
#17
opened May 11, 2026 by
0xClandestine
Member
Loading…
6 tasks
feat: TurboQuant+ KV cache compression
#16
opened May 11, 2026 by
0xClandestine
Member
Loading…
6 tasks
Previous Next
ProTip!
Add no:assignee to see everything that’s not assigned.