FreeToken — Edge-Native MoE Serving Engine for Consumer Hardware
- URL: https://www.threads.com/share/_wJroozaw/
- Date Saved: 2026-08-21
- Source: Threads (@myps6415)
- Tags: local-llm, ai-tools
- Repo: https://github.com/FlashML-org/FreeToken (⭐98)
- Paper: https://arxiv.org/abs/2608.16157
Summary
FreeToken is an open-source MoE inference engine designed specifically for edge/consumer hardware. Instead of treating laptops as “small GPUs,” it dynamically adapts to actual available resources (memory, bandwidth, CPU/GPU split).
Key Claims
- 8GB laptop GPU → runs 35B models
- Gaming desktop (single GPU) → 284B models
- Workstation single GPU → 753B (GLM-5.2)
- 20+ MoE models supported (coding agents, tool-use)
Core Tech
- “Expert residency” + “bandwidth-adaptive execution” (q* policy)
- System continuously monitors agent workload patterns
- Dynamically maps model state to available resources instead of fixed offloading strategy
- Full-layer double-buffered prefill streaming
- Global LRU expert caching
- Semantic anchor checkpoints for agentic context (avoid redundant recomputation on tool calls)
Practical Details
- Desktop app (Windows/Linux) at flashml.ai
- CLI via
uv pip install "freetoken[accel]" - Supports: DeepSeek-V4-Flash, Qwen3.6-35B-A3B, GLM-5.2
- Quantization: MXFP4, NVFP4, FP8, BF16
- OpenAI/Anthropic-compatible API
- Works with Codex, Claude Code, OpenCode, OpenClaw, DeepSeek Harness
- NVIDIA RTX 30/40/50 series
- Apache 2.0 license
Authors
UC Berkeley / Stanford group (Ion Stoica, Matei Zaharia, Song Han, Kurt Keutzer among authors)
Thread Discussion
Reply asks about actual tokens/sec speed and precision loss — valid concerns not addressed in the post.