Threads (@myps6415)
local-llmai-tools
Original source

FreeToken — Edge-Native MoE Serving Engine for Consumer Hardware

Summary

FreeToken is an open-source MoE inference engine designed specifically for edge/consumer hardware. Instead of treating laptops as “small GPUs,” it dynamically adapts to actual available resources (memory, bandwidth, CPU/GPU split).

Key Claims

  • 8GB laptop GPU → runs 35B models
  • Gaming desktop (single GPU) → 284B models
  • Workstation single GPU → 753B (GLM-5.2)
  • 20+ MoE models supported (coding agents, tool-use)

Core Tech

  • “Expert residency” + “bandwidth-adaptive execution” (q* policy)
  • System continuously monitors agent workload patterns
  • Dynamically maps model state to available resources instead of fixed offloading strategy
  • Full-layer double-buffered prefill streaming
  • Global LRU expert caching
  • Semantic anchor checkpoints for agentic context (avoid redundant recomputation on tool calls)

Practical Details

  • Desktop app (Windows/Linux) at flashml.ai
  • CLI via uv pip install "freetoken[accel]"
  • Supports: DeepSeek-V4-Flash, Qwen3.6-35B-A3B, GLM-5.2
  • Quantization: MXFP4, NVFP4, FP8, BF16
  • OpenAI/Anthropic-compatible API
  • Works with Codex, Claude Code, OpenCode, OpenClaw, DeepSeek Harness
  • NVIDIA RTX 30/40/50 series
  • Apache 2.0 license

Authors

UC Berkeley / Stanford group (Ion Stoica, Matei Zaharia, Song Han, Kurt Keutzer among authors)

Thread Discussion

Reply asks about actual tokens/sec speed and precision loss — valid concerns not addressed in the post.