FastFlowLM is unclaimed —
F
FastFlowLM
The fastest, most efficient LLM inference on NPUs·fastflowlm.com
FastFlowLM (FLM) is an NPU-first LLM inference runtime built exclusively for AMD Ryzen AI NPUs. It offers an Ollama-style CLI and an OpenAI-compatible API server, supporting text, vision, audio, reasoning, embedding, and MoE models with context windows up to 256k tokens. It ships as a ~16MB installer and runs models locally on-device with over 10x better power efficiency than GPU-first stacks.
What it's for
Run LLMs locally on AMD Ryzen AI NPUsStream tokens via an OpenAI-compatible APIRun vision models to understand and describe imagesTranscribe and summarize long-form audio locallyBuild and run Retrieval-Augmented Generation (RAG) workflows on-deviceRun agentic chains with deterministic latency
Features 13
- Context windows up to 256k tokens
- Full-stack telemetry: NPU, CPU, and memory counters
- NPU-first architecture optimized for AMD Ryzen AI NPUs
- Ollama-style CLI (flm run, flm list, flm serve)
- On-device security with local tokens and full offline mode
- OpenAI-compatible API server
- Over 10x power efficiency vs GPU-first stacks
- Retrieval-Augmented Generation (RAG) workflow support
- Scenario-driven benchmark suites (instruction tuning, RAG, chat, multimodal)
- Support for models including Qwen3.6-MoE, GPT-OSS, Gemma3 (Vision), Whisper, Llama 3.2, DeepSeek-R1, Qwen3-VL
- Supports LLMs, Vision, Audio, Reasoning, Embeddings, and MoE models
- Zero-conf signed installer for Ryzen AI 300 laptops
- ~16MB runtime size
At a glance
API
TypeCLI tool
DeploymentInstalled (local)
PlatformsWindows, Linux
Fordevelopers
CompanyAdvanced Micro Devices, Inc.
Integrations
Open WebUI
Resources