Skip to content
AppClap
FastFlowLM is unclaimed —
F

FastFlowLM

The fastest, most efficient LLM inference on NPUs·fastflowlm.com

Visit

FastFlowLM (FLM) is an NPU-first LLM inference runtime built exclusively for AMD Ryzen AI NPUs. It offers an Ollama-style CLI and an OpenAI-compatible API server, supporting text, vision, audio, reasoning, embedding, and MoE models with context windows up to 256k tokens. It ships as a ~16MB installer and runs models locally on-device with over 10x better power efficiency than GPU-first stacks.

What it's for

Run LLMs locally on AMD Ryzen AI NPUsStream tokens via an OpenAI-compatible APIRun vision models to understand and describe imagesTranscribe and summarize long-form audio locallyBuild and run Retrieval-Augmented Generation (RAG) workflows on-deviceRun agentic chains with deterministic latency

Features 13

  • Context windows up to 256k tokens
  • Full-stack telemetry: NPU, CPU, and memory counters
  • NPU-first architecture optimized for AMD Ryzen AI NPUs
  • Ollama-style CLI (flm run, flm list, flm serve)
  • On-device security with local tokens and full offline mode
  • OpenAI-compatible API server
  • Over 10x power efficiency vs GPU-first stacks
  • Retrieval-Augmented Generation (RAG) workflow support
  • Scenario-driven benchmark suites (instruction tuning, RAG, chat, multimodal)
  • Support for models including Qwen3.6-MoE, GPT-OSS, Gemma3 (Vision), Whisper, Llama 3.2, DeepSeek-R1, Qwen3-VL
  • Supports LLMs, Vision, Audio, Reasoning, Embeddings, and MoE models
  • Zero-conf signed installer for Ryzen AI 300 laptops
  • ~16MB runtime size

At a glance

API
TypeCLI tool
DeploymentInstalled (local)
PlatformsWindows, Linux
Fordevelopers
CompanyAdvanced Micro Devices, Inc.

Integrations

Open WebUI

Resources

Last checked 28 days ago·