Clone this repo:
  1. 270ab2f refactor: simplify fallback Jinja templates in model_type_utils.cc by Wai Hon Law · 37 minutes ago upstream/main
  2. 2e1941b Add Decode method to ImagePreprocessor and use it in MiniCPM-V image preprocessing. by Yi-Chun Kuo · 60 minutes ago
  3. 4e3d7d1 Add set_activation_data_type support to LiteRT-LM embedding runtime. by Andrew Zhang · 2 hours ago
  4. 13bb895 Add support for lazy loading multimodal encoders in EmbeddingEngine. by Yi-Chun Kuo · 2 hours ago
  5. dc911c3 Return an error during signature selection when target capacity exceeds available signature lengths in LiteRT-LM embedding engine. by Matthew Chan · 4 hours ago

LiteRT-LM

LiteRT-LM is Google's production-ready orchestration layer to run LLMs with LiteRT, engineered for high-performance, cross-platform execution.

🔗 Product Website | 🌐✨ Web Demo

🔥 What's New: v0.16.0

This release is a quick follow up to v0.15.0 (which brought Apple Foundation Framework integration, CLI configuration, and JavaScript API Updates).

  • 📦 C API Prebuilts: Added the first versioned C API shared library prebuilts for all supported platforms. This allows natively integrating LiteRT-LM into your applications and creating language bindings without the hassle of building shared libraries.
  • 🚀 Experimental YNNPACK Delegate: Added the experimental YNNPACK delegate, enabled for linux arm64 builds in the LiteRT-LM CLI and Python API.

👉 Try Gemma4-E4B with MTP on Linux, macOS, Windows or Raspberry Pi with the LiteRT-LM CLI:

litert-lm run  \
   --from-huggingface-repo=litert-community/gemma-4-E4B-it-litert-lm \
   gemma-4-E4B-it.litertlm \
   --backend=gpu \
   --enable-speculative-decoding=true \
   --prompt="What is the capital of France?"

🌟 Key Features

  • 📱 Cross-Platform Support: Android, iOS, Web, Desktop, and IoT (e.g. Raspberry Pi).
  • 🚀 Hardware Acceleration: Peak performance via GPU and NPU accelerators.
  • 👁️ Multi-Modality: Support for vision and audio inputs.
  • 🔧 Tool Use: Function calling support for agentic workflows.
  • 📚 Broad Model Support: Gemma, Llama, Phi-4, Qwen, and more.


🚀 Production-Ready for Google's Products

LiteRT-LM powers on-device GenAI experiences in Chrome, Chromebook Plus, Pixel Watch, and more.

You can also try the Google AI Edge Gallery app to run models immediately on your device.

Install the app today from Google PlayInstall the app today from App Store

📰 Blogs & Announcements

LinkDescription
Bringing Gemma 4 12B to your Laptop: Unlocking Local, Agentic Workflows with Google AI EdgeBring agentic, multimodal AI capabilities to everyday laptops, enabling local data processing and visual insight generation.
Blazing-fast on-device GenAI with LiteRT-LMUnlock Gemma 4's full potential with blazing speed and incredible efficiency using newly added Swift, JavaScript, and Flutter APIs.
Accelerating Gemma 4: faster inference with multi-token prediction draftersAn overview of how Multi-Token Prediction (MTP) drafters are making Gemma 4 models up to 3x faster at inference.
Bring state-of-the-art agentic skills to the edge with Gemma 4Deploy Gemma 4 in-app and across a broader range of devices with stellar performance and broad reach using LiteRT-LM.
On-device GenAI in Chrome, Chromebook Plus and Pixel WatchDeploy language models on wearables and browser-based platforms using LiteRT-LM at scale.
On-device Function Calling in Google AI Edge GalleryExplore how to fine-tune FunctionGemma and enable function calling capabilities powered by LiteRT-LM Tool Use APIs.
Google AI Edge small language models, multimodality, and function callingLatest insights on RAG, multimodality, and function calling for edge language models.

🏃 Quick Start

🔗 Key Links

⚡ Quick Try (No Code)

Try LiteRT-LM immediately from your terminal without writing a single line of code using uv:

uv tool install litert-lm

litert-lm run \
  --from-huggingface-repo=google/gemma-3n-E2B-it-litert-lm \
  gemma-3n-E2B-it-int4 \
  --prompt="What is the capital of France?"

📚 Supported Language APIs

Ready to get started? Explore our language-specific guides and setup instructions.

LanguageStatusBest For...Documentation
Python✅ StablePrototyping & ScriptingPython Guide
Kotlin✅ StableAndroid apps & JVMKotlin Guide
Swift🚀 Early PreviewNative iOS & macOSSwift Guide
JavaScript (web)🚀 Early PreviewBrowser environmentsJavaScript Guide
Flutter🚀 CommunityCross-platform mobileFlutter Guide
C++✅ StableHigh-performance nativeC++ Guide

🏗️ Build From Source

This guide shows how you can compile LiteRT-LM from source. If you want to build the program from source, you should checkout the stable Latest Release tag.