Runanywhere Sdks

Production ready toolkit to run AI locally

Last verified:

Visit Runanywhere Sdks

What is Runanywhere Sdks?

Runanywhere Sdks is a privacy-first, on-device AI SDK that enables developers to run large language models directly on iOS, Android, React Native, and Flutter applications without requiring cloud connectivity for inference. The SDK provides a unified API for text generation (LLM), speech-to-text (STT), text-to-speech (TTS), voice activity detection (VAD), and tool calling, all processing data locally on the device to ensure maximum privacy and minimal latency.

Runanywhere Sdks pricing

Pricing model: Freemium

RunAnywhere offers a free tier to get started for developers and teams experimenting with on-device AI. The open-source SDK is free with iOS, Android, React Native, Flutter support and includes MetalRT runtime and on-device inference. The Enterprise Control Plane with OTA model updates, policy-based routing, fleet analytics, and hybrid cloud fallback requires contact for pricing. Some sources indicate subscription plans starting at $49/mo. Enterprise pricing details are not explicitly outlined publicly and are likely based on usage and support requirements.

Runanywhere Sdks pros

  • Runs AI models entirely on-device with no network required for inference
  • Zero marginal inference cost once models are downloaded
  • <7ms time-to-first-token on M4 Max with Qwen3-0.6B
  • 668 tok/s LLM decode speed on Apple Silicon with MetalRT
  • Privacy-first architecture keeps all data on device
  • Cross-platform support: Swift, Kotlin, React Native, Flutter
  • Smart hybrid routing automatically switches between on-device and cloud
  • Built-in model management with downloading, caching, and lifecycle control
  • Streaming token support via async iterators for LLM
  • Multi-backend architecture with LlamaCPP and ONNX Runtime options
  • Real-time voice activity detection with Silero VAD
  • Whisper-based speech-to-text with multi-language support
  • Piper TTS for neural voice synthesis
  • Tool calling with typed tool definitions and parseToolCall
  • Production-ready with built-in analytics, logging, and observability
  • OTA model updates without app store releases
  • Policy-based routing for privacy, cost, and performance control
  • Fleet dashboard for managing thousands of devices
  • Open-source SDK available on GitHub under Apache 2.0 license
  • Backed by Y Combinator with enterprise-ready workflow

Runanywhere Sdks cons

  • Android SDK still under active development, not yet production-ready
  • Apple Silicon devices recommended (M1/M2/M3, A14+) for best performance
  • Android devices need 6GB+ RAM for optimal performance
  • iOS 15.1+ and Android API 24+ minimum requirements limit older devices
  • Voice AI workflow marked as experimental feature
  • Structured outputs feature still experimental
  • Small model quality may not match larger cloud models for complex tasks
  • Enterprise pricing not transparent, requires contact for quotes
  • Limited model ecosystem compared to established cloud platforms
  • Control Plane features require enterprise pricingtier

Frequently asked questions about Runanywhere Sdks

What is RunAnywhere?

RunAnywhere is a privacy-first, on-device AI SDK and control plane that runs large language models directly on mobile devices with a single API. It enables developers to build AI chatbots, voice assistants, and offline AI tools without cloud dependency, using smart routing to intelligently switch between on-device and cloud based on privacy, cost, and performance requirements.

How does on-device AI work with RunAnywhere?

The RunAnywhere SDK analyzes each request and device capabilities, then intelligently routes to on-device processing or cloud based on requirements. All AI inference runs locally by default, ensuring low latency and data privacy. Once models are downloaded, no network connection is required for inference. Sensitive data never leaves the device unless explicitly configured.

What platforms does RunAnywhere support?

RunAnywhere provides cross-platform SDKs for Swift (iOS/macOS), Kotlin (Android), React Native, and Flutter with one unified API. The iOS SDK is production-ready and available now. The Android SDK is under active development. The SDK supports iOS 15.1+, macOS 10.15+, Android API 24+, React Native 0.74+.

What AI capabilities does the SDK provide?

The SDK provides LLM text generation with streaming support, speech-to-text transcription with Whisper models, neural text-to-speech with Piper TTS, real-time voice activity detection with Silero VAD, and tool calling with typed tool definitions. It also supports full voice agent pipeline with VAD to STT to LLM to TTS orchestration.

What is MetalRT?

MetalRT is RunAnywhere's custom GPU inference engine with hand-written Metal shaders for Apple Silicon. It achieves 668 tok/s LLM decode, 101ms speech-to-text latency, and 287 tok/s vision inference on a single MacBook. Every kernel is hand-designed with custom memory layouts and fused operators, bypassing generic abstraction layers for record-setting speeds on Apple Silicon.

Is RunAnywhere free to use?

RunAnywhere offers a free tier to get started, making it accessible for developers experimenting with on-device AI. The open-source SDK is free on GitHub under Apache 2.0 license. Enterprise Control Plane features with fleet dashboard, OTA updates, and policy-based routing require contact for pricing, with some sources indicating plans starting at $49/mo.

How does RunAnywhere handle privacy?

RunAnywhere uses privacy-by-design architecture where all processing happens on-device by default. Audio and text data never leaves the device unless explicitly configured. Only anonymous analytics are collected by default. The SDK offers strict privacy mode for on-device-only processing, ensuring complete privacy for sensitive applications.

What models does RunAnywhere support?

RunAnywhere supports GGUF models via LlamaCPP backend, ONNX Runtime models for speech pipelines, and Apple Foundation Models on iOS 18+. The SDK supports small models like Qwen3-0.6B that now match quality of models 250x their size. Model delivery includes downloading with resume support, extraction, and storage management.

What is the Control Plane?

The Control Plane is RunAnywhere's enterprise-ready fleet dashboard for managing on-device AI at scale. It provides OTA model updates without app store releases, policy-based routing to define when to run locally vs cloud, inference analytics and performance dashboards, and the ability to manage thousands of devices. It turns edge AI into an enterprise workflow.

Who founded RunAnywhere?

RunAnywhere was founded in 2025 by Sanchit Monga (Co-Founder & CEO) and Shubham Malhotra (Co-Founder & CTO). Sanchit built SDKs used by 50M+ users at Intuit and leads product and go-to-market. Shubham formerly worked on AWS EC2 Spot and Microsoft Azure Arc, is a published ML researcher, and writes the custom kernels powering MetalRT. The company is backed by Y Combinator (YC W26).

Categories

Use cases

Browse all AI tools on NeedAnAI